Encoding method and decoding method, encoder and decoder, and storage medium
By incorporating subblock-based spatial motion vector prediction in video coding, spatial correlation is utilized to enhance compression efficiency, addressing the limitations of existing techniques and improving decoding and encoding processes.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-07-23
AI Technical Summary
Existing video coding techniques fail to effectively utilize spatial correlation in subblock-based prediction, limiting the efficiency of video compression.
Implementing subblock-based spatial motion vector prediction (SbSMVP) to determine merge candidates based on spatial neighboring blocks, enhancing the decoding and encoding processes.
Improves video coding efficiency by leveraging spatial correlation, leading to more effective compression and decoding of high-quality video data.
Smart Images

Figure CN2025141198_23072026_PF_FP_ABST
Abstract
Description
ENCODING METHOD AND DECODING METHOD, ENCODER AND DECODER, AND STORAGE MEDIUMCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 746,198, filed January 16, 2025, entitled “ENCODING METHOD AND DECODING METHOD, ENCODER AND DECODER, AND STORAGE MEDIUM, ” which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Embodiments of the present disclosure relate to video coding.
[0003] Digital video has become mainstream and is being used in a wide range of applications including digital television, video telephony, and teleconferencing. These digital video applications are feasible because of the advances in computing and communication technologies as well as efficient video coding techniques. Various video coding techniques may be used to compress video data, such that coding on the video data may be performed using one or more video coding standards. Exemplary video coding standards may include, but not limited to, versatile video coding (H. 266 / VVC) , high-efficiency video coding (H. 265 / HEVC) , advanced video coding (H. 264 / AVC) , moving picture expert group (MPEG) coding, enhanced video coding model (ECM) , to name a few.SUMMARY
[0004] According to one aspect of the present disclosure, a method of decoding is provided. The method may include determining, by a processor, a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block. The method may include determining, by the processor, a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The method may include decoding, by the processor, the current block based on the subblock-based merge candidate list.
[0005] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The memory storing instructions, which when executed by the processor, may cause the processor to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The memory storing instructions, which when executed by the processor, may cause the processor to decode the current block based on the subblock-based merge candidate list.
[0006] According to a further aspect of the present disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The memory storing instructions, which when executed by the processor, may cause the processor to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The memory storing instructions, which when executed by the processor, may cause the processor to decode the current block based on the subblock-based merge candidate list.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a processor of a decoder is provided. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to decode the current block based on the subblock-based merge candidate list.
[0008] According to one aspect of the present disclosure, a method of encoding is provided. The method may include determining, by a processor, an SbSMVP candidate based on a spatial neighboring block of a current block. The method may include determining, by the processor, a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The method may include encoding, by the processor, the current block based on the subblock-based merge candidate list.
[0009] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The memory storing instructions, which when executed by the processor, may cause the processor to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The memory storing instructions, which when executed by the processor, may cause the processor to encode the current block based on the subblock-based merge candidate list.
[0010] According to a further aspect of the present disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The memory storing instructions, which when executed by the processor, may cause the processor to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The memory storing instructions, which when executed by the processor, may cause the processor to encode the current block based on the subblock-based merge candidate list.
[0011] According to another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a processor of an encoder is provided. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to determine an SbSMVP candidate based on a spatial neighboring block of a current block. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to encode the current block based on the subblock-based merge candidate list.
[0012] According to a further aspect of the present disclosure, a method of transmitting a bitstream is provided. The method may include executing the encoding method described herein to generate a bitstream. The method may include transmitting the bitstream.
[0013] According to still another aspect of the present disclosure, a non-transitory computer-readable storage medium, having a computer program and a bitstream stored thereon is provided. The computer program, when executed by a processor, enables the processor to perform the operations of the encoding method described herein to generate the bitstream.
[0014] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding thereof. Additional embodiments are described in the Detailed Description, and further description is provided there.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the pertinent art to make and use the present disclosure.
[0016] FIG. 1A illustrates a block diagram of an exemplary encoding system, according to some embodiments of the present disclosure.
[0017] FIG. 1B illustrates a block diagram of an exemplary decoding system, according to some embodiments of the present disclosure.
[0018] FIG. 2 illustrates a block diagram of an exemplary encoder, according to some embodiments of the present disclosure.
[0019] FIG. 3A illustrates an exemplary technique of quadtree splitting of a coding unit, according to some embodiments of the present disclosure.
[0020] FIG. 3B illustrates an exemplary technique of binary splitting and ternary splitting of a coding unit, according to some embodiments of the present disclosure.
[0021] FIG. 3C illustrates an exemplary technique of splitting of a coding unit into various split types, according to some embodiments of the present disclosure.
[0022] FIG. 4 illustrates an exemplary technique of inter prediction based on template matching, according to some embodiments of the present disclosure.
[0023] FIG. 5A illustrates first exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0024] FIG. 5B illustrates second exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0025] FIG. 5C illustrates third exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0026] FIG. 6 illustrates an exemplary technique of intra prediction based on template matching, according to some embodiments of the present disclosure.
[0027] FIGs. 7A-7E illustrate example syntax elements for signaling controlling parameters for template based prediction modes, according to some embodiments of the present disclosure.
[0028] FIGs. 8A-8G illustrate example syntax elements for signaling controlling parameters for template based prediction modes, according to some embodiments of the present disclosure.
[0029] FIGs. 9A-9E illustrate example syntax elements for signaling controlling parameters for template based prediction modes, according to some embodiments of the present disclosure.
[0030] FIG. 10 illustrates an example implementation of decoder, according to some embodiments of the present disclosure.
[0031] FIG. 11 illustrates an example of source device, according to some embodiments of the present disclosure.
[0032] FIG. 12 illustrates an example of receiving device, according to some embodiments of the present disclosure.
[0033] FIG. 13 illustrates an example of communication system, according to some embodiments of the present disclosure.
[0034] FIG. 14 illustrates an example of video codec system, according to some embodiments of the present disclosure.
[0035] FIG. 15 illustrates an example of communication system, according to some embodiments of the present disclosure.
[0036] FIG. 16 illustrates an example of current block and reference sample, according to some embodiments of the present disclosure.
[0037] FIG. 17 illustrates a diagram of non-adjacent spatial neighboring candidates for occurrence-based intra coding (OBIC) mode, according to some embodiments of the present disclosure.
[0038] FIG. 18 illustrates an example of the histogram-of-occurrences (HoC) of intra prediction mode, according to some embodiments of the present disclosure.
[0039] FIG. 19 illustrates an example of region transform of current block, according to some embodiments of the present disclosure.
[0040] FIG. 20A illustrates a first example of region transform of current block, according to some embodiments of the present disclosure.
[0041] FIG. 20B illustrates a second example of region transform of current block, according to some embodiments of the present disclosure.
[0042] FIG. 21 illustrates an example of region transform of current block, according to some embodiments of the present disclosure.
[0043] FIG. 22 illustrates an example of adjacent blocks of a current block, according to some embodiments of the present disclosure.
[0044] FIG. 23 illustrates a first example of an adjacent block of block vector (BV) -based prediction mode, according to some embodiments of the present disclosure.
[0045] FIG. 24 illustrates a second example of an adjacent block of BV-based prediction mode, according to some embodiments of the present disclosure.
[0046] FIG. 25 is a flowchart of an example method of predicting a current block using a BV-based technique, according to some embodiments of the present disclosure.
[0047] FIG. 26 illustrates an example of spatial merge candidates of a current block, according to some embodiments of the present disclosure.
[0048] FIG. 27 illustrates an example of block vector (BV) refinement, according to some embodiments of the present disclosure.
[0049] FIG. 28A illustrates a first example of BV flip, according to some embodiments of the present disclosure.
[0050] FIG. 28B illustrates a second example of BV flip, according to some embodiments of the present disclosure.
[0051] FIG. 29 illustrates an example of NN-based intra prediction, according to some embodiments of the present disclosure.
[0052] FIG. 30 is a flow chart of an example method of extrapolation filter-based intra prediction (EIP) prediction, according to some embodiments of the present disclosure.
[0053] FIG. 31 illustrates a first example of EIP filter shapes, according to some embodiments of the present disclosure.
[0054] FIG. 32 illustrates a second example of EIP filter shapes, according to some embodiments of the present disclosure.
[0055] FIG. 33 illustrates a third example of EIP filter shapes, according to some embodiments of the present disclosure.
[0056] FIG. 34 illustrates a fourth example of EIP filter shapes, according to some embodiments of the present disclosure.
[0057] FIG. 35 illustrates a fifth example of EIP filter shapes, according to some embodiments of the present disclosure.
[0058] FIG. 36 illustrates a sixth example of EIP filter shapes, according to some embodiments of the present disclosure.
[0059] FIG. 37 illustrates example templates of a reconstructed area used to derive EIP filter parameters, according to some embodiments of the present disclosure.
[0060] FIG. 38 illustrates an example of a block vector-guided EIP (bvEIP) mode, according to some embodiments of the present disclosure.
[0061] FIG. 39 illustrates an example of a reconstructed area, according to some embodiments of the present disclosure.
[0062] FIG. 40 illustrates an example of subblock template generation of subblock-based temporal motion vector prediction (SbTMVP) , according to some embodiments of the present disclosure.
[0063] FIG. 41 illustrates an example of subblock-based spatial motion vector prediction (SbSMVP) candidate types, according to some embodiments of the present disclosure.
[0064] FIG. 42A illustrates a first example Sobel filter, according to some embodiments of the present disclosure.
[0065] FIG. 42B illustrates a second example Sobel filter, according to some embodiments of the present disclosure.
[0066] FIG. 42C illustrates a first example Edge filter, according to some embodiments of the present disclosure.
[0067] FIG. 42D illustrates a second example Edge filter, according to some embodiments of the present disclosure.
[0068] FIG. 43 illustrates a flowchart of an example method of decoding, according to some embodiments of the present disclosure.
[0069] FIG. 44 illustrates a flowchart of an example method of encoding, according to some embodiments of the present disclosure.
[0070] Embodiments of the present disclosure may be described with reference to the accompanying drawings.DETAILED DESCRIPTION
[0071] Although some configurations and arrangements are discussed, it should be understood that this is done for illustrative purposes only. A person skilled in the pertinent art will recognize that other configurations and arrangements may be used without departing from the spirit and scope of the present disclosure. It may be apparent to a person skilled in the pertinent art that the present disclosure can also be employed in a variety of other applications.
[0072] It is noted that references in the specification to “one embodiment, ” “an embodiment, ” “an example embodiment, ” “some embodiments, ” “certain embodiments, ” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of a person skilled in the pertinent art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0073] In general, terminology may be understood at least in part from usage in context. For example, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a, ” “an, ” or “the, ” again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0074] Various aspects of point cloud coding systems will now be described with reference to various apparatus and methods. These apparatus and methods may be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as “elements” ) . These elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends upon the particular application and design constraints imposed on the overall system. The techniques described herein may be used for various point cloud coding applications. As described herein, point cloud coding includes both encoding and decoding a point cloud.
[0075] In August 2020, ITU-T finalized a standardization project namely H. 266 / VVC (Versatile Video Coding) and published the first version of the ITU-T H. 266 standard. Then, the standardization committee started exploration work aiming at achieve a performance superior to the latest H. 266 / VVC standard in coding high quality video with one or more features of high resolution, high frame rate, high bit depth, high dynamic range, wide color gamut and omnidirectional video. JVET (Joint Video Expert Group of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29) is in charge of this exploration work. Various prediction and transform modes have been verified to achieve high compression efficiency in coding high quality video and thus adopted in a software platform for this exploration work.
[0076] Currently, subblock-based prediction method considers temporal correlation. Motion information of a collocated block in a reference picture may be considered for subblock-based prediction. There is an unmet need for considering spatial correlation in subblock prediction.
[0077] FIG. 1A illustrates a block diagram of an exemplary encoding system 100, according to some embodiments of the present disclosure. FIG. 1B illustrates a block diagram of an exemplary decoding system 150, according to some embodiments of the present disclosure. Each system 100 or 150 may be applied or integrated into various systems and apparatuses capable of data processing, such as computers and wireless communication devices. For example, system 100 or 150 may be the entirety or part of a mobile phone, a desktop computer, a laptop computer, a tablet, a vehicle computer, a gaming console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an argument reality (AR) device, or any other suitable electronic devices having data processing capability. As shown in FIGs. 1A and 1B, system 100 or 150 may include a processor 102, a memory 104, and an interface 106. These components are shown as connected one to another by a bus, but other connection types are also permitted. It is understood that system 100 or 150 may include any other suitable components for performing functions described here.
[0078] Processor 102 may include microprocessors, such as graphic processing unit (GPU) , image signal processor (ISP) , central processing unit (CPU) , digital signal processor (DSP) , tensor processing unit (TPU) , vision processing unit (VPU) , neural processing unit (NPU) , synergistic processing unit (SPU) , or physics processing unit (PPU) , microcontroller units (MCUs) , application-specific integrated circuits (ASICs) , field-programmable gate arrays (FPGAs) , programmable logic devices (PLDs) , state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout the present disclosure. Although only one processor is shown in FIGs. 1A and 1B, it is understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Software can include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for instructing hardware are also permitted under the broad category of software.
[0079] Memory 104 can broadly include both memory (a.k.a, primary / system memory) and storage (a.k.a. secondary memory) . For example, memory 104 may include random-access memory (RAM) , read-only memory (ROM) , static RAM (SRAM) , dynamic RAM (DRAM) , ferro-electric RAM (FRAM) , electrically erasable programmable ROM (EEPROM) , compact disc read-only memory (CD-ROM) or other optical disk storage, hard disk drive (HDD) , such as magnetic disk storage or other magnetic storage devices, Flash drive, solid-state drive (SSD) , or any other medium that may be used to carry or store desired program code in the form of instructions that may be accessed and executed by processor 102. Broadly, memory 104 may be embodied by any computer-readable medium, such as a non-transitory computer-readable medium. Although only one memory is shown in FIGs. 1A and 1B, it is understood that multiple memories may be included.
[0080] Interface 106 can broadly include a data interface and a communication interface that is configured to receive and transmit a signal in a process of receiving and transmitting information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although only one memory is shown in FIGs. 1A and 1B, it is understood that multiple interfaces may be included.
[0081] Processor 102, memory 104, and interface 106 may be implemented in various forms in system 100 or 150 for performing point cloud coding functions. In some embodiments, processor 102, memory 104, and interface 106 of system 100 or 150 are implemented (e.g., integrated) on one or more system-on-chips (SoCs) . In one example, processor 102, memory 104, and interface 106 may be integrated on an application processor (AP) SoC that handles application processing in an operating system (OS) environment, including running point cloud encoding and decoding applications. In another example, processor 102, memory 104, and interface 106 may be integrated on a specialized processor chip for point cloud coding, such as a GPU or ISP chip dedicated to graphic processing in a real-time operating system (RTOS) .
[0082] As shown in FIG. 1A, in encoding system 100, processor 102 may include one or more modules, such as an encoder 101. Although FIG. 1A shows that encoder 101 is within one processor 102, it is understood that encoder 101 may include one or more sub-modules that may be implemented on different processors located closely or remotely with each other. Encoder 101 (and any corresponding sub-modules or sub-units) may be hardware units (e.g., portions of an integrated circuit) of processor 102 designed for use with other components or software units implemented by processor 102 through executing at least part of a program, i.e., instructions. The instructions of the program may be stored on a computer-readable medium, such as memory 104, and when executed by processor 102, it may perform a process having one or more functions related to point cloud encoding, such as voxelization, transformation, quantization, arithmetic encoding, etc., as described below in detail.
[0083] Similarly, as shown in FIG. 1B, in decoding system 150, processor 102 may include one or more modules, such as a decoder 120. Although FIG. 2 shows that decoder 120 is within one processor 102, it is understood that decoder 120 may include one or more sub-modules that may be implemented on different processors located closely or remotely with each other. Decoder 120 (and any corresponding sub-modules or sub-units) may be hardware units (e.g., portions of an integrated circuit) of processor 102 designed for use with other components or software units implemented by processor 102 through executing at least part of a program, i.e., instructions. The instructions of the program may be stored on a computer-readable medium, such as memory 104, and when executed by processor 102, it may perform a process having one or more functions related to point cloud decoding, such as arithmetic decoding, dequantization, inverse transformation, reconstruction, synthesis, as described below in detail.
[0084] FIG. 2 illustrates a block diagram of an exemplary encoder 200, according to some embodiments of the present disclosure. As shown, the input to the encoder 200 may be a video including a sequence of pictures or a still picture, while the output of the encoder 200 may be a bitstream representing a compressed version of the input video. As also shown, encoder 200 may include, e.g., a partition unit 201, a prediction unit 202, a block partition unit 203, an inter prediction unit 204, an intra prediction unit 205, a first adder 206, a transform unit 207, a quantization unit 208, an inverse quantization unit 209, an inverse transform unit 210, a second adder 211, a filtering unit 212, and a decoded picture buffer (DPB) 213. The various operations of encoder 200 will now be described.
[0085] For example, referring to FIG. 2, partition unit 201 divides a picture in an input video into one or more coding tree units (CTUs) . Partition unit 201 divides the picture into tiles, and optionally may further divide a tile into one or more bricks. A tile or a brick may contain one or more integral and / or partial CTUs. Partition unit 201 forms one or more slices, where a slice may contain one or more tiles in a raster order of tiles in the picture, or one or more tiles covering a rectangular region in the picture. Partition unit 201 may also form one or more sub-pictures, which may contain one or more slices, tiles, or bricks.
[0086] During the encoding process, partition unit 201 passes CTUs to prediction unit 202. Generally, prediction unit 202 is composed of block partition unit 203, inter prediction unit 204, and intra prediction unit 205. Block partition unit 203 further divides an input CTU into smaller coding units (CUs) using various split or partition types, such as quadtree split, binary split, and ternary split iteratively. Examples of quadtree split, binary split, and ternary split of a CU or a coding block are described below in connection with FIGs. 3A, 3B, and 3C.
[0087] FIG. 3A illustrates an exemplary technique of quadtree splitting 300 of a coding unit, according to some embodiments of the present disclosure. As shown in FIG. 3A, quadtree split is applied to CU or coding block 301. Blocks 3010, 3011, 3012 and 3013, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. The processing order of the four CUs obtained by partitioning coding block 301 is 3010, 3011, 3012, and 3013.
[0088] FIG. 3B illustrates an exemplary technique of binary splitting and ternary splitting 325 of a coding unit, according to some embodiments of the present disclosure. As shown in FIG. 3B, various examples of binary splitting and / or ternary splitting is applied to a CU are depicted. For instance, CU 302 is partitioned using vertical binary split. Blocks 3020 and 3021, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 303 is partitioned using horizontal binary split. Blocks 3030 and 3031, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 304 is partitioned using vertical ternary split. Blocks 3040, 3041, and 3042, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 305 is partitioned using horizontal ternary split. Blocks 3003, 3051, and 3052, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc.
[0089] FIG. 3C illustrates an exemplary technique of splitting 350 of a coding unit into various split types, according to some embodiments of the present disclosure. As shown in FIG. 3C, an example of partitioning or splitting a CU 306 iteratively using various split or partition types, e.g., such as quadtree split, binary split, ternary split, etc. is depicted.
[0090] Referring again to FIG. 2, prediction unit 202 may derive inter prediction block of a CU using inter prediction unit 204, and may derive intra prediction block of a CU using intra prediction unit 205. In an example, prediction unit 202 may use a rate-distortion mode decision process to determine a prediction mode of a CU.
[0091] Generally, inter prediction unit 204 performs motion estimation to derive motion parameters of a CU. The motion parameters include motion vector (MV) and reference index (refIdx) . MV indicates a relative location of a matching block in a reference picture indicated by refIdx in a specific reference list (e.g., list 0 and list 1) . Generally, list 0 mainly includes reference pictures that is ahead of the current picture in an output order or a displaying order, while list 1 mainly includes reference pictures that is behind the current picture in an output order or a displaying order. Inter prediction unit 204 may derive the MV of the CU using the samples in the CU and find the block in the reference with least cost according to a rate-distortion motion estimation method. Inter prediction unit 204 may derive motion parameters and / or prediction samples using an NN-based method or process. Inter prediction unit 204 may derive the MV of the CU using spatial reference samples of the CU.
[0092] FIG. 4 illustrates an exemplary technique of inter prediction 400 based on template matching, according to some embodiments of the present disclosure.
[0093] As shown in FIG. 4, CU 403 is the current CU in the current picture 401. A reference picture of the current CU 403 is a reconstructed or decoded reference picture 402. Template 404 includes the above neighboring samples and / or left samples. Instead of deriving a matching block of the current CU 403 in a search range in a reference picture 402, inter prediction unit 204 may derive a matching template 405 of template 404 in reference picture 402. An initial MV is derived based on the displacement between template 404 and matching template 405. In an example, the block 406 in a same size of the current block 403 may be used as an inter prediction block of the current CU 403. FIGs. 5A-5C shows examples of templates.
[0094] FIG. 5A illustrates first exemplary templates 500 used in template matching, according to some embodiments of the present disclosure. Referring to FIG. 5A, examples of candidate templates using neighboring samples of the current CU (Curr CU) are depicted. As an example, a template of the current CU may be one of AboveLeft, Above, AboveRight, Left and BottomLeft. As another example, a template of the current CU may be a region of a combination of two or more of AboveLeft, Above, AboveRight, Left and BottomLeft. Take template 404 for example. Template 404 in FIG. 4 is a combination of Above and Left templates in FIG. 5A.
[0095] FIG. 5B illustrates second exemplary templates 525 used in template matching, according to some embodiments of the present disclosure. As shown in FIG. 5B, examples of candidate templates using samples that are not directly neighboring the current CU (Curr CU) are depicted. As an example, a template of the current CU may be one of AboveLeft, Above-Left, Above, AboveRight, Left-Above, Left and BottomLeft. As another example, a template of the current CU may be a region of a combination of two or more of AboveLeft, Above-Left, Above, AboveRight, Left-Above, Left and BottomLeft. As another example, a template of the current CU may be a combination of two or more of the example candidate templates in FIGs. 5A and 5B.
[0096] FIG. 5C illustrates third exemplary templates 550 used in template matching, according to some embodiments of the present disclosure. As shown in FIG. 5C examples of templates from the candidate templates in FIG. 5A are shown. In some other implementations, templates using non-neighboring samples may be derived in a similar way in FIG. 5C using examples of candidate templates in FIG. 5B.
[0097] Referring again to FIG. 2, inter prediction unit 204 may refine an existing MV based on a displacement between a template of a current block and a matching block. Additionally and / or alternatively, inter prediction unit 204 may refine an existing MV based on a displacement between two template blocks. In some embodiments, inter prediction unit 204 may also use the matching template to derive certain compensation to the current block.
[0098] Generally, inter prediction unit 204 may determine a template shape. Then, inter prediction unit 204 may calculate a cost between various reference templates and the template. Finally, inter prediction unit 204 may determine a reference template that results in an optimal cost function as a matching template of the template.
[0099] A cost between two templates may be represented as an error between a template and a reference template. As an example, the cost may be a Sum of Absolute Differences (SAD) calculated according to equation (1) shown below. where Ti, m and Tm are samples in two templates, respectively; and M is a number of samples in a template.
[0100] As an example, the cost may be a Sum of Absolute Transformed Differences (SATD) calculated according to equation (2) shown below. where X represents a matrix of a difference between two template samples, M is the size of the matrix, and H is a normalized MxM Hadamard matrix.
[0101] As an example, the cost may be a Mean Reduced Sum of Absolute Differences (MR-SAD) calculated according to equation (3) shown below. where Ti, m and Tm are samples in two templates, respectively; M is the number of samples in a template; Avgi is the average value of a first template containing Ti, m; and Avg is the average of a second template containing Tm.
[0102] In some implementations, other cost measures that may be used by inter prediction unit 204 may include one or more of, e.g., Mean Squared Error (MSE) , Sum of Squared Differences (SSD) , Mean Absolute Difference (MAD) , Mean Squared Difference (MSD) , Normalized Cross-Correlation (NCC) , Structural Similarity Index Measure (SSIM) , or Multi-Scale Structural Similarity Index Measure (MS-SSIM) , just to name a few.
[0103] Still referring to FIG. 2, besides directly using the matching block 406 as a prediction of the current CU 403, inter prediction unit 204 may be enabled with a number of prediction modes that utilizes template matching to derive a matching template. Table 1 below shows example prediction modes, which may be carried out by the inter prediction unit 204, with default cost functions of the modes. As one example, the intra block copying (IBC) mode may be carried out by inter prediction unit 204 by setting available reconstructed area of the current picture as “reference picture” to derive a matching block of the current block. Table 1: Example Prediction Modes for Inter Prediction
[0104] Inter prediction unit 204 may determine a cost different from the default cost for one or more prediction modes in Table 1. That is, inter prediction unit 204 may choose a cost function from several candidate cost functions. In some implementations, inter prediction unit 204 may determine whether default costs are used for one or more of the inter prediction modes. In some implementations, inter prediction unit 204 can determine that for an inter prediction mode that utilizes template matching, a first cost function (e.g., SAD) may be used for coding all pictures in a video sequence, and / or pictures within a certain period in a video sequence, and / or a picture, and / or a tile, and / or a slice, and / or a CTU, and / or a CU. In an embodiment, inter prediction unit 204 may use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 1, inter prediction unit 204 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, inter prediction unit 204 can use SAD or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In an embodiment, inter prediction unit 204 can use a single cost function for all prediction modes.
[0105] Generally, intra prediction unit 205 may derive an intra prediction block of a CU using various intra prediction modes including, e.g., direct copying (DC) mode, planar mode, angular prediction mode, Matrix-based Intra Prediction (MIP) mode, cross-component linear model intra prediction (CCLM) mode, intra block copying (IBC) mode, intra template matching prediction (IntraTMP) mode, NN-based intra prediction mode, etc. In an example, rate-distortion optimized motion estimation may be invoked by intra prediction unit 205 to derive the intra prediction mode for a current block; and according to the intra prediction mode, intra prediction unit 205 determines the intra prediction block for the current block.
[0106] FIG. 6 illustrates an exemplary technique of intra prediction 600 based on template matching, according to some embodiments of the present disclosure.
[0107] Referring to FIG. 6, an example of IntraTMP mode is shown. Picture 601 is a current picture, and block 603 is current block or current CU. Region 602 in picture 601 is an already reconstructed area. Current block 603 is within region 608, which is an area to be reconstructed or encoded. Intra prediction unit 205 first determines a template of the current block 603, where a template may be one of the template types as shown in one or more of FIGs. 5A-5C. In the non-limiting example shown in FIG. 6, intra prediction unit 205 uses a template that is a combination of Above and Left neighboring samples of the current block 603. Intra prediction unit 205 may then search in a search range within region 602 to find a matching template that results in an optimal cost as a matching template of the template.
[0108] The cost between two templates may be represented as an error between a template and a reference template. As an example, the cost may be the SAD calculated according to equation (1) , shown above. As another example, the cost may be an SATD calculated according to equation (2) , shown above. As a further example, the cost may be a MR-SAD calculated according to equation (3) shown above. In some implementations, other cost measures that may be used by intra prediction unit 205 may include one or more of, e.g., MSE, SSD, MAD, MSD, NCC, SSIM, or MS-SSIM, just to name a few.
[0109] By way of example and not limitation, intra prediction unit 205 may use the SAD function to determine a matching reference template 605 for the template 604. The reference sample 607, which may be in a same size as that of the current block 603, is used by the intra prediction unit 205 to derive a prediction of the current block 603. A block vector (BV) 606 represents a displacement between the template 604 and its matching reference template 605, and / or represents a displacement between the current block 603 and the reference sample 607.
[0110] In some non-limiting examples, referring to FIGs. 2 and 6, intra prediction unit 205 can use the reference sample 607 as the prediction of the current block 603. In some non-limiting examples, intra prediction unit 205 can use a result of filtering the reference sample 607 to be the prediction of the current block 603. In some non-limiting examples, intra prediction unit 205 can use a result of the reference sample 607 adjusted with weighting factor to be the prediction of the current block 603.
[0111] Still referring to FIGs. 2 and 6, intra prediction unit 205 can derive more than one reference sample, in some implementations. For example, an optimal matching block and a sub-optimal matching block may be derived by the intra prediction unit 205, and two reference samples may be available. Intra prediction unit 205 may perform a fusion operation on the multiple reference samples derived in template matching process to derive a prediction of the current block 603. An example fusion operation is a weighted combination of the multiple reference samples, where the weights may be determined based on a matching error between the two reference templates and the template 604.
[0112] Still referring to FIGs. 2 and 6, in one example, besides deriving the prediction of the current block, intra prediction unit 205 may also derive coding parameters for the current block based on the reference template 605. In one embodiment, intra prediction unit 205 may derive model parameters for performing cross-component prediction on the current block 603 based on one or more reference templates 605.
[0113] Table 3 shows example prediction modes, which may be carried out by the intra prediction unit 205, with default cost functions of the modes. In some implementations, the intra block copying (IBC) mode may be carried out by intra prediction unit 205 as the prediction block is derived only based on the samples in the current picture and no temporal reference picture or inter-layer reference picture is used. Table 3: Example Prediction Modes for Intra Prediction
[0114] Referring to FIG. 2, intra prediction unit 205 can determine a cost different from the default cost for one or more prediction modes in Table 3. That is, intra prediction unit 205 can choose a cost function from several candidate cost functions. In some implementations, intra prediction unit 205 can determine whether default costs are used for one or more of the intra prediction modes. In some implementations, intra prediction unit 205 can determine that for an intra prediction mode that utilizes template matching, a first cost function (e.g., SAD) may be used for coding all pictures in a video sequence, and / or pictures within a certain period in a video sequence, and / or a picture, and / or a tile, and / or a slice, and / or a CTU, and / or a CU. In some implementations, intra prediction unit 205 may use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 3, intra prediction unit 205 may use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 3, intra prediction unit 205 can use SAD, or any one of the abovementioned cost functions for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In some implementations, intra prediction unit 205 can use a single cost function for all prediction modes.
[0115] Prediction unit 202 may also derive intra prediction mode or angular prediction direction of the current CU based on one or more reference samples, as described below in connection with FIG. 16.
[0116] FIG. 16 illustrates an example visualization 1600 of a current block 1601 and reference samples 1602, 1603, according to some embodiments of the present disclosure.
[0117] Referring to FIG. 16, the reference sample 1602, 1603, which are marked as black dot outside the current block 1601, may be one or more samples in a template ( “L-shape” template consisting of the black dots in FIG. 16) as shown in FIGs. 5A-5C. For example, a gradient of reference sample 1602, 1603 may be derived by applying one or more filters to process one or more samples in a template as shown in FIGs. 5A-5C. Prediction unit 202 first may derive a gradient of a reference sample using an operator. Generally, the operator may be used to detect an edge or a gradient in a picture. The operator may be a 2-dimentional (2D) M x N filter, where M and N are positive integers, and M may be equal to or different from N.
[0118] One example of the operator is a Sobel filter. A first example of a 3x3 Sobel filter 4200 is depicted in FIG. 42A, and a second example of a 3x3 Sobel filter 4225 is depicted in FIG. 42B.
[0119] Another example of the operator is an Edge filter. A first example of an Edge filter 4250 is depicted in FIG. 42C, and a second example of an Edge filter 4275 is depicted in FIG. 42D.
[0120] In some implementations, prediction unit 202 can choose different operators according to a width and / or a height of the current block. For example, prediction unit 202 uses smaller operator for smaller block, and uses larger operator for larger block. One example would be that prediction unit 202 uses the abovementioned Edge filter when a size (e.g., width x height) of the current block is 4x4, 4x8 or 8x4, and uses the abovementioned Sobel filter for other sizes of the current block.
[0121] Prediction unit 202 may derive a Histogram of Gradients (HoG) by analyzing one or more reference samples 1602, 1603 marked in black dots in FIG. 16. The template in FIG. 16 may include three reference sample lines above and three reference sample columns on the left of the current block 1601. The HoG is derived by accumulating the magnitudes of one or more gradients at one or more given directions, for one, a part of, or all of the reference samples 1602, 1603 as shown in FIG. 16. One or more directions indicated by the gradients with the highest or higher cumulative magnitudes may be identified as the angular prediction direction or intra prediction mode of the current block 1601.
[0122] As an example, when prediction unit 202 uses one reference sample (e.g., any black dot above and / or to the left of the current block 1601) as shown in FIG. 16 to derive a direction, the prediction unit 202 can determine the direction as the one indicated by a gradient derived at this reference sample.
[0123] As an example, when prediction unit 202 uses a part of reference samples (e.g., one or more of the black dots above and / or to the left of the current block 1601) as shown in FIG. 16 to derive directions, prediction unit 202 can choose a preset number of samples from the reference samples. For example, prediction unit 202 choose J reference samples above the current block and K reference samples left to the current block, wherein J and K are integers greater than or equal to 0. For example, both J and K are equal to 2, J equal to 4 and K equal to 8, or J plus K equal to 4. Prediction unit 202 may derive the gradients at the selected reference samples using the operator and may derive the HoG. One or more directions indicated by the gradients with highest or higher cumulative magnitudes may be identified as the angular prediction direction or intra prediction mode of the current block 1601.
[0124] As an example, prediction unit 202 can adaptively determine one or more reference samples used for deriving a HoG. Prediction unit 202 uses a preset scanning order of the reference samples. When scanning a reference sample, prediction unit 202 may derive a gradient at this reference sample, and updates the accumulation of the magnitudes according to this gradient at one or more given directions in the HoG. When prediction unit 202 determines that the total cumulative amplitude is greater or equal than a given threshold, prediction unit 202 will stop scanning the remaining reference sample and deriving new gradient. The resulting HoG at the termination of scanning by prediction unit 202 is determined as the HoG, which may be used to derive intra prediction mode or angular prediction direction.
[0125] Examples of the abovementioned preset scanning order of the reference samples may be the following.
[0126] For example, a scanning order (A) may be scanning the left column of reference samples as shown in FIG. 16 from bottom to top; a scanning order (B) may be scanning the left column of reference samples as shown in FIG. 16 from top to bottom; a scanning order (C) may be scanning the above line of reference samples as shown in FIG. 16 from left to right; a scanning order (D) may be scanning the above line of reference samples as shown in FIG. 16 from right to left.
[0127] An example of the preset scanning order may be one or more of “first order (A) then order (C) , ” “first order (A) then order (D) , ” “first order (B) then order (C) , ” “first order (B) then order (D) , ” “first order (C) then order (A) , ” “first order (D) then order (A) , ” “first order (C) then order (B) , ” and / or “first order (D) then order (B) . ”
[0128] In some implementations, the preset scanning order may be performed in an interleaving manner. In one example, one or multiple reference samples may be scanned from the left column and then one or multiple second reference samples from the above line and then one or multiple third reference sample from left column. In another example, one or multiple reference samples may be scanned from the above line and then one or multiple second reference samples from the left column and then one or multiple third reference sample from above line. Additionally, as an example, the scanning order of samples in the left column may be one or more of order (A) and (B) , and the scanning order of samples in the above line may be one or more of order (C) and (D) .
[0129] Prediction unit 202 may use the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction also may be called as “intra prediction direction” ) to derive a prediction of the current block. For example, prediction unit 202 may pass the derived one or more intra prediction modes or one or more angular prediction directions to intra prediction unit 205. In some implementations, intra prediction unit 205 may derive a prediction of the current block by fusing one or more predictions corresponding to the derived intra prediction modes. In one embodiment, Intra prediction unit 205 may derive a prediction of the current block by fusing one or more predictions corresponding to the derived intra prediction modes. In one embodiment, Intra prediction unit 205 may derive a prediction of the current block by fusing one or more predictions determined according to the derived intra prediction modes and one or more predictions determined according to one or more preset mode (e.g., Planar mode, DC mode and cross-component prediction mode and etc. ) .
[0130] Prediction unit 202 may use Occurrence-based intra coding (OBIC) to derive the intra prediction modes of the current block based on the sample-wise occurrence of the intra modes in the spatial neighborhood of the block. For this, adjacent and non-adjacent spatial neighboring blocks are checked, and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradient (HoG) as in DIMD, the OBIC method uses the Histogram of occurrence (HoC) , which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an IPM mode, the occurrence of the mode in that particular block is calculated as: HoC[IPM] += uiWidth * uiHeight, where uiWidth and uiHeight are the width and height of a spatial neighboring block.
[0131] The occurrences of the existing modes from the spatial neighborhood blocks are accumulated into the histogram, adjacent and non-adjacent spatial neighboring blocks are checked, and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradient (HoG) as in DIMD, the OBIC method uses the Histogram of occurrence (HoC) , which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an Intra Predication Mode (IPM) mode, the occurrence of the mode in that particular block is calculated as: HoC [IPM] += uiWidth * uiHeight, where uiWidth and uiHeight are the width and height of a spatial neighboring block.
[0132] The occurrences of the existing modes from the spatial neighborhood blocks are accumulated into the histogram.
[0133] FIG. 17 illustrates a diagram of non-adjacent spatial neighboring candidates 1700 for occurrence-based intra coding (OBIC) mode, according to some embodiments of the present disclosure.
[0134] Referring to FIG. 17, one or multiple (e.g., up to five angular modes) with the highest occurrence along with the planar mode or block vector based prediction (same as in DIMD) are selected from the HoC and used for final prediction by blending the prediction of the selected modes.
[0135] Some blocks, mentioned below, use more than one intra mode for prediction. In such cases, all the intra modes of such blocks are selected and used when creating the OBIC histogram. For example, DIMD may use up to 5 angular modes, TIMD may use up to 2 modes, SGPM may use up to 2 modes, and OBIC may use up to 5 angular modes.
[0136] Moreover, the virtual intra prediction modes (VIPMs) of the following blocks are considered only in inter slices when creating the histogram of OBIC mode: MIP block, IntraTMP block, IBC block, and EIP block.
[0137] The blending weights are calculated similarly to the DIMD mode, but instead of using gradient values from the template, the occurrence values are used for OBIC. Moreover, the planar mode’s weight is also decided similarly to the DIMD mode.
[0138] FIG. 18 illustrates an example of the histogram-of-occurrences (HoC) 1800 of intra prediction mode, according to some embodiments of the present disclosure.
[0139] Referring to FIG. 18, in an embodiment, in order to decrease the buffer memory, the OBIC method may use non-sample-wise occurrences, such as block-wise occurrences. For example, if a uiWidth× uiHeight block is coded with an IPM mode, the block-wise occurrence of the mode in that particular block is calculated as: HoC [IPM] += (uiWidth>>shift1) *(uiHeight>>shift2) , where the variable shift1 and shift2 are both positive integers greater than or equal to 1. When the variable shift1 and shift2 are set to 1, the occurrence type 2×2 block-wise occurrence is used. The variable shift 1 or shift 2 can also be set to 2, 3, 4, 5, and so on. It is not limited that the shift1 is equal to shift2.
[0140] In an embodiment, the shift1 or shift 2 may be determined according to the size (width or height) of the current block. For example, if the size of the current block is 4×4, the 2×2 block-wise occurrence may be used.
[0141] In an embodiment, the shift1 or shift2 may be determined by using the flag sps_log2_min_luma_coding_block_size_minus2. For example, parse a bitstream and obtain the value of sps_log2_min_luma_coding_block_size_minus2. If the value of sps_log2_min_luma_coding_block_size_minus2 is equal to 2, the size of the current block is determined to 4×4. Then the occurrence type may be obtained through the determined size of the current block.
[0142] In an embodiment, the block-wise occurrence may be determined through a look-up table. Tables 4A-4E illustrate various non-limiting examples of look-up tables that correlate CU size to occurrence type. Table 4A: First Example Occurrence-Type Look-up Table Table 4B: Second Example Occurrence-Type Look-up Table Table 4C: Third Example Occurrence-Type Look-up Table Table 4D: Fourth Example Occurrence-Type Look-up Table Table 4E: Fifth Example Occurrence-Type Look-up Table
[0143] In an embodiment, a confidence level can also be used to calculate the occurrence. If there is a very large size block neighboring a small size block, the IPM of the large size block is considered as a low confidence level block and the large size block is not used to calculate the occurrence of the current block. For example, if a 64×64 block neighbors a 2×2 block, the 64×64 block may not be used to calculate the occurrence of the 2×2 block for the reason that the 2×2 block is too small than the 64×64 block. If the 64×64 block is used to calculate, it will negatively interfere with the histogram statistics. The size of current block is a factor the determined the confidence level of a neighbor block.
[0144] In an embodiment, the OBIC mode may be used to luma blocks.
[0145] In an embodiment, the OBIC mode may be used to chroma blocks.
[0146] FIG. 22 illustrates an example of adjacent blocks 2200 of current block 2201, according to some embodiments of the present disclosure. The numbers on adjacent reference samples shows an example of scanning order of the adjacent blocks by prediction unit 202. In one example, intra prediction unit 205 scans prediction modes of the adjacent reference samples to derive one or more MPM. If the mode of an adjacent block is not a BV based intra prediction mode (e.g., planar mode, DC mode, angular prediction mode, etc. ) , intra prediction unit 205 may include this intra mode as a candidate mode in MPM list. In an example, spatial reference samples can also be one or more non-adjacent blocks shown in FIG. 17.
[0147] FIG. 23 illustrates a first example of an adjacent block of a BV-based prediction mode, according to some embodiments of the present disclosure. In FIG. 23, 2300 is a current picture 2300 and 2301 is a current block 2301.
[0148] For example, current block 2301 is a block in an intra coded slice. For example, current block 2301 is a block in an intra coded picture (e.g., a picture that may be employed as an access point, such as Instantaneous Decoding Refresh picture, Clean Random Access picture, Broken Link Access picture, etc. ) .
[0149] Block 2302 is an adjacent block of current block 2301, and block 2302 is coded using a BV based coding mode, e.g., IBC or IntraTMP. BV (2302) is a BV of block 2302 and indicates a reference block 2303. If block 2303 is an intra coded block which only references to reconstructed samples in the current picture 2300 and block 2303 is not using a BV based coding mode, for example, block 2303 is of an angular mode, DC mode or planar mode, the width and / or height of block 2303 may be used to derive an HoC of the current block 2301. As block 2302 is of a same size as that of 2303, equivalently, the width and / or height of block 2302 may be used to derive an HoC of the current block 2301. In one example, coding information of the block 2303 may be used to derive an HoC of the current block 2301. In one example, the coding information of the block 2303 may include one or more of the width of block 2303, the height of block 2303, or the coding mode of block 2303. In one example, the width and / or height of block 2303 as well as the coding mode of block 2303 may be used to derive an HoC of the current block 2301.
[0150] If block 2303 is also coded using a BV based coding mode, prediction unit 202 will check a coding mode of block 2304, which is indicated by BV (2303) of block 2303. If block 2304 is an intra coded block which only references to reconstructed samples in the current picture 2300 and block 2304 is not using a BV based coding mode, for example, block 2304 is of an angular mode, DC mode or planar mode, the width and / or height of block 2304 may be used to derive an HoC of the current block 2301. As block 2302 is of a same size as that of 2303 and 2304, equivalently, the width and / or height of block 2302 may be used to derive an HoC of the current block 2301. In one example, coding information of the block 2304 may be used to derive an HoC of the current block 2301. In one example, the coding information of the block 2304 may include one or more of the width of block 2304, the height of block 2304, or the coding mode of block 2304. In one example, the width and / or height of block 2304 as well as the coding mode of block 2304 may be used to derive an HoC of the current block 2301.
[0151] Optionally, if prediction unit 202 cannot determine an intra coding mode that is not BV based mode after recursively searching using “BV links, ” for example using N linked BVs, prediction unit 202 will stopped and does not use any information from block 2302 to derive HoC of the current block 2301. As an example, N is a non-negative integer. When N is equal to 0, only block 2302 is checked by prediction unit 202. As an example, when N is equal to 1, block 2302 and block 2303 may be checked if block 2302 is coded using BV-based intra mode. As an example, when N is equal to 2, blocks 2302, 2303 and 2304 may be checked by prediction unit 202 if blocks 2302 and 2303 is coded using BV-based intra mode.
[0152] FIG. 24 illustrates a second example of an adjacent block of BV-based prediction mode, according to some embodiments of the present disclosure. For example, current block 2401 may be a block in an intra coded slice. For example, current block 2401 is a block in an intra coded picture (e.g., a picture that may be employed as an access point, such as Instantaneous Decoding Refresh picture, Clean Random Access picture, Broken Link Access picture, etc. ) .
[0153] Block 2402 is an adjacent block of current block 2401, and block 2402 is coded using a BV based coding mode, e.g. IBC or IntraTMP. BV (2402) is a BV of block 2402. Block 2403 is a block which contains a reference sample indicated by BV (2402) . In an example, (x0, y0) denotes a sample in current block 2401, and the reference sample is derived as (x0 + x (2402) , y0 + y (2402) ) , wherein BV (2402) has a horizontal component equal to x (2402) and a vertical component equal to y (2402) . For example, (x0, y0) may be a top-left sample in the current block 2401. For example, (x0, y0) may be a sample located at any corner of the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-right position (e.g., (width (2401) / 2, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-right position (e.g., (width (2401) / 2, height (2401) / 2 - 1) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-left position (e.g., (width (2401) / 2 - 1, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2 - 1) ) in the current block 2401. width (2401) and height (2401) are width and height, respectively, of the current block 2401.
[0154] If block 2403 is an intra coded block which only references to reconstructed samples in the current picture 2400 and block 2403 is not using a BV based coding mode, for example, block 2403 is of an angular mode, DC mode or planar mode, in one example, the width and / or height of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 may be used to derive a HoC of the current block 2401.
[0155] In one example, coding information of the block 2403 may be used to derive a HoC of the current block 2401. In one example, the coding information of the block 2403 may include one or more of the width of block 2403, the height of block 2403, or the coding mode of block 2403. In one example, coding information of the block 2403 and one of coding information of the block 2402, coding information of the current block 2401 may be used to derive a HoC of the current block 2401. In one example, the width and / or height of block 2403 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401.
[0156] If block 2403 is also coded using a BV based coding mode, prediction unit 202 will check a coding mode of block 2404, which is indicated by BV (2403) of block 2403. In one example, Block 2404 is a block which contains another reference sample indicated further by BV (2403) . In one example, (x0, y0) denotes a sample in current block 2401, and the another reference sample is derived as (x0 + x (2402) + x (2403) , y0 + y (2402) + x (2403) ) , wherein BV (2402) has a horizontal component equal to x (2402) and a vertical component equal to y (2402) , and BV (2403) has a horizontal component equal to x (2403) and a vertical component equal to y (2403) . For example, (x0, y0) may be a top-left sample in the current block 2401. For example, (x0, y0) may be a sample located at any corner of the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-right position (i.e. (width (2401) / 2, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-right position (i.e. (width (2401) / 2, height (2401) / 2 - 1) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2 - 1) ) in the current block 2401. width (2401) and height (2401) are width and height, respectively, of the current block 2401.
[0157] If block 2404 is an intra coded block which only references to reconstructed samples in the current picture 2400 and block 2404 is not using a BV based coding mode, for example, block 2404 is of an angular mode, DC mode or planar mode, in one example, the width and / or height of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 may be used to derive a HoC of the current block 2401.
[0158] In one example, coding information of the block 2404 may be used to derive a HoC of the current block 2401. In one example, the coding information of the block 2404 may include one or more of the width of block 2404, the height of block 2404, or the coding mode of block 2404. In one example, coding information of the block 2404 and one of coding information of the block 2403, coding information of the block 2402, or coding information of the current block 2401 may be used to derive a HoC of the current block 2401. In one example, the width and / or height of block 2404 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2403 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401.
[0159] Optionally, if prediction unit 202 cannot determine an intra coding mode that is not BV based mode after recursively searching using “BV links, ” for example using N linked BVs, prediction unit 202 will stopped and does not use any information from block 2402 to derive HoC of the current block 2401. As an example, N is a non-negative integer. When N is equal to 0, only block 2402 is checked by prediction unit 202. As an example, when N is equal to 1, block 2402 and block 2403 may be checked if block 2402 is coded using BV-based intra mode. As an example, when N is equal to 2, blocks 2402, 2403 and 2404 may be checked by prediction unit 202 if blocks 2402 and 2403 is coded using BV-based intra mode.
[0160] FIG. 25 illustrates a flowchart of an example method 2500 of predicting a current block using a BV-based prediction technique, according to some implementations. The BV-based prediction technique may be used in intra prediction unit 205 and inter prediction unit 204. For example, BV-based intra prediction technology may include intra template matching prediction (IntraTMP) , intra block copy (IBC) , spatial geometric partitioning mode (SGPM) , and so on. Additionally, the BV-based prediction technique may be used in inter prediction technology, e.g., such as the Geometric partitioning mode (GPM) . It should be noted that the following BV-based prediction may be applied to either intra or inter prediction.
[0161] Referring to FIG. 25, at block 2502, a block vector of a current block may be determined as follows.
[0162] Methods for determining the BV of the current block include, but are not limited to, 1) performing motion search in the reconstructed area of the current picture to obtain the best BV for the current block and 2) constructing a BV Merge list for the current block by using the BV of the reconstructed block in the spatial and temporal domains. In general, BV has two precision options: integer-pel precision and sub-pel precision. The sub-pel precision may include 1 / 2-pel precision, 1 / 4-pel precision, or 1 / 16-pel precision, etc.
[0163] In some implementations, the BV of the current block may be determined based on a motion search. To perform a motion search, one type of template shown in FIGs. 5A-5C may be determined for the current block. Then, a search may be performed to identify a matching template with the smallest matching cost with the template of the current block in the reconstructed area of the current picture. An area with the same size as the current block corresponding to the matching template may be determined as the reference block. The BV of the current block may be determined based on the current block and the reference block.
[0164] In some implementations, the BV of the current block may be determined based on a BV merge list. The BV merge list may include at least one of spatial merge candidate, temporal merge candidate, history-based merge candidate, pairwise average merge candidate, and default merge candidate.
[0165] In some implementations, the BV merge list may be constructed based on spatial merge candidates. As shown in FIG. 26, the spatial merge candidates 2600 include merge candidates A1, B1, B0, A0, and B2. Spatial merge candidates A1, B1, B0, A0, and B2 are sequentially checked. If one of the spatial merge candidates A1, B1, B0, A0 and B2 is available and the available spatial merge candidate applies a BV-based prediction mode, then the BV of that spatial merge candidate may be added to the BV merge list. If the number of BV merge candidates added in the BV merge list is smaller than the allowable maximum value of the BV merge list after checking the spatial merge candidates, then one or more temporal merge candidates may be checked. One or more collocated positions in the collocated picture may be determined and temporal merge candidates corresponding to the collocated positions may be checked. If a temporal merge candidate is available and applies BV-based prediction mode, then the BV of the temporal merge candidate may be added to the BV merge list. If the number of BV merge candidates added in the BV merge list is smaller than the allowable maximum value of the BV merge list after checking the temporal merge candidates, the history-based merge candidates, pairwise average merge candidate, or default merge candidate may be checked. The BV for the current block may be determined from the BV merge list based on corresponding cost values. In some implementations, the BV merge list may be reordered, and the BV for the current block may be determined based on cost values after the reordering.
[0166] It should be noted that, at operation 2502, other methods may be used to determine the BV for the current block. The following embodiments may be applied as long as a BV is used in the process of prediction, regardless of how the BV is obtained. In some implementations, the BV may be determined having a integer-pel precision.
[0167] In some implementations, BV may be determined having a sub-pel precision, including 1 / 2-pel precision, 1 / 4-pel precision, or 1 / 16-pel precision, etc. A BV with a sub-pel precision means the BV may refer to a fractional-pel position of the reconstruction area of the current picture. A BV with an integer-pel precision means the BV may refer to an integer-pel position of the reconstruction area of the current picture.
[0168] If an integer-pel precision BV is determined, to improve the precision of the BV, or considering the unity of BV precision, the integer-pel precision BV may be converted to sub-pel precision. In one example, the integer-pel precision BV may be converted to the sub-pel precision BV according to the following functions: xfrac=xint<< (FRAC_BITS-INT_BITS) , and yfrac=yint<< (FRAC_BITS-INT_BITS) .
[0169] The sub-pel precision BV is represented by (xfrac, yfrac) , and integer-pel precision BV is represented by (xint, yint) , where xfrac or xint indicates a horizontal value of the BV and yfrac or yint indicates a vertical value of the BV.
[0170] The sub-pel precision may be represented by FRAC_BITS and the integer-pel precision may be represented by INT_BITS. INT_BITS indicates a number of bits representing a BV with integer-pel precision, and FRAC_BITS indicates a number of bits representing a BV with sub-pel precision. In an example, the value of INT_BITS is equal to 2. The value of FRAC_BITS of 1 / 2-pel precision is equal to 3. The value of FRAC_BITS of 1 / 4-pel precision is equal to 4. The value of FRAC_BITS of 1 / 16-pel precision is equal to 6. Of course, the values of INT_BITS and values of FRAC_BITS of different sub-pel precisions may be determined as different preset values; or the values of INT_BITS and values of FRAC_BITS of different sub-pel precisions may be determined adaptively based on the network conditions.
[0171] At operation 2504, the block vector of the current block may be adjusted.
[0172] Different prediction modes may utilize BVs with different pel precisions. Therefore, for the BV of the current block, an adjustment of the BV may be determined based on the prediction mode. In some implementations, if the determined BV of the current block has a sub-pel precision and prediction mode utilizes a BV with integer-pel precision, a rounding operation may be performed on the determined BV of the current block to obtain an adjusted BV with integer-pel precision. The rounding operation may include, but is not limited to, rounding, rounding up and rounding down.
[0173] In one example, a rounding process may be performed on the determined BV (xfrac, yfrac) of the current block to determine an adjusted BV. An adjusted sub-pel precision BV (xint, yint) of the current block may be determined according to the following function (s) : If xfrac≥0, xint=(xfrac+nOffset-1)>>rightShift; If xfrac<0, xint=(xfrac+nOffset)>>rightShift; If yfrac≥0, yint=(yfrac+nOffset-1)>>rightShift; and If yfrac<0, yint=(yfrac+nOffset)>>rightShift, where rightShift and nOffset may be determined as follows: rightShift=FRAC_BITS-INT_BITS; and nOffset=1<< (rightShift-1) .
[0174] If BV (xfrac, yfrac) has a 1 / 2-pel precision, FRAC_BITS is equal to 3, and INT_BITS is equal to 2; thus, rightShift=1 and nOffset=1. If BV (xfrac, yfrac) has a 1 / 4-pel precision, FRAC_BITS is equal to 4, and INT_BITS is equal to 2; thus, rightShift=2 and nOffset=2. If BV (xfrac, yfrac) has a 1 / 16-pel precision, FRAC_BITS is equal to 6, and INT_BITS is equal to 2; thus, rightShift=4 and nOffset=8.
[0175] In one example, a rounding down process may be performed on the determined BV (xfrac, yfrac) to determine an adjusted BV (xint, yint) of the current block using the following functions: xint=xfrac>> (FRAC_BITS-INT_BITS) , and yint=yfrac>> (FRAC_BITS-INT_BITS) .
[0176] In one example, a rounding up process may be performed on the determined BV (xfrac, yfrac) to determine an adjusted BV (xint, yint) of the current block using the following functions: xtemp= (xfrac>> (FRAC_BITS-INT_BITS) ) << (FRAC_BITS-INT_BITS) ; ytemp= (yfrac>> (FRAC_BITS-INT_BITS) ) << (FRAC_BITS-INT_BITS) ; If xfrac=xtemp, xint=xfrac>> (FRAC_BITS-INT_BITS) ; If xfrac≠xtemp, xint=xfrac>> (FRAC_BITS-INT_BITS) +1; If yfrac=ytemp, yint=yfrac>> (FRAC_BITS-INT_BITS) ; and If yfrac≠ytemp, yint=yfgrac>> (FRAC_BITS-INT_BITS) +1.
[0177] If a sub-pel precision BV is obtained for the current block, the sub-pel precision BV may be converted to an integer-pel precision BV for use in predicting the current block. In this way, a reference block may be determined in the reference area based on the integer-pel precision BV with no sample interpolation, thereby improving coding efficiency.
[0178] In some implementations, the BV of the current block may be adjusted based on refinement. In one example, if the determined BV of the current block has a sub-pel precision, the determined BV may be adjusted based on a rounding operation to obtain an integer-pel precision BV, and the integer-pel precision BV may be further refined. The rounding operation may include, but is not limited to, rounding, rounding up, or rounding down. Using a refined BV for the prediction process may improve prediction accuracy.
[0179] FIG. 27 illustrates example sub-pel and integer-pel positions 2700, according to some implementations. Sub-pel positions include 1 / 4-pel positions, 1 / 2-pel positions and 3 / 4-pel positions. The sub-pel position may be represented by FracPre. The sub-pel direction may include eight directions: left (LEFT_POS) , above left (ABOVE_LEFT_POS) , left bottom (LEFT_BOTTOM_POS) , right (RIGHT_POS) , above right (ABOVE_RIGHT_POS) , right bottom (RIGHT_BOTTOM_POS) , above (ABOVE_POS) and bottom (BOTTOM_POS) . The sub-pel direction may be represented by FracDir. The sub-pel position may also be referred to as “fractional-pel position” or “fractional sample position. ” The integer-pel position may also be referred to as “integer sample position. ”
[0180] An integer-pel precision BV may be adjusted to a sub-pel precision BV. For example, a sub-pel precision BV, BV′int (x′int, y′int) , may be determined based on an integer-pel precision BV, BVint (xint, yint) , according to the following functions: x′int=xint<< (FRAC_BITS-INT_BITS) , and y′int=yint<< (FRAC_BITS-INT_BITS) .
[0181] A sub-pel precision BV (a refined BV) , BVfrac, may be determined based on the sub-pel precision BV, BV′int, the FracPre (e.g., sub-pel position) , and FracDir (e.g., the sub-pel direction) .
[0182] A sub-pel step absDistance may be determined based on sub-pel position FracPre. The sub-pel step absDistance may also be referred to as “fractional sub step absDistance” or “fractional sample step absDistance. ”
[0183] If FracPre is 1 / 4-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>2.
[0184] If FracPre is 1 / 2-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>1.
[0185] If FracPre is 3 / 4-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>2×3.
[0186] For example, if sub-pel precision is 1 / 16-pel precision, FRAC_BITS is equal to 6 and INT_BITS is equal to 2. If FracPre is 1 / 4-pel position, sub-pel step absDistance is equal to 4. If FracPre is 1 / 2-pel position, sub-pel step absDistance is equal to 8. If FracPre is 3 / 4-pel position, sub-pel step absDistance is equal to 12.
[0187] A horizontal offset xDistance and a vertical offset yDistance may be determined based on the sub-pel step absDistance and sub-pel direction FracDir.
[0188] If the FracDir is LEFT_POS, ABOVE_LEFT_POS, or LEFT_BOTTOM_POS, xDistance=-absDistance.
[0189] If FracDir is RIGHT_POS, ABOVE_RIGHT_POS, or RIGHT_BOTTOM_POS, xDistance=absDistance.
[0190] If FracDir is ABOVE_POS, ABOVE_LEFT_POS or ABOVE_RIGHT_POS, yDistance=-absDistance.
[0191] If FracDir is BOTTOM_POS, LEFT_BOTTOM_POS or RIGHT_BOTTOM_POS, yDistance=absDistance.
[0192] A refined BV, BVfrac (xfrac, yfrac) , may be determined based on the sub-pel precision BV, BV′int (x′int, y′int) , the horizontal offset xDistance, and the vertical offset yDistance, where xfrac=x′int+xDistance and yfrac=y′int+yDistance.
[0193] As shown in FIG. 27, for example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2702) is a 1 / 4-pel position, absDistance=4, FracDir is ABOVE_LEFT_POS, xDistance=-4, and yDistance=-4, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int-4 and yfrac=y′int-4.
[0194] In another example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2703) is a 1 / 2-pel position, absDistance=8, FracDir is ABOV_RIGHT_POS, xDistance=8, and yDistance=-8, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int+8 and yfrac=y′int-8.
[0195] In a further example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2704) is a 3 / 4-pel position, absDistance=12, FracDir is LEFT_BOTTOM_POS, xDistance=-12, and yDistance=12, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int-12 and yfrac=y′int+12.
[0196] All or a preset number of sub-pel positions may be traversed in the reconstructed area of the current frame based on template matching cost to determine a refined BV of the current block. For example, the matching cost of each sub-pel precision BV, BVfrac, may be computed, and the matching cost of the sub-pel precision BV, BV′int , corresponding to integer-pel precision BV, BVint, may be determined, and determine a BVfrac with the smallest matching cost in the reconstructed area may be used as the refined BV used for predicting the current block. The type of template may be determined based on the availability of neighboring reference samples.
[0197] Referring to FIG. 5C, when the above left, above and left reference samples are all available, the template shape may be the template shown in (a) , e.g., refTemplateType=1.
[0198] When only the left reference samples are available, the template shape may be the template shown in (b) , e.g., refTemplateType=2.
[0199] When only the above reference samples are available, the template shape may be the template shown in (c) , e.g., refTemplateType=3.
[0200] When the left and above left reference samples are available, the template shape may be the template shown in (d) , e.g., refTemplateType=4.
[0201] When the left and left bottom reference samples are available, the template shape may be the template shown in (e) , e.g., refTemplateType=5.
[0202] When the above and above right reference samples are available, the template shape may be the template shown in (f) , e.g., refTemplateType=6.
[0203] When the above and above left reference samples are available, the template shape may be the template shown in (g) , e.g., refTemplateType=7.
[0204] When the above and left reference samples are available, the template shape may be the template shown in (h) , e.g., refTemplateType=8.
[0205] A cost between two templates may be represented as an error between a template and a reference template. As an example, the cost may be a SAD, which may be calculated using equation (1) , shown above. In another example, the cost may be a SATD, which may be calculated according to equation (2) , shown above. Coding accuracy may be improved using BV refinement.
[0206] In another embodiment, predicting current block may consider BV information of a neighboring reconstructed block. Additional information corresponding to the adjacent block is comprehensively considered and may improve the prediction accuracy. In this process, a BV flip operation 2800, 2801 may be applied, as shown in FIGs. 28A and 28B.
[0207] In some implementations, BV of current block may be determined based on a neighboring reconstructed block. The determined BV of the neighboring reconstructed block may be represented by, and the flip indication, which indicates whether the BV flip operation is a horizontal flip (see FIG. 28A) or a vertical flip (see FIG. 28B) , may be represented by, rribcFlipTypenbr. The adjusted BV determined based on the flip operation may be represented by, and the flip indication may be represented by, rribcFlipTypecur.
[0208] If rribcFlipTypenbr is equal to 0, no flip may be performed. In this example, and rribcFlipTypecur=rribcFlipTypenbr=0.
[0209] If rribcFlipTypenbr is equal to 1, a horizontal flip (see FIG. 28A) may be performed. In this example, and rribcFlipTypecur=rribcFlipTypenbr=1.
[0210] A schematic diagram of the horizontal flip operation 2800 is illustrated in FIG. 28A. Assume the current block is a chroma block and xnbr is the X coordinate of the center position of the adjacent reconstructed block. If the current BV information is determined based on the luma block, then xcur is the X coordinate of the center position of the co-located luma block of the current block. If the current BV information is determined based on the chroma block, then xcur is the X coordinate of the center position of the current block. Assuming the current block is a luma block, the method may be the same.
[0211] If rribcFlipTypenbr is equal to 2, a vertical flip may be performed. In this example, andrribcFlipTypecur=rribcFlipTypenbr=2.
[0212] A schematic diagram of the vertical flip operation 2801 is illustrated in FIG. 28B. Assume the current block is a chroma block and ynbr is the Y coordinate of the center position of the adjacent reconstructed block. If the current BV information is determined based on the luma block, then ycur is the Y coordinate of the center position of the co-located luma block of the current block. If the current BV information is determined based on the chroma block, then ycur is the Y coordinate of the center position of the current block. Assuming the current block is a luma block, the method may be the same.
[0213] Referring again to FIG. 25, at operation 2506, the current block may be predicted based on the adjusted BV.
[0214] A reference block in the current frame may be determined based on the adjusted BV, and a prediction of the current block may be determined based on the reference block.
[0215] The BV used for predicting current block may be stored for use in the prediction of a subsequent block. The BV may be stored in sub-pel precision.
[0216] In one example, the BV may be stored in 1 / 2-pel precision. In one example, the BV may be stored in 1 / 4-pel precision. In one example, the BV may be stored in 1 / 16-pel precision. In one example, the BV may be stored in at least two kinds of sub-pel precisions of 1 / 2-pel precision, 1 / 4-pel precision, and / or 1 / 16-pel precision. In one example, the encoder and decoder may use a unified default sub-pel precision. In one example, the encoder may determine one or more sub-pel precision and signal a precision indication in the bitstream. The decoder may determine to store the BV with a precision based on the precision indication. The precision indication may be sequence level, picture level, slice level and block level. The precision indication may be associated with prediction mode.
[0217] It is possible to store both the integer-pel precision BV and the sub-pel precision BV of the current block. In some implementations, integer-pel precision BV may be derived based on the sub-pel precision BV. However, storing both integer-pel precision BV and sub-pel precision BV consumes more memory than only storing one of them. Thus, to reduce memory consumption sub-pel precision BV rather than integer-pel precision BV may be stored. This improves memory usage efficiency.
[0218] In some implementations, intra prediction unit 205 may use an extrapolation filter-based intra prediction (EIP) mode to derive a prediction of the current block. FIG. 30 illustrates a flow chart of an example method of EIP 3000, in accordance with some embodiments of the present disclosure.
[0219] Referring to FIG. 30, at operation 3002, neighboring reference samples of a current block may be determined. For example, in the EIP prediction process, the samples in the current block may be predicted from the top-left position to the bottom-right position by applying an extrapolation filter to neighboring reconstructed samples or predicted samples. In this process, the input of the EIP filter may be neighboring reference samples, which include neighboring reconstructed samples, neighboring prediction samples, or padding samples.
[0220] At operation 3004, filter parameters for the current block may be determined. Filter parameters of the current block may include one or more of EIP filter length, EIP filter shape, or EIP filter coefficients. Some candidate EIP filters are provided in this disclosure.
[0221] FIG. 31 illustrates a first example of EIP filter shapes 3100, in accordance with some embodiments of the present disclosure, For example, EIP filter shapes 3100 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) in the current block, ci is the filter coefficient, and the is the neighboring reconstructed sample or neighboring predicted sample.
[0222] FIG. 32 illustrates a second example of EIP filter shapes 3200, according to some embodiments of the present disclosure. The EIP filter shapes 3200 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 32) , ci is the filter coefficient, and is the neighboring reconstructed sample or neighboring predicted sample (shown at position X in FIG. 32) , and c14 is the offset term.
[0223] In all three EIP filter shapes, their support areas cover samples above and to the left of the predicting sample, e.g., offsetXi and offsetYi are greater than or equal to zero.
[0224] FIG. 33 demonstrates a third example of EIP filter shapes 3300, according to some embodiments of the present disclosure. The EIP filter shapes 3300 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. Samples above-right and below-left of the predicting sample are considered. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 33) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample or padding sample, and c14 is the offset term.
[0225] Still referring to FIG. 33, samples at position A may be combined with a same coefficient as one input of the EIP filter, and samples at position X may use an independent coefficient. In some implementations, the average of samples at position A may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A may be determined as an input of the EIP filter.
[0226] FIG. 34 illustrates a fourth example of EIP filter shapes 3400, according to some embodiments of the present disclosure. The EIP filter shapes 3400 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. Samples above-right and below-left of the predicting sample may be considered. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 34) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample or padding sample, and c14 is an offset term.
[0227] FIG. 35 illustrates a fifth example of EIP filter shapes 3500, according to some embodiments of the present disclosure. The EIP filter shapes 3500 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 9-tap filters, and the filter length is 9. The prediction samples of the current block are determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , is computed using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 35) , ci is the filter coefficient, the is the neighboring reconstructed sample or neighboring predicted sample, and c8 is an offset term.
[0228] Still referring to FIG. 35, samples at position A may be combined with a same coefficient as one input of the EIP filter, and samples at position X may use an independent coefficient. In some implementations, the average of samples at position A may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A may be determined as an input of the EIP filter.
[0229] FIG. 36 illustrates a sixth example of EIP filter shapes 3600, according to some embodiments of the present disclosure. The EIP filter shapes 3600 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 4-tap filters, and the filter length is 4. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 36) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample, and c4 is an offset term.
[0230] Still referring to FIG. 36, samples at positions A may be combined with a same coefficient. Samples at positions B may be combined with a same coefficient. Samples at positions C may be combined with a same coefficient. Samples at positions D may combined with a same coefficient. In some implementations, the average of samples at position A / B / C / D may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A / B / C / D may be determined as an input of the EIP filter.
[0231] The EIP filter length and / or EIP filter shape may be a default configuration for a current block with a size within a predefined range or for a current block in a selected EIP prediction mode. In one example, for a selected EIP prediction mode, one of the above types of EIP filter shapes may be the default configuration. For example, for regularEIP mode and mergeEIP mode, the EIP filter shapes shown in FIG. 32 may be the default configuration. For bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be the default configuration. In one example, for a current block with a size of 8*8 and the EIP prediction mode being bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be the default configuration.
[0232] The EIP filter length and / or EIP filter shape may be determined based on one or more of the size parameter (s) of the current block, the EIP prediction mode, and / or the type of EIP models. The size parameter of the current block may include, but is not limited to, the number of samples included in the current block, the width and / or the height of the current block, the ratio of the width and height of the current block, the ratio of the height and width of the current block, and so on. The EIP prediction mode may include, but is not limited to, regularEIP, mergeEIP, or bvEIP. The type of EIP model may be a multi-model EIP or a single-model EIP.
[0233] The EIP filter length may be determined based on the size parameter of the current block, where the EIP filter length refers to the number of EIP filter coefficients included in the EIP filter. For example, the length of 15-tap EIP filter is 15, and the length of 9-tap EIP filter is 9. The size parameter of the current block may include, but is not limited to, the number of samples included in the current block, the width and / or the height of the current block, the ratio of the width and height of the current block, the ratio of the height and width of the current block, and so on.
[0234] For example, if the number of samples included in the current block is greater than a predefined threshold, a filter with a first length may be used; if the number of samples included in the current block is less than or equal to the predefined threshold, a filter with a second length may be used, where the first length is less than the second length. For a coding block with a larger number of samples, a filter with a shorter length may be used to limit the computational complexity of the EIP mode.
[0235] The EIP filter length may be determined based on the type of EIP model. If the current block utilizes a multi-model EIP, a filter with a third length may be used; if the current block utilizes a single-model EIP, a filter with a fourth length may be used, where the third length is less than the fourth length. For a coding block using a multi-model EIP, a filter with shorter length may be used to limit the computational complexity of the EIP mode.
[0236] The EIP filter length may be determined based on the type of EIP prediction mode, where EIP prediction mode may include, but is not limited to, regularEIP mode, mergeEIP mode or bvEIP mode. In an example, if the EIP mode is bvEIP mode, a filter with a fifth length may be used; if the current block uses regularEIP mode or mergeEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, if the EIP prediction mode is regularEIP mode, a filter with a fifth length may be used; if the current block uses bvEIP mode or mergeEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, if the EIP prediction mode is mergeEIP mode, a filter with a fifth length may be used; if the current block uses bvEIP mode or regularEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, two EIP modes selected from regularEIP mode, mergeEIP mode, and bvEIP mode may use a filter with a fifth length and the remaining EIP mode may use a filter with a sixth length, where the fifth length is less than the sixth length. The length of the EIP filter may be adaptively determined based on the type of EIP prediction mode to limit the computational complexity of EIP mode.
[0237] It should be noted that the first length, the third length, and the fifth length may be the same value, such as 9; and the second length, the fourth length, and the sixth length may be the same value, such as 15. In some implementations, the first length, the third length, and the fifth length may be different values; and the second length, the fourth length and the sixth length may be the different values.
[0238] The length of EIP filter may include multiple candidate lengths including 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more. The length of EIP filter may be adaptively determined based on one or more size parameters of the current block, the EIP prediction mode, and / or the type of EIP model (s) to achieve a balance between prediction accuracy and coding complexity.
[0239] The type of EIP filter shape (s) may be determined based on one or more the size parameters of the current block, the type of EIP models, the type of EIP prediction mode, and so on.
[0240] In some implementations, the type of EIP filter shape (s) may be determined based on the type of EIP prediction mode. For example, if the EIP prediction mode of the current block is bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be used. Otherwise, the EIP filter shapes shown in FIG. 32 may be used.
[0241] The EIP prediction mode includes one of regularEIP mode, mergeEIP mode, or bvEIP mode.
[0242] The regularEIP mode refers to deriving EIP filter coefficients based on neighboring reconstructed samples in a reconstructed area. The mergeEIP mode refers to deriving EIP filter parameters by inheriting EIP filter parameters from the reconstructed blocks. The bvEIP mode refers to using a block vector to determine the reference area for calculating the EIP filter parameters instead of directly using the adjacent spatial reference area.
[0243] FIG. 37 illustrates example templates 3700 of a reconstructed area used to derive EIP filter parameters, according to some embodiments of the present disclosure. In FIG. 37, three example templates of the reconstructed area, which are used to derive EIP filter parameters, are shown.
[0244] Referring to the example template shown on the left-hand side of FIG. 37, the reconstructed area includes the left and the above neighboring reconstructed area, which may be called the L-shaped reconstructed area. The reconstructed area of this shape has the largest area and may cover the surrounding adjacent above, above-left, above-right, left, and left-bottom samples.
[0245] Referring to the example template shown in the middle of FIG. 37, the reconstructed area includes the above neighboring reconstructed area, which may be called the T-shaped reconstructed area. The reconstructed area of this shape may cover the above, above-left, and above-right samples.
[0246] Referring to the example template shown on the right-hand side of FIG. 37, the reconstructed area includes the left neighboring reconstructed area, which may be called the L-shaped reconstructed area. The reconstructed area of this shape may cover the above-left, left, and left-bottom samples.
[0247] As shown in FIG. 37, taking the square filter as an example, fWidth represents the width of the square filter, and fHeight represents the height of the square filter.
[0248] In the example template shown on the left-hand side of FIG. 37, leftSize represents the width of the left template, that is, the number of columns on the left side of the current block; aboveSize represents the height of the upper template, that is, the number of rows on the upper side of the current block; the size of the template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0249] In the example template shown in the middle of FIG. 37 (e.g., the above template) , leftSize represents the number of columns on the left side of the current block and aboveSize represents the number of rows on the above side of the current block; the size of the upper template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0250] In the example template shown on the right-hand side of FIG. 37 (e.g., the left template) , leftSize represents the number of columns on the left side of the current block and aboveSize represents the number of rows on the above side of the current block; the size of the left template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0251] RegularEIP modes allows the use of three reconstructed areas and three filter shapes; when constructing the regularEIP candidate list, the maximum allowable length is 9. The length of the regularEIP candidate list and the EIP candidate information constructed by the coding units of different shapes or different width and height sizes may be different. For example, a 32×32 coding unit allows the use of all combinations of EIP candidate information; that is, the length of the regularEIP candidate list is 9; while a 4×4 coding unit allows the use of an L-shaped reconstructed area to combine 3 filter shapes, which means the length of the regular EIP candidate list of the coding unit is 3.
[0252] The selected filter moves in the selected reconstructed area either horizontally or vertically with a one-pixel step to construct the auto-correlation matrix and the cross-correlation vector. The calculation of coefficients from the auto-correlation matrix and the cross-correlation vector is the same as that in CCCM.
[0253] To decrease coding complexity, a subsampling process may be optionally performed on the reconstructed area used to derive EIP filter parameters.
[0254] In some implementations, the reconstructed area may be downsampled in the horizontal direction or vertical direction with a step size to obtain a downsampled reconstructed area. In some implementations, the step size may be a predefined value, such as 2. In some implementations, the step size may be determined based on size parameters of the reconstructed area, or size parameters of the current block. The size parameters may include, but are not limited to, the number of samples, the width and / or height, the ratio of width and height, or the ratio of height and width. In some implementations, when obtaining the input of the EIP model training, the reference area samples may be subsampled with a step size, such as 2.
[0255] In some implementations, the reconstructed area may be downsampled by a low-pass filter.
[0256] FIG. 38 illustrates an example of a bvEIP mode 3800, according to some embodiments of the present disclosure. Referring to FIG. 38, a BV may be determined for the current block, and the reference block pointed to by the BV may be used as the reconstructed area for deriving EIP filter parameters for the current block.
[0257] Non-limiting examples for determining the BV of the current block may include the following: 1) performing a motion search in the reconstructed area of the current picture to obtain the BV for the current block, and 2) constructing a BV merge list for the current block by inheriting the BV of the reconstructed block from the spatial and temporal domains.
[0258] In an embodiment, the BV of the current block can be determined based on a motion search. First, one of the templates shown in FIG. 5 may be determined for the current block. Then, a search may be performed to determine the matching template with the smallest template cost corresponding to the template of the current block in the reconstructed area of the current picture.
[0259] FIG. 39 illustrates an example of a reconstructed area 3900, according to some embodiments of the present disclosure.
[0260] Referring to FIG. 39, H represents the height of the current block, and W represents the width of the current block. A template search in the search area from R1 to R6 may be performed, the distortion cost value of the template area may be obtained, and a candidate list of block vector information may be updated and maintained. For example, when encoding the current block, a coarse search in the R1 to R6 area may be performed with a step of 3 or 4 samples using the template shown in FIG. 39; the distortion cost value of each search template area may be determined using, e.g., SAD, MR-SAD, SATD, etc. Based on the principle of ascending cost value, a coarse search block vector list with a length of 33 or other natural value length may be maintained and updated such that the smallest value is located at the top of the list and the largest cost value is located at the bottom of the list.
[0261] Taking each block vector information in the coarse search block vector information list obtained by the coarse search process as the starting point, a fine search with a step of 1 sample point in the search area (R1~R6) corresponding to each block vector information may be performed, and a fine search block vector list with a length of 19 or other natural number may be updated and maintained. The fine search process is similar to the coarse search process in that evaluation indicators such as SAD, MR-SAD, SATD, or MR-SATD may be used to determine the distortion cost, and the fine search list may be maintained with the same sorting principle.
[0262] A matching template with the smallest template cost may be determined, and the area with the same size as the current block corresponding to the matching template as the reference block may be determined. The BV of the current block may be determined based on the current block and the reference block.
[0263] In some implementations, the BV of the current block may be determined by constructing a BV merge list. The BV merge list may include at least one of a spatial merge candidate, a temporal merge candidate, a history-based merge candidate, a pairwise average merge candidate, and / or a default merge candidate. Additional details for constructing the BV merge list are provided above in connection with FIG. 26.
[0264] A reference block as the reconstructed area may be used to derive EIP parameters of current block. To decrease the coding complexity, a subsampling process may be optionally performed on the reconstructed area used for deriving EIP filter parameters.
[0265] In some implementations, the reconstructed area may be downsampled in the horizontal direction or vertical direction with a step size to obtain a downsampled reconstructed area. In some implementations, the step size may be a predefined value, such as 2. In some implementations, the step size may be determined based on size parameters of the reconstructed area or size parameters of the current block. The size parameters may include, but are not limited, to the number of samples, the width and / or height, the ratio of width and height, or the ratio of height and width. In some implementations, while obtaining the input of the EIP model training, the reference area samples may be subsampled with a step size, such as 2.
[0266] In some implementations the reconstructed area is downsampled by a low-pass filter.
[0267] For a current block coded in the EIP mode, an EIP filter may be determined based on rate distortion optimization (RDO) or other cost computations associated with candidate filters in regularEIP prediction mode, mergeEIP prediction mode, and / or bvEIP prediction mode. An EIP merge flag may be signaled to indicate whether the EIP filter is inherited from previous blocks coded in EIP mode. When the EIP merge flag is true, an EIP merge list may be constructed based on one or more of spatial adjacent candidates, spatial non-adjacent candidates, temporal candidates, and / or historical candidates. The position and inclusion order of these candidates may be the same as those used in a cross-component prediction (CCP) merge candidate list. An EIP merge index may be signaled to indicate which EIP merge candidate is selected. The filter parameters, which include the filter shape and the filter coefficients of the selected candidate, may be inherited to code the current block.
[0268] A bvEIP flag may be signaled to indicate whether bvEIP prediction mode is used. If regularEIP prediction mode is used, the relevant syntax element may be signaled to indicate which of the three types of reconstructed area and which of the three filter shapes are used for the current block.
[0269] Referring again to FIG. 30, at operation 3006, prediction samples of the current block may be determined based on the neighboring reference samples and the filter parameters.
[0270] After determining a EIP filter for the current block, prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring prediction samples and the EIP filter.
[0271] After generating the prediction samples of the current block using the EIP filter, an intra prediction mode may be derived by applying the DIMD process to the prediction samples. For instance, a horizontal gradient and a vertical gradient may be calculated for each predicted sample to build a histogram of gradients (HoG) . The intra prediction mode corresponding to the largest histogram count may be used to determine the low-frequency non-separable transform (LFNST) , non-separable primary transform (NSPT) , or multiple transform selection (MTS) transform set.
[0272] In one example, intra prediction unit 205 may use an NN-based intra prediction mode to derive a prediction of the current block. FIG. 29 illustrates an example of NN-based intra prediction 2900, according to some embodiments.
[0273] Block 2901 (block Y) is a current block having w×h samples. The samples 2902 adjacent to block 2901 are reference samples. Let “X” be an input of the NN-based intra prediction process. In one example, “X” may be one or more samples among the samples 2902. In one example, “X” may be derived by filtering one or more samples among the samples 2902. Block 2903 is an output of the NN-based intra prediction process. In one example, block 2903 is an intra prediction of block 2901.
[0274] In one example, inter prediction unit 204 may also derive motion parameters and / or prediction samples using an NN-based method or process.
[0275] In some implementations, prediction unit 202 may conduct a subblock-based prediction process. A subblock-based merge candidate list may be constructed that includes one or more of subblock-based temporal motion vector prediction (SbTMVP) candidate (s) , subblock-based spatial motion vector prediction (SbSMVP) candidate (s) , and / or affine merge candidate (s) .
[0276] FIG. 40 illustrates an example of subblock template generation of subblock-based temporal motion vector prediction (SbTMVP) 4000, according to some embodiments of the present disclosure.
[0277] The temporal motion vector prediction (TMVP) for advanced motion vector prediction (AMVP) and merge mode may be derived by obtaining the motion information from the center or the bottom-right of the collocated block in a signaled collocated picture. For the subblock-based temporal motion vector prediction (SbTMVP) mode, the motion information from the left neighboring position may be used as a motion shift, which may be used to obtain TMVPs at sub-CU level. In some implementations, two collocated pictures, which are the two reference pictures with the least picture-order-count (POC) distance relative to the to-be-coded picture, may be utilized. The motion shift to locate TMVP may be adaptively determined from multiple locations according to template costs. Two motion shift candidate lists may be constructed respectively for the two collocated frames. The motion shifts with the minimum template matching cost may be used to derive SbTMVP or TMVP candidates. In a non-limiting example, at most 4 SbTMVP candidates may be included in the subblock merge candidate list. The SbTMVP candidate with the least template matching cost derived from the first collocated frame may be placed in the first entry without reordering, while other SbTMVP candidates may be sorted together with other candidates. Further, the prediction direction of each subblock template may be determined based on the center subblock. As illustrated in FIG. 40, if the center subblock is uni-predicted, then all the subblock templates are uni-predicted, and vice versa. If the motion vector of the corresponding adjacent subblock at the determined reference list is not available for a subblock template, zero MV may be used for that subblock template. In some implementations, if the motion vector of a corresponding adjacent subblock at the determined reference list is not available for a subblock template, a derived MV may be used for that subblock template, where the derived MV may be derived based on other available subblocks in the determined reference list. Derivation of an MV may be performed based on a weighted average or other combination technique (s) .
[0278] FIG. 41 illustrates an example of subblock-based spatial motion vector prediction (SbSMVP) candidate types 4100, according to some embodiments of the present disclosure.
[0279] Referring to FIG. 41, the subblock motion field of the current block may be inherited based on the motion of the spatial neighboring blocks. FIG. 41 shows an example of derivation of SbSMVP candidates from spatial neighboring blocks. Examples of different SbSMVP candidate types, in which MVs of subblocks are inherited in a directional way, are shown in FIG. 41.
[0280] As shown in diagram (a) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along a horizontal direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the horizontal direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the horizontal direction.
[0281] As shown in diagram (b) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along a vertical direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the vertical direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the vertical direction.
[0282] As shown in diagrams (c) - (e) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along an angular direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the angular direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the angular direction. It should be noted that angular directions may include, but are not limited to, the example angular directions shown in (c) - (e) of FIG. 41. Additional angular directions are available than those shown, and an angular direction may be adaptively determined.
[0283] In some implementations, if a spatial neighboring block in which a spatial neighboring subblock is located is coded based on an intra prediction mode or is not coded based on an inter prediction mode, or if one of spatial neighboring subblocks has no available MV, an MV of the spatial neighboring subblock may be determined to be a default value, such as zero.
[0284] In some implementations, if a spatial neighboring block in which a spatial neighboring subblock is located is coded based on an intra prediction mode or is not coded based on an inter prediction mode, or if one of spatial neighboring subblocks has no available MV, an MV of the spatial neighboring subblock may be determined based on one or more MVs of other spatial neighboring subblocks. In some implementations, the MV of the spatial neighboring subblock may be determined by copying one of the adjacent subblocks of the spatial neighboring subblock. In some implementations, the MV of the spatial neighboring subblock may be determined by a weighted average of two or more adjacent subblocks of the spatial neighboring subblock.
[0285] Using diagram (a) of FIG. 41 as a non-limiting example, if the spatial neighboring block in which the spatial neighboring subblock tagged with MV6 is located is coded based on intra prediction, MV6 of the spatial neighboring subblock is not available. MV6 may be determined to be a zero MV. Otherwise, MV6 may be determined by copying one of MV5, MV7, MV0, or MV8, or MV6 may be determined to be an average MV of two or more of MV5, MV7, MV0, and MV8.
[0286] In some implementations, if an SbSMVP candidate is selected, the motion information, such as motion vector, reference index, and prediction direction of the corresponding neighboring subblock may be copied to the current subblock along a predefined direction, as depicted by the arrows in FIG. 41. In some implementations, if an SbSMVP candidate is selected, the motion information, such as motion vector, reference index, and prediction direction of the current block may be determined based on the motion information of the corresponding neighboring subblock along a predefined direction.
[0287] Up to five SbSMVP candidates may be added to the subblock merge candidate list. SbSMVP candidates may be added between SbTMVP candidates and affine merge candidates.
[0288] All or part of the subblock merge candidates included in the subblock merge candidate list may be reordered. Twenty candidates may be sorted, in some implementations. A candidate may be determined from the list and a merge index may be signaled to indicate the merge candidate.
[0289] Prediction unit 202 may pass the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction also may be called as “intra prediction direction” ) to transform unit 207. In one embodiment, transform unit 207 may use such information to determine transform kernel or a set of transform kernels in the primary transform and / or secondary transform.
[0290] Transform unit 207 performs a first transform, for example, integer transform which is originally designed based on discrete cosine transform (DCT) , on the residual block. Transform unit 207 determines whether a secondary transform is enabled to be applied on a block or not. When enabled, transform unit 207 further determines whether to apply the secondary transform to the coefficients obtained after performing the first transform. Transform unit 207 may generate transform coefficients by applying a transform technique to the residual signal, which is derived by the adder 206 as a difference between the original samples of the current block and the prediction of the current block. For example, the transform technique may include at least one of a discrete cosine transform (DCT) , a discrete sine transform (DST) , a karhunen-loève transform (KLT) , a graph-based transform (GBT) , or a conditionally non-linear transform (CNT) . Here, the GBT means transform obtained from a graph when relationship information between pixels is represented by the graph. The CNT refers to transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks having the same size or may be applied to blocks having a variable size rather than square.
[0291] Transform unit 207 in encoder 200 may determine a region or a sub-block in a current block and performs transform on the samples in this region or sub-block. FIG. 19 demonstrates an example of region transform of a current block 1900. The current block 1900 may be a coding block, a coding unit, a transform unit or a sub-block. Region 1901 is a region in the current block 1900. Transform unit 207 will perform a transform on the samples in region 1901, and set a value of a sample in the remaining region 1902 in the current block 1900 to be equal to 0. The sample in region 1901 may be residual sample of the current block, or original sample of the current block. Transform unit 207 may determine a region parameter indicating the region 1901 including at least one of the following: [Parameter 1] : position of region 1901 and / or [Parameter 2] : size of region 1901.
[0292] As an option, transform unit 207 can also determine that more than one region in a current block may be transformed. FIG. 19 also shows an example in which two regions are in a current block. Transform unit 207 will perform transform on samples in regions 1911 and 1913 in a current block 1910, and set a value of a sample in the remaining region 1912 in the current block 1910 to be equal to 0. The said sample in regions 1911 and 1913 may be residual sample of the current block or the original sample of the current block. Transform unit 207 may determine one or more region parameters indicating region 1911 and 1913 including one or both of [Parameter 1] and [Parameter 2] .
[0293] In the following descriptions, current block 1900 may be taken as an example. The implementation with multiple regions containing non-zero samples (e.g., current block 1911) is carried out using similar method to indicate the regions.
[0294] FIGs. 20A and 20B demonstrates examples 2000, 2001 of region transform. Transform unit 207 performs transform on a sample in a gray region in a current block and sets a value of a sample in the remaining region in a current block to be equal to 0, wherein the sample may be a residual sample after prediction or an original sample. The “arrays” of each gray region demonstrates the transform directions and transform kernel of each direction.
[0295] The gray region in FIGs. 20A and 20B is at a pre-defined position with pre-defined size. For example in FIGs. 20A and 20B, [Parameter 2] may be one or more parameters indicating a split type of a current block (e.g., quad, triple, horizontal or vertical) , and [Parameter 1] may be one or more parameters indicating which one of the regions according to the split type of a current block as indicated by [Parameter 2] is the region on a sample of which transform unit 207 performs a transform. Transform unit 207 may derive a width and a height (e.g., a size) of a gray region according to the abovementioned parameters. For example, given that a size (e.g., width x height) of the current block is 4Wx4H. The size (e.g., represented width x height of a region in the current block) and position (e.g., represented by a location of top-left sample of a region in the current block) of a gray region in FIGs. 20A and 20B are shown below in Tables 5A and 5B. Table 5A: Example sizes of gray region in FIG. 20A Table 5B: Example sizes of gray region in FIG. 20B
[0296] In addition to the abovementioned examples, [Parameter 1] may also include one or more offset parameters (e.g., dx and / or dy in the following examples) , and for each example of a gray region in FIG. 20A, a position of a gray region is shown in Tables 5C and 5D. Table 5C: Example positions of gray region in FIG. 20A Table 5D: Example positions of gray region in FIG. 20B
[0297] FIG. 21 illustrates examples of region transform. Transform unit 207 performs transform on a sample in a gray region in a current block and sets a value of a sample in the remaining region in a current block to be equal to 0, wherein the sample may be a residual sample after prediction or an original sample. The “arrays” of each gray region demonstrates the transform directions and transform kernel of each direction.
[0298] Referring to FIG. 21, a size of a gray region 2101, 2111 or 2121 may be represented as gW x gH, wherein gW is a width of the gray region, and gH is a height of the gray region, and a position of a gray region may be represented by a location of a top-left sample in the gray region in the current block, e.g., (dx, dy) .
[0299] Transform unit 207 may determine values of dx and dy of [Parameter 1] , and gW and gH of [Parameter 2] .
[0300] Optionally, transform unit 207 may first determine a split type of a gray region. For example, a split type of gray region 2101 is “arbitrary type, ” which indicates that gW and gH are smaller than a with and a height of the current block 2100, respectively. In this case, transform unit 207 determines values of dx and dy of [Parameter 1] , and gW and gH of [Parameter 2] for gray region 2101. For example, a split type of gray region 2111 is “vertical type, ” transform unit 207 determines values of dx of [Parameter 1] , and gW of [Parameter 2] for gray region 2111, as dy may be inferred to be 0 and gH may be inferred to be equal to the height of the current block 2110. For example, a split type of gray region 2121 is “horizontal type, ” transform unit 207 determines values of dy of [Parameter 1] , and gH of [Parameter 2] for gray region 2121, as dx may be inferred to be 0 and gW may be inferred to be equal to the width of the current block 2120.
[0301] Transform unit 207 may select an optimal region transform for a current block, and passes the parameters indicating “gray region” of such optimal region transform to entropy coding unit 214 for a signaling in output bitstream of encoder 200.
[0302] In an example, dx and dy are represented in a precision of integral sample in a bitstream.
[0303] In another example, dx and dy are represented in a precision of multiple samples. For example, dx is represented as dx>>shift in a bitstream, wherein shift is an non-negative integer, and “dx>>shift” is arithmetic right shift of a two’s complement integer representation of dx by shift binary digits. “shift” may be a fixed value, for example 1, 2, 3, 4, …, Log2 (MaxCuSize) - 1, wherein MaxCuSize is the maximum value of a width or height of a coding unit and Log2 (MaxCuSize ) is a base-2 logarithm of MaxCuSize. Examples of a representation of dy in a bitstream is the same as those of dx.
[0304] Quantization unit 208 quantizes the coefficients from the transform unit 207.
[0305] Inverse quantization unit 209 performs scaling operations on the quantized coefficients to output reconstructed coefficients. Inverse transform unit 210 performs one or more inverse transforms corresponding to the transforms in transform unit 207 and output reconstructed residual.
[0306] When transform unit 207 uses region transform to code the current block, inverse transform unit 210 gets parameters indicating a position and size of a region in the current block, and performs inverse transform to obtain the reconstructed samples in the region. The reconstructed samples are residual samples. The inverse transform unit 210 sets a value of a sample in the remaining region of the current block to be equal to 0, wherein the said sample is a residual sample.
[0307] Adder 211 calculates reconstructed CU by adding the reconstructed residual and the prediction block of the CU from prediction unit 202. Adder 211 also forwards its output to prediction unit 202 to be used as intra prediction reference. After all the CUs in a picture or a sub-picture have been reconstructed, filtering unit 212 performs in-loop filtering on the reconstructed picture or sub-picture. Filtering unit 212 contains one or more filters, for example, deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF) , luma mapping with chroma scaling (LMCS) filter and neural network based filters. Alternatively, when filtering unit 212 determines that the CU is not used as reference for encoding other CUs, filtering unit 212 performs in-loop filtering on one or more target samples in the CU.
[0308] In one embodiment, filtering unit 212 would process filtering on the reconstructed samples of one or more color components of the current block (e.g., a CU) . The encoder 200 stores the filtered reconstructed samples of one or more color components of the current block in a picture buffer for a picture in which the current block locates. Thus, the prediction unit 202 can use the filtered samples of the current block in encoding the succeeding block of the current block in encoding order. For example, the prediction unit 202 can use the filtered samples of the current block to derive a prediction of succeeding block of the current block in encoding order. For example, the prediction unit 202 as well as other units in encoder 200, would include the filtered samples of the current block in a template and derive of a prediction, reordering candidate modes or parameters, and / or coding parameters using template matching approach. Since the filtering unit 212 suppresses reconstruction distortion of the current block introduced by the lossy source coding, when the filtered sample of the current block is used to encode the succeeding block, the prediction efficiency of the succeeding coding block has been improved, and thus the coding performance of encoder 200 is greatly improved.
[0309] In one embodiment, the filtering unit 212 uses one or more fixed 1D or 2D filters to process the reconstruct sample of the current block. For example, the 1D filter may be a symmetry filter. For example, the 1D filter may be an asymmetry filter. For example, the 2D filter may be a symmetry filter. For example, the 2D filter may be an asymmetry filter. For example, the 2D filter may be a separable filter. For example, the 2D filter may be a non-separable filter.
[0310] In one embodiment, the filtering unit 212 uses one or more adaptive 1D or 2D filters to process the reconstruct sample of the current block. For example, the 1D filter may be a symmetry filter. For example, the 1D filter may be an asymmetry filter. For example, the 2D filter may be a symmetry filter. For example, the 2D filter may be an asymmetry filter. For example, the 2D filter may be a separable filter. For example, the 2D filter may be a non-separable filter.
[0311] In one embodiment, the filtering unit 212 uses one or more neural-network (NN) based filters to process the reconstruct sample of the current block.
[0312] In one embodiment, the filtering unit 212 can use one or more filters of the spatial and / or temporal neighboring blocks of the current block. In one example, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks before filtering which is invoked after reconstructing a picture where the neighboring block locates. In one example, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks after reconstructing a picture where the neighboring block locates. One example is that filtering unit 212 may use the adaptive loop filter (ALF) which is used to filter a temporal neighboring block of the current block. In one example, the filtering unit 212 may select one or more existing filters which are available before filtering the current block. One example is that the filters with parameters are signaled at block layer (e.g., coding tree unit or coding unit) or a layer higher than a block layer of the current block (e.g., video parameter set, sequence parameter set, picture parameter set, adaption parameter set, picture header and / or slice header) . In one example, filtering unit 212 adaptively may derive parameter of one or more filters to process the sample in the current block using spatial and / or temporal samples in one or more templates. In one example, filtering unit 212 adaptively may derive parameter of one or more filters to process the sample in the current block based on the reconstructed samples and the original samples, and filtering unit 212 will pass the filter parameters to entropy coding unit 214 to signal the filter parameters in the bitstream.
[0313] In one embodiment, filtering unit 212 may derive an indication parameter to indicate whether the reconstructed sample in the current block is needed to be filtered or not. For example, the indication parameter may be a 1 bit flag. For example, the indication parameter may be a variable with a number of values indicating not only whether the reconstruct sample is needed to be filtered but also which filter is used. When the variable is equal to 0, the reconstructed sample of the current block will not be filtered; otherwise, the reconstructed sample of the current block is filtered with a filter with an index equal to the value of this variable. Filter unit 212 will pass this indication parameter to entropy coding unit 214 to signal the parameter value in the bitstream.
[0314] In one embodiment, filtering unit 212 may also set indication parameter to indicate which color component may be filtered. Filtering unit 212 can choose to filter one or more of the luma and two chroma components. Filter unit 212 may pass this indication parameter to entropy coding unit 214 to signal the parameter value in the bitstream.
[0315] Output of filtering unit 212 is a decoded picture or sub-picture, which is forwarded to DPB (decoded picture buffer) 213. DPB 213 outputs decoded pictures according to timing and controlling information. Pictures stored in DPB 213 may also be employed as reference for performing inter or intra prediction by prediction unit 202.
[0316] Entropy coding unit 214 converts parameters from units in encoder 200 that are necessary for deriving decoded picture as well as control parameters and supplemental information into binary representations, and writes such binary representations according to syntax structure of each data unit into a generated video bitstream.
[0317] Generally, an NN-based process consume more computational resources or use a dedicated hardware module to support calculations in an NN structure, as compared to a non-NN-based process. For example, one NN-based process may request Neural Processing Unit (NPU) , Graphics Processing Unit (GPU) , and / or any other modules that support the NN-based process may be support an implementation of a NN-based process (e.g., one or more of NN-based inter prediction, NN-based intra prediction, NN-based filtering, etc. ) . In one example, when an encoder determines that NN-based process are used to decode a bitstream generated by the encoder, the encoder signals an indication of this ability in the bitstream for a receiving device or decoder. In one example, when an encoder or sending device has an ability to perform an NN process (e.g., the encoder or sending device is integrated dedicated module, such as NPU, GPU or other modules or chips, to conduct NN process) , the encoder or sending device may signal this to a receiving device or a decoder during session negotiation process.
[0318] To signal the ability of performing NN-based process, in one example, the encoder 200 may signal a Profile in the bitstream, where the Profile specifies that an NN-based process may be used in decoding the bitstream. The Profile may be signaled in a data unit containing Profile data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Profile may specify whether an NN-based process may be used in decoding a bitstream. The Profile may further specify a quantization-error bounds for performing the NN-based process.
[0319] In one example, the encoder 200 can signal a Level in the bitstream, where the Level specifies that an NN-based process may be used in decoding the bitstream. The Level may be signaled in a data unit containing Level data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Level may specify whether an NN-based process may be used in decoding a bitstream. The Level may further specify a quantization-error bounds for performing the NN-based process.
[0320] In one example, the encoder 200 can signal a Tier in the bitstream, where the Tier specifies that an NN-based process may be used in decoding the bitstream. The Tier may be signaled in a data unit containing Tier data. The data unit may be within one or more parameter sets, e.g. decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Tier may specify whether an NN-based process may be used in decoding a bitstream. The Tier may further specify a quantization-error bounds for performing the NN-based process.
[0321] In one example, the encoder 200 can signal a General Constraints Information (GCI) in the bitstream, where the GCI specifies that an NN-based process may be used in decoding the bitstream. The GCI may be signaled in a data unit containing GCI data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The GCI may specify whether NN-based process may be used in decoding a bitstream. The GCI may further specify a quantization-error bounds for performing the NN-based process.
[0322] Encoder 200 may provide controlling parameters to prediction unit 202 to instruct inter prediction unit 204 and intra prediction unit 205 whether one or more template based prediction modes are enabled in the encoding process, or equivalently whether one or more template based prediction modes are disabled in the encoding process. In this disclosure, the embodiments are describes from an aspect way of “enabling” a template based prediction mode. The embodiments from an aspect way of “disabling” a template based prediction mode may be directly derived based on the following descriptions.
[0323] In one example implementation, encoder 200 may perform pre-analysis process on the input video or picture to identify whether template based prediction modes would bring benefit to the coding efficiency, especially perceptual quality. For example, template based prediction modes would hurt the perceptual quality of a picture or video containing complex texture and / or motion. One example of such picture or video would be waterfront with random waves, and the reflection on the surface of the waterfront is random because of the small waves. In this case, encoder 200 will determine to disable all template based prediction modes or enable only one or several template based prediction modes in the encoding process. In one embodiment, encoder 200 can use rate-distortion based method to make such decisions. In one embodiment, encoder 200 can use multi-pass encoding method to make such decisions, wherein encoder 200 codes the input video or picture will all or several template based prediction modes enabled, and then determines which of the template based prediction modes are used in the second pass encoding.
[0324] In one embodiment, the controlling parameters are obtained from configurations for encoder 200. For example, such controlling parameters are set in an encoder configuration file according to a pre-analysis on the input video or picture. For example, controlling parameters are set in an encoder configuration file according to complexity restrictions on encoder 200 and / or a decoder. For example, controlling parameters are set according to conformance indications, such as indications by one or more of Profile, Tier and Level. One example implementation would be a Profile may specify that all of or several of template based prediction modes are enabled or disabled.
[0325] Encoder 200 passes such controlling parameters to the entropy coding unit 214. Entropy coding unit 214 performs entropy coding on the value of such controlling parameters, and then writes the coding bits in to the output bitstream. To avoid any doubt, another implementation is that if the controlling parameters are set according to conformance indications, such as indications by one or more of Profile, Tier and Level, encoder 200 has one option of not passing such controlling parameters to the entropy coding unit 214 to write such controlling parameters into the output bitstream, as the entropy coding unit 214 has written the conformance indications such as Profile, Tier and Level into the output bitstream.
[0326] FIGs. 7A-7E, 8A-8G, and 9A-9E depict examples of the syntax elements for signaling the controlling parameters for template based prediction modes.
[0327] FIGs. 7A-7E illustrate example syntax elements of “high level” or “collective” indications. Entropy coding unit 214 can write the example syntax elements into the output bitstream.
[0328] In one embodiment as shown in FIG. 7A, encoder 200 can set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, prediction unit 202 may use the template based prediction modes in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_flag is equal to 0, prediction unit 202 will not use the template based prediction modes in encoding one or more blocks in the input video or picture.
[0329] In one embodiment as shown in FIG. 7B, encoder 200 can set an indication information for inter prediction. As shown in FIG. 7B, the indication information may be syntax element “tm_prediction_enable_for_inter_flag” to indicate whether template based prediction modes may be used for inter predictions. When tm_prediction_enable_for_inter_flag is equal to 1, inter prediction unit 204 within prediction unit 202 may use the template based inter prediction modes (as shown as example in Table 1) in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_for_inter_flag is equal to 0, inter prediction unit 204 within prediction unit 202 will not use the template based inter prediction modes (as shown as example in Table 1) in encoding one or more blocks in the input video or picture.
[0330] In one embodiment as shown in FIG. 7C, encoder 200 can set an indication information for intra prediction. As shown in FIG. 7C, the indication information may be syntax element “tm_prediction_enable_for_intra_flag” to indicate whether template based prediction modes may be used for intra predictions. When tm_prediction_enable_for_intra_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template based intra prediction modes (as shown as example in Table 3) in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_for_intra_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template based intra prediction modes (as shown as example in Table 3) in encoding one or more blocks in the input video or picture.
[0331] In one embodiment as shown in FIG. 7D, encoder 200 can set an indication information for a set of prediction modes. As shown in FIG. 7D, the indication information may be syntax element “tm_mode_setA_enable_flag” to indicate whether several template based prediction modes may be used for intra and / or inter predictions. The said several template based prediction modes may be viewed or classified as “setA. ” “setA” can contains one or more prediction modes. One example would be that IntraTMP and DMVR are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template based intra prediction modes in “setA” (e.g., IntraTMP in this example) , and inter prediction unit 204 within prediction unit 202 may use the template based inter prediction modes in “setA” (e.g., DMVR in this example) in encoding one or more blocks in the input video or picture. Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template based intra prediction modes in “setA” (e.g., IntraTMP in this example) and inter prediction unit 204 within prediction unit 202 will not use the template based inter prediction modes in “setA” (e.g., DMVR in this example) in encoding one or more blocks in the input video or picture. One example would be that IntraTMP and TIMD (that is, two intra prediction modes) are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in the above example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in the above example) . One example would be that DMVR and BDOF (that is, two inter prediction modes) are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, inter prediction unit 204 within prediction unit 202 may use the template based intra prediction modes in “setA” (e.g., DMVR and BDOF in the above example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, inter prediction unit 204 within prediction unit 202 will not use the template based intra prediction modes in “setA” (e.g., DMVR and BDOF in the above example) .
[0332] In one embodiment as shown in FIG. 7E, encoder 200 can set an indication information for one prediction mode in Table 1 and / or Table 3 (e.g., referred to as “modeA” in this description) . As shown in FIG. 7D, the indication information may be syntax element “tm_modeA_enable_flag” to indicate whether the template based prediction mode ( “modeA” ) may be used for intra and / or inter predictions. When tm_modeA_enable_flag is equal to 1, prediction unit 202 may use the template based prediction mode ( “modeA” ) in encoding one or more blocks in the input video or picture. Otherwise, when tm_modeA_enable_flag is equal to 0, prediction unit 202 will not use the template based prediction mode ( “modeA” ) in encoding one or more blocks in the input video or picture.
[0333] FIGs. 8A-8G illustrate example syntax elements, which may be implemented as a combination of the example syntax elements shown in FIGs. 7A-7E to enable sophisticated controlling options. Entropy coding unit 214 can write the example syntax elements into the output bitstream.
[0334] In one embodiment as shown in FIG. 8A, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for one or more template based prediction modes. FIG. 8A provides an adaption of different template based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A and FIG. 7E.
[0335] In one embodiment as shown in FIG. 8B, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_for_inter_flag” to indicate whether template based prediction modes may be used for inter predictions (as shown as example in Table 1) . When tm_prediction_enable_for_inter_flag is equal to 1, encoder 200 can further set separate indications for one or more template based inter prediction modes. FIG. 8B provides an adaption of different template based inter prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7B and FIG. 7E.
[0336] In one embodiment as shown in FIG. 8C, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_for_intra_flag” to indicate whether template based prediction modes may be used for intra predictions (as shown as example in Table 3) . When tm_prediction_enable_for_intra_flag is equal to 1, encoder 200 can further set separate indications for one or more template based intra prediction modes. FIG. 8C provides an adaption of different template based intra prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7C and FIG. 7E.
[0337] In one embodiment as shown in FIG. 8D, encoder 200 has already set an indication information by syntax element “tm_mode_setA_enable_flag” to indicate whether template based prediction modes may be used for inter and / or intra predictions in “setA. ” When tm_mode_setA_enable_flag is equal to 1, encoder 200 can further set separate indications for one or more template based inter and / or intra prediction modes in “setA. ” FIG. 8D provides an adaption of different template based inter and / or intra prediction modes in “setA” based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7D and FIG. 7E.
[0338] In one embodiment as shown in FIG. 8E, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template based inter prediction modes and template based intra prediction modes. FIG. 8E provides an adaption of different template based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B and FIG. 7C.
[0339] In one embodiment as shown in FIG. 8F, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template based inter prediction modes using “tm_prediction_enable_for_inter_flag” and several template based prediction modes in “setA. ” For example, as template based inter prediction mode may be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” “setA” may contain one or more template based intra prediction modes. FIG. 8E provides an adaption of different template based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B and FIG. 7E.
[0340] In one embodiment as shown in FIG. 8G, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template based intra prediction modes using “tm_prediction_enable_for_intra_flag” and several template based prediction modes in “setA. ” For example, as template based intra prediction mode may be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” “setA” may contain one or more template based inter prediction modes. FIG. 8F provides an adaption of different template based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B and FIG. 7E.
[0341] FIGs. 9A-9E illustrate example syntax elements, which may be implemented as a combination of the example syntax elements shown in FIGs. 7A-7E and / or FIG. 8A-8G to enable sophisticated controlling options. Entropy coding unit 214 can write the example syntax elements into the output bitstream.
[0342] In one embodiment as shown in FIG. 9A, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of FIG. 7E for one template based prediction mode (with the single mode being represented as “modeC” ) , FIG. 7D for one or more template based prediction modes in “setA, ” and FIG. 8D for separate controlling or indication for the one or more template based prediction modes in “set A. ” FIG. 9A provides an adaption of different template based prediction modes based on characteristics of input video or picture.
[0343] In one embodiment as shown in FIG. 9B, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_for_inter_flag” to indicate whether template based inter prediction modes may be used for inter predictions (as shown as example in Table 1) as in FIG. 7B. When tm_prediction_enable_for_inter_flag is equal to 1, encoder 200 may use a combination of FIG. 7E for one template based inter prediction mode (with the single mode being represented as “modeC” ) , FIG. 7D for one or more template based inter prediction modes in “setA, ” and FIG. 8D for separate controlling or indication for the one or more template based inter prediction modes in “set A. ” FIG. 9B provides an adaption of different template based prediction modes based on characteristics of input video or picture.
[0344] In one embodiment as shown in FIG. 9C, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_for_intra_flag” to indicate whether template based intra prediction modes may be used for intra predictions (as shown as example in Table 3) as in FIG. 7C. When tm_prediction_enable_for_intra_flag is equal to 1, encoder 200 may use a combination of FIG. 7E for one template based intra prediction mode (with the single mode being represented as “modeC” ) , FIG. 7D for one or more template based intra prediction modes in “setA, ” and FIG. 8D for separate controlling or indication for the one or more template based intra prediction modes in “set A. ” FIG. 9C provides an adaption of different template based prediction modes based on characteristics of input video or picture.
[0345] In one embodiment as shown in FIG. 9D, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of FIG. 7B for template based inter prediction modes, and as template based inter prediction modes may be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” FIG. 7D for one or more template based intra prediction modes in “setA, ” FIG. 8D for separate controlling or indication for the one or more template based intra prediction modes in “set A, ” and FIG. 7E for one template based intra prediction mode (with the single mode being represented as “modeC” ) .
[0346] In one embodiment as shown in FIG. 9E, encoder 200 has already set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of FIG. 7B for template based intra prediction modes, and as template based intra prediction modes may be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” FIG. 7D for one or more template based inter prediction modes in “setA, ” FIG. 8D for separate controlling or indication for the one or more template based inter prediction modes in “set A, ” and FIG. 7E for one template based inter prediction mode (with the single mode being represented as “modeC” ) .
[0347] In FIGs. 7A-7E, FIGS. 8A-8G and FIG. 9, the descriptor refers to an entropy coding method for the corresponding syntax element. The definition and algorithms of u (1) , u (n) , ue (v) and ae (v) are the same as those in the H. 265 / HEVC standard.
[0348] In an embodiment, entropy coding unit 214 in encoder 200 writes the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in one or more of the following data units in the output bitstream. In one embodiment, an order of the levels from high to low is sequence level, picture level, slice level and block level. The indications or controlling parameters in lower levels may overwrite or override the counterparts in higher levels. Prediction unit 202 in encoder 200 will follow the final valid instruction or controlling parameters in deriving the prediction of the current coding block using the template based prediction modes.
[0349] A sequence level data unit which is valid for all pictures in a coded video sequence. An example of sequence level data unit may be one or more of video parameter set (VPS) , sequence parameter set (SPS) , picture parameter set (PPS) with consistent parameters for all pictures in a coded video sequence, adaptation parameter set (APS) with consistent parameters for all pictures in a coded video sequence, picture header with consistent parameters for all pictures in a coded video sequence, supplemental enhancement information (SEI) with consistent parameters for all pictures in a coded video sequence. Entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in one or more sequence level data units. For example, entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in the same one sequence level data unit. For example, entropy coding unit 214 writes FIGs. 7A-7E syntax elements in SPS, and FIGs. 8A-8G and / or FIGs. 9A-9E syntax elements in PPS and / or APS directly or indirectly referring to the said SPS.
[0350] A picture level data unit which is valid for one picture. An example of picture level data unit may be one or more of picture parameter set (PPS) , adaptation parameter set (APS) with consistent parameters for one picture, picture header, slice header (s) of one picture with consistent parameters, supplemental enhancement information (SEI) with consistent parameters for one picture. Entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in one or more picture level data units, for example, in picture header, PPS and / or APS. For example, entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in the same one picture level data units. For example, entropy coding unit 214 writes FIGS. 7A-7E syntax elements in PPS, and FIGS. 8A-8G and / or FIGS. 9A-9E syntax elements in picture header, slice header, and / or APS directly or indirectly referring to the said PPS. For example, entropy coding unit 214 writes FIGS. 7A-7E syntax elements in picture header, and FIGS. 8A-8G and / or FIGS. 9A-9E syntax elements in slice header, and / or APS directly or indirectly referred to by the said picture header. For example, entropy coding unit 214 writes FIGS. 7A-7E syntax elements in APS, and FIGS. 8A-8G and / or FIGS. 9A-9E syntax elements in picture header and / or slice header directly or indirectly referring to the said APS.
[0351] Slice level data unit which is valid for one slice. An example of slice level data unit may be one or more of slice header, adaptation parameter set (APS) , supplemental enhancement information (SEI) for a slice. Entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in one or more slice level data units. For example, entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in the same one slice level data units. For example, entropy coding unit 214 writes FIGs. 7A-7E syntax elements in slice header, and FIGs. 8A-8G and / or FIGS. 9A-9E syntax elements in APS directly or indirectly referred to by the said slice header. For example, entropy coding unit 214 writes FIGs. 7A-7E syntax elements in slice header, and FIGs. 8A-8G and / or FIGs. 9A-9E syntax elements in slice header directly or indirectly referring to the said APS.
[0352] Block level data unit, which is valid for one or more of CTU, CU, coding block, transform block. Entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in one or more block level data units. For example, entropy coding unit 214 may write the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G and / or FIGs. 9A-9E in the same one block level data units. For example, entropy coding unit 214 writes FIGs. 7A-7E syntax elements in CTU, and FIGs. 8A-8G and / or FIGs. 9A-9E syntax elements in CU that is within the said CTU. For example, entropy coding unit 214 writes FIGs. 7A-7E syntax elements in a first CU, and FIGs. 8A-8G and / or FIGs. 9A-9E syntax elements in the CU that is within the said first CU.
[0353] Encoder 200 could be a computing device with a processor and a storage medium recording an encoding program. When the processor reads and executes the encoding program, the encoder 200 reads an input video and generates corresponding video bitstream.
[0354] Encoder 200 could be a computing device with one or more chips. The units, implemented as integrated circuits, on the chip are of similar functionalities with similar connections as well as data exchangings as the corresponding ones in FIG. 2.
[0355] FIG. 10 illustrates an example of an example implementation of a decoder 1000. Input of the decoder 1000 a bitstream representing a compressed version of a video or a still picture. Output of the decoder 1000 may be a decoded video consisting of a sequence of pictures or a decoded still picture.
[0356] Input bitstream of a decoder 1000 may be a bitstream generated by the encoder 200. Parsing unit 1001 parses the input bitstream and obtains values of syntax elements from the input bitstream. Parsing unit 1001 converts binary representations of syntax elements to numerical values and forwards the numerical values to the units in the decoder 1000 to derive one or more decoded pictures. Parsing unit 1001 may also parse one or more syntax elements from the input bitstream for rendering the decoded pictures.
[0357] Generally, an NN-based process consume more computational resources or use a dedicated hardware module to support calculations in an NN structure, as compared to a non-NN-based process. For example, one NN-based process may request Neural Processing Unit (NPU) , Graphics Processing Unit (GPU) , and / or any other modules that support the NN-based process may be support an implementation of a NN-based process (e.g., one or more of NN-based inter prediction, NN-based intra prediction, NN-based filtering, etc. ) . In one example, when an encoder determines that NN-based process are used to decode a bitstream generated by the encoder, the encoder signals an indication of this ability in the bitstream for a receiving device or decoder. In one example, when an encoder or sending device has an ability to perform an NN process (e.g., the encoder or sending device is integrated dedicated module, such as NPU, GPU or other modules or chips, to conduct NN process) , the encoder or sending device may signal this to a receiving device or a decoder during session negotiation process.
[0358] To signal the ability of performing NN-based process, in one example, the decoder 1000 may signal a Profile in the bitstream, where the Profile specifies that an NN-based process may be used in decoding the bitstream. The Profile may be signaled in a data unit containing Profile data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Profile may specify whether an NN-based process may be used in decoding a bitstream. The Profile may further specify a quantization-error bounds for performing the NN-based process.
[0359] In one example, the decoder 1000 can signal a Level in the bitstream, where the Level specifies that an NN-based process may be used in decoding the bitstream. The Level may be signaled in a data unit containing Level data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Level may specify whether an NN-based process may be used in decoding a bitstream. The Level may further specify a quantization-error bounds for performing the NN-based process.
[0360] In one example, the decoder 1000 can signal a Tier in the bitstream, where the Tier specifies that an NN-based process may be used in decoding the bitstream. The Tier may be signaled in a data unit containing Tier data. The data unit may be within one or more parameter sets, e.g. decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The Tier may specify whether an NN-based process may be used in decoding a bitstream. The Tier may further specify a quantization-error bounds for performing the NN-based process.
[0361] In one example, the decoder 1000 can signal a General Constraints Information (GCI) in the bitstream, where the GCI specifies that an NN-based process may be used in decoding the bitstream. The GCI may be signaled in a data unit containing GCI data. The data unit may be within one or more parameter sets, e.g., decoder parameter set, video parameter set, sequence parameter set, adaption parameter set, etc. In one example, the data unit may also be within a packet of session unit data, e.g., a payload of session negotiation packet. The GCI may specify whether NN-based process may be used in decoding a bitstream. The GCI may further specify a quantization-error bounds for performing the NN-based process.
[0362] As mentioned above, FIGs. 7A-7E illustrate example syntax elements of “high level” or “collective” indications. Parsing unit 1001 can process the example syntax elements in the bitstream to determine corresponding values of such syntax elements.
[0363] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as show in FIG. 7A. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, prediction unit 1002 may use the template based prediction modes in decoding one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_flag is equal to 0, prediction unit 1002 will not use the template based prediction modes in decoding one or more blocks in the video or picture bitstream.
[0364] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 7B. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag” which indicates whether template based prediction modes may be used for inter predictions. When tm_prediction_enable_for_inter_flag is equal to 1, inter prediction unit 1003 within prediction unit 1002 may use the template based inter prediction modes (as shown as example in Table 1) in decoding one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_for_inter_flag is equal to 0, inter prediction unit 1003 within prediction unit 1002 will not use the template based inter prediction modes (as shown as example in Table 1) in decoding one or more blocks in the video or picture bitstream.
[0365] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 7C. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag” which indicates whether template based prediction modes may be used for intra predictions. When tm_prediction_enable_for_intra_flag is equal to 1, intra prediction unit 1004 within prediction unit 1002 may use the template based intra prediction modes (as shown as example in Table 3) in decoding one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_for_intra_flag is equal to 0, intra prediction unit 1004 within prediction unit 1002 will not use the template based intra prediction modes (as shown as example in Table 3) in decoding one or more blocks in the video or picture bitstream.
[0366] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 7D. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_mode_setA_enable_flag” which indicates whether several template based prediction modes may be used for intra and / or inter predictions. The said several template based prediction modes may be viewed or classified as “setA” . “setA” can contains one or more prediction modes. One example would be that IntraTMP and DMVR are within “setA” . When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 1004 within prediction unit 1002 may use the template based intra prediction modes in “setA” (e.g., IntraTMP in this example) , and inter prediction unit 1003 within prediction unit 1002 may use the template based inter prediction modes in “setA” (e.g., DMVR in this example) in decoding one or more blocks in the video or picture bitstream. Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 1004 within prediction unit 1002 will not use the template based intra prediction modes in “setA” (e.g., IntraTMP in this example) and inter prediction unit 1003 within prediction unit 1002 will not use the template based inter prediction modes in “setA” (e.g., DMVR in this example) in decoding one or more blocks in the video or picture bitstream. One example would be that IntraTMP and TIMD (that is, two intra prediction modes) are within “setA” . When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 1004 within prediction unit 1002 may use the template based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in the above example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 1004 within prediction unit 1002 will not use the template based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in the above example) . One example would be that DMVR and BDOF (that is, two inter prediction modes) are within “setA” . When tm_mode_setA_enable_flag is equal to 1, inter prediction unit 1003 within prediction unit 1002 may use the template based intra prediction modes in “setA” (e.g., DMVR and BDOF in the above example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, inter prediction unit 1003 within prediction unit 1002 will not use the template based intra prediction modes in “setA” (e.g., DMVR and BDOF in the above example) .
[0367] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 7E. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_modeA_enable_flag” which indicates whether the template based prediction mode ( “modeA” ) may be used for intra and / or inter predictions. When tm_modeA_enable_flag is equal to 1, prediction unit 1002 may use the template based prediction mode ( “modeA” ) in decoding one or more blocks in the video or picture bitstream. Otherwise, when tm_modeA_enable_flag is equal to 0, prediction unit 1002 will not use the template based prediction mode ( “modeA” ) in decoding one or more blocks in the video or picture bitstream.
[0368] FIGs. 8A-8G illustrate example syntax elements, which may be implemented as a combination of the example syntax elements shown in FIGs. 7A-7E to enable sophisticated controlling options. Parsing unit 1001 can process the example syntax elements in the bitstream to determine corresponding values of such syntax elements.
[0369] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8A. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may further determine separate indications for one or more template based prediction modes according to the additional syntax elements in the bitstream. Parsing unit 1001 determines further the controlling or indication by the example syntax elements in identical or similar way to those in FIG. 7A and FIG. 7E.
[0370] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8B. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag” which indicates whether template based prediction modes may be used for inter predictions (as shown as example in Table 1) . When tm_prediction_enable_for_inter_flag is equal to 1, parsing unit 1001 may further determine separate indications for one or more template based inter prediction modes according to the additional syntax elements in the bitstream. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7B and FIG. 7E.
[0371] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8C. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag” which indicates whether template based prediction modes may be used for intra predictions (as shown as example in Table 3) . When tm_prediction_enable_for_intra_flag is equal to 1, parsing unit 1001 may further determine separate indications for one or more template based intra prediction modes according to the additional syntax elements in the bitstream. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7C and FIG. 7E.
[0372] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8D. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_mode_setA_enable_flag” which indicates whether template based prediction modes may be used for inter and / or intra predictions in “setA” . When tm_mode_setA_enable_flag is equal to 1, parsing unit 1001 may further determine separate indications for one or more template based inter and / or intra prediction modes in “setA” according to the additional syntax elements in the bitstream. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7D and FIG. 7E.
[0373] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8E. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may further determine separate indications for template based inter prediction modes and template based intra prediction modes according to the additional syntax elements in the bitstream. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B, and FIG. 7C.
[0374] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8F. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may further determine separate indications for template based inter prediction modes using “tm_prediction_enable_for_inter_flag” and several template based prediction modes in “setA” according to the additional syntax elements in the bitstream. For example, as template based inter prediction mode may be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” “setA” may contain one or more template based intra prediction modes. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B and FIG. 7D.
[0375] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 8G. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) . When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may further determine separate indications for template based intra prediction modes using “tm_prediction_enable_for_intra_flag” and several template based prediction modes in “setA” according to the additional syntax elements in the bitstream. For example, as template based intra prediction modes may be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” “setA” may contain one or more template based inter prediction modes. Parsing unit 1001 determines further the controlling or indication by the example syntax elements are identical or similar to those in FIG. 7A, FIG. 7B and FIG. 7D.
[0376] FIGs. 9A-9E illustrate example syntax elements, which may be implemented as a combination of the example syntax elements shown in FIG. 7 and / or FIG. 8 to enable sophisticated controlling options. Parsing unit 1001 can process the example syntax elements in the bitstream to determine corresponding values of such syntax elements.
[0377] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 9A. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” which indicates whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may, according to a combination of syntax elements in FIG. 7E for one template based prediction mode (with the single mode being represented as “modeC” ) , in FIG. 7D for one or more template based prediction modes in “setA, ” and in FIG. 8D for separate controlling or indication for the one or more template based prediction modes in “set A, ” further determine corresponding indication or controlling parameters.
[0378] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 9B. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag” which indicates whether template based inter prediction modes may be used for inter predictions (as shown as example in Table 1) as in FIG. 7B. When tm_prediction_enable_for_inter_flag is equal to 1, parsing unit 1001 may, according to a combination of syntax elements in FIG. 7E for one template based inter prediction mode (with the single mode being represented as “modeC” ) , in FIG. 7D for one or more template based inter prediction modes in “setA, ” and in FIG. 8D for separate controlling or indication for the one or more template based inter prediction modes in “set A, ” further determine corresponding indication or controlling parameters.
[0379] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 9C. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag” which indicates whether template based intra prediction modes may be used for intra predictions (as shown as example in Table 3) as in FIG. 7C. When tm_prediction_enable_for_intra_flag is equal to 1, parsing unit 1001 may, according to a combination of syntax elements in FIG. 7E for one template based intra prediction mode (with the single mode being represented as “modeC” ) , in FIG. 7D for one or more template based intra prediction modes in “setA, ” and in FIG. 8D for separate controlling or indication for the one or more template based intra prediction modes in “set A, ” further determine corresponding indication or controlling parameters.
[0380] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 9D. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may, according to a combination of syntax elements in FIG. 7B for template based inter prediction modes, and as template based inter prediction modes may be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” in FIG. 7D for one or more template based intra prediction modes in “setA, ” in FIG. 8D for separate controlling or indication for the one or more template based intra prediction modes in “set A, ” and in FIG. 7E for one template based intra prediction mode (with the single mode being represented as “modeC” ) .
[0381] In an embodiment, parsing unit 1001 can process the bitstream containing one or more of the syntax elements as shown in FIG. 9E. Parsing unit 1001 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” to indicate whether template based prediction modes may be used for inter and intra predictions (as shown as example in Table 1 and Table 3) as in FIG. 7A. When tm_prediction_enable_flag is equal to 1, parsing unit 1001 may, according to a combination of syntax elements in FIG. 7B for template based intra prediction modes, and as template based intra prediction modes may be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” in FIG. 7D for one or more template based inter prediction modes in “setA, ” in FIG. 8D for separate controlling or indication for the one or more template based inter prediction modes in “set A, ” and in FIG. 7E for one template based inter prediction mode (with the single mode being represented as “modeC” ) .
[0382] In FIGs. 7A-7E, FIGs. 8A-8G and FIGs. 9A-9E, the descriptor refers to an entropy decoding method for the corresponding syntax element. The definition and algorithms of u (1) , u (n) , ue(v) and ae (v) are the same as those in the H. 265 / HEVC standard.
[0383] In an embodiment, parsing unit 1001 determines the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G, and / or FIGs. 9A-9E from one or more of the following data units in the input bitstream. To avoid any doubt, the “parameters” mentioned below refers to the example syntax elements in FIGs. 7A-7E, FIGs. 8A-8G, and / or FIGs. 9A-9E indicating whether one or more template based prediction modes are enabled in decoding one or more blocks in the input bitstream. In one embodiment, an order of the levels from high to low is sequence level, picture level, slice level and block level. The indications or controlling parameters in lower levels may overwrite or override the counterparts in higher levels. Parsing unit 1001 will pass the final valid indication or controlling parameters to prediction unit 1002 in decoder 1000. Prediction unit 1002 will follow the final valid instruction or controlling parameters in deriving the prediction of one or more blocks using the template based prediction modes.
[0384] A sequence level data unit which is valid for all pictures in a coded video sequence. An example of sequence level data unit may be one or more of video parameter set (VPS) , sequence parameter set (SPS) , picture parameter set (PPS) with consistent parameters for all pictures in a coded video sequence, adaptation parameter set (APS) with consistent parameters for all pictures in a coded video sequence, picture header with consistent parameters for all pictures in a coded video sequence, supplemental enhancement information (SEI) with consistent parameters for all pictures in a coded video sequence.
[0385] If the parsing unit 1001 determines the instruction or controlling parameters according to syntax elements in a sequence level data unit, prediction unit 1002 in decoder 1000 may use the template based prediction modes that may be enabled according to the instruction or controlling parameters from the parsing unit 1001 in decoding one or more blocks in this coded video sequence in the input bitstream. In an embodiment, the said instruction or controlling parameters determined by the parsing unit 1001 from a sequence level data unit can also be overwritten or overridden by the instruction or controlling parameters determined by parsing lower level (e.g., one or more of picture level, slice level and block level) data units.
[0386] A picture level data unit which is valid for one picture. An example of picture level data unit may be one or more of picture parameter set (PPS) , adaptation parameter set (APS) with consistent parameters for one picture, picture header, slice header (s) of one picture with consistent parameters, supplemental enhancement information (SEI) with consistent parameters for one picture.
[0387] If the parsing unit 1001 determines the instruction or controlling parameters according to syntax elements in a picture level data unit, prediction unit 1002 in decoder 1000 may use the template based prediction modes that may be enabled according to the instruction or controlling parameters from the parsing unit 1001 in decoding one or more blocks in this picture in the input bitstream. In an embodiment, the said instruction or controlling parameters determined by the parsing unit 1001 from a picture level data unit can also be overwritten or overridden by the instruction or controlling parameters determined by parsing lower level (e.g., one or more of slice level and block level) data units. In an embodiment, the instruction or controlling parameters determined by the parsing unit 1001 from a picture level data unit can also overwrite or override the instruction or controlling parameters determined by parsing higher level (e.g., sequence level) data unit.
[0388] Slice level data unit which is valid for one slice. An example of slice level data unit may be one or more of slice header, adaptation parameter set (APS) , supplemental enhancement information (SEI) for a slice. In a slice level data unit, ae (v) would not be used.
[0389] If the parsing unit 1001 determines the instruction or controlling parameters according to syntax elements in a slice level data unit, prediction unit 1002 in decoder 1000 may use the template based prediction modes that may be enabled according to the instruction or controlling parameters from the parsing unit 1001 in decoding one or more blocks in this slice in the input bitstream. In an embodiment, the said instruction or controlling parameters determined by the parsing unit 1001 from a slice level data unit can also be overwritten or overridden by the cost function determined by parsing lower level (e.g., block level) data units. In an embodiment, the said instruction or controlling parameters determined by the parsing unit 1001 from a picture level data unit can also overwrite or override the said instruction or controlling parameters determined by parsing higher level (e.g., one or more of sequence level, picture level) data units.
[0390] Block level data unit, which is valid for one or more of CTU, CU, coding block, transform block. In a block level data unit, ae (v) would be used.
[0391] If the parsing unit 1001 determines the instruction or controlling parameters according to syntax elements in a block level data unit, prediction unit 1002 in decoder 1000 may use the template based prediction modes that may be enabled according to the instruction or controlling parameters from the parsing unit 1001 in decoding the block in the bitstream. In an embodiment, the instruction or controlling parameters determined by the parsing unit 1001 from a block level data unit can also overwrite or override the cost function determined by parsing higher level (e.g., one or more of sequence level, picture level, slice level) data units.
[0392] Parsing unit 1001 forwards the instruction or controlling parameters to indicate template based prediction modes that may be enabled in deriving prediction of a block by the prediction unit 1002, the values of other syntax elements, as well as one or more variables set or determined according the values of syntax elements, for deriving one or more decoded pictures to the units in the decoder 1000. Prediction unit 1002 determines a prediction block of a current decoding block (e.g., a CU) . When it is indicated that an inter prediction mode is used to decoding the current decoding block, prediction unit 1002 passes relative parameters from parsing unit 1001 to inter prediction unit 1003 to derive inter prediction block. When it is indicated that an intra prediction mode is used to decoding the current decoding block, prediction unit 1002 passes relative parameters from parsing unit 1001 to intra prediction unit 1004 to derive intra prediction block.
[0393] In one embodiment, parsing unit 1001 may determine controlling parameters according to conformance indications, such as indications by one or more of Profile, Tier and Level. One example implementation would be a Profile may specify that all of or several of template based prediction modes are enabled or disabled.
[0394] In one embodiment, if a template based inter prediction mode is enabled and used in decoding a block, inter prediction unit 1003 may derive a reference template. One example of deriving the reference template of a current decoding block is the same as that shown in FIG. 4 using the template as an example shown in FIG. 5. In one example, the prediction block of the current decoding block may be derived based on the reference template, for example, in a same way as that described for encoder 200. In an embodiment, inter prediction unit 1003 can use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 1, inter prediction unit 1003 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, inter prediction unit 1003 can use SAD or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In an embodiment, inter prediction unit 1003 can use a single cost function for all prediction modes.
[0395] In one embodiment, if a template based intra prediction mode is enabled and used in decoding a block, intra prediction unit 1004 may derive a reference template with the determined cost function. One example of deriving the reference template of a current decoding block is the same as that shown in FIG. 6 using the template as an example shown in FIG. 5. In one example, the prediction block of the current decoding block may be derived based on the reference template, for example, in a same way as that described for encoder 200. In an embodiment, intra prediction unit 1004 can use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 3, intra prediction unit 1004 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, intra prediction unit 1004 can use SAD, or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In an embodiment, intra prediction unit 1004 can use a single cost function for all prediction modes.
[0396] Prediction unit 1002 can also derive intra prediction mode or angular prediction direction of the current CU based on one or more reference samples. This derivation process is identical to the counterpart of Prediction unit 202. FIG. 16 illustrates an example of a current block 1601 and its reference sample 1602, 1603. For example, the reference sample 1602, 1603, which are marked as black dot outside the current block 1601, may be one or more samples in a template ( “L-shape” template consisting of the black dots in FIG. 16) as shown in FIGs. 5A-5C. For example, a gradient of reference sample 1602, 1603 may be derived by applying one or more filters to process one or more samples in a template as shown in FIGs. 5A-5C. Prediction unit 1002 first may derive a gradient of a reference sample using an operator. Generally, the operator may be used to detect an edge or a gradient in a picture. The operator may be a 2-dimentional (2D) M x N filter, wherein M and N are positive integers, and M may be equal to or different from N.
[0397] One example of the operator is a Sobel filter. A first example of a 3x3 Sobel filter 4200 is depicted in FIG. 42A, and a second example of a 3x3 Sobel filter 4225 is depicted in FIG. 42B.
[0398] Another example of the operator is an Edge filter. A first example of an Edge filter 4250 is depicted in FIG. 42C, and a second example of an Edge filter 4275 is depicted in FIG. 42D.
[0399] In an example, prediction unit 1002 can choose different operators according to a width and / or a height of the current block. For example, prediction unit 1002 uses smaller operator for smaller block, and uses larger operator for larger block. One example would be that prediction unit 1002 uses the abovementioned Edge filter when a size (width by height or “width x height” ) of the current block is 4x4, 4x8 or 8x4, and uses the abovementioned Sobel filter for other sizes of the current block.
[0400] Prediction unit 1002 may derive a Histogram of Gradients (HoG) based on analyzing one or more reference sample marked in black dots in FIG. 16. The template in FIG. 16 comprises three reference sample lines above and three reference sample columns on the left of the current block. The HoG is derived by accumulating the magnitudes of one or more gradients at one or more given direction, for one, a part of, or all of the reference samples as shown in FIG. 16. One or more directions indicated by the gradients with highest or higher cumulative magnitudes are to be angular prediction direction or intra prediction mode of the current block.
[0401] As an example, when prediction unit 1002 uses one reference sample as shown in FIG. 16 to derive a direction, the prediction unit 1002 can determine the direction as the one indicated by a gradient derived at this reference sample.
[0402] As an example, when prediction unit 1002 uses a part of reference samples as shown in FIG. 16 to derive directions, prediction unit 1002 can choose to a preset number of samples from the reference samples. For example, prediction unit 1002 choose J reference samples above the current block and K reference samples left to the current block, wherein J and K are integers greater than or equal to 0. For example, both J and K are equal to 2, J equal to 4 and K equal to 8, or J plus K equal to 4. Prediction unit 1002 may derive the gradients at the selected reference samples using the operator, and derive the HoG. One or more directions indicated by the gradients with highest or higher cumulative magnitudes are to be angular prediction direction or intra prediction mode of the current block.
[0403] As an example, prediction unit 1002 can adaptively determine one or more reference samples used for deriving a HoG. Prediction unit 1002 uses a preset scanning order of the reference samples. When scanning a reference sample, prediction unit 1002 may derive a gradient at this reference sample, and updates the accumulation the magnitudes according to this gradient at one or more given direction in HoG. When prediction unit 1002 determines that the total cumulative amplitude is greater or equal than a given threshold, prediction unit 1002 will stop scanning the remaining reference sample and deriving new gradient. The resulted HoG at the termination of prediction unit 1002’s scanning, is determined as the HoG to derive intra prediction mode or angular prediction direction.
[0404] Examples of the abovementioned preset scanning order of the reference samples may be the following.
[0405] For example, a scanning order (A) may be scanning the left column of reference samples as shown in FIG. 16 from bottom to top; a scanning order (B) may be scanning the left column of reference samples as shown in FIG. 16 from top to bottom; a scanning order (C) may be scanning the above line of reference samples as shown in FIG. 16 from left to right; a scanning order (D) may be scanning the above line of reference samples as shown in FIG. 16 from left to right.
[0406] An example of the preset scanning order may be one or more of “first order (A) then order (C) , ” “first order (A) then order (D) , ” “first order (B) then order (C) , ” “first order (B) then order (D) , ” “first order (C) then order (A) , ” “first order (D) then order (A) , ” “first order (C) then order (B) ” and “first order (D) then order (B) ” .
[0407] An example of the preset scanning order may be an interleaving manner. For example, one or multiple reference samples from the left column and then one or multiple second reference samples from the above line and then one or multiple third reference sample from left column. For example, one or multiple reference samples from the above line and then one or multiple second reference samples from the left column and then one or multiple third reference sample from above line. Additionally, as an example, the scanning order of samples in the left column may be one or more of order (A) and (B) , and the scanning order of samples in the left column may be one or more of order (C) and (D) .
[0408] Prediction unit 1002 may use the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction also may be called as “intra prediction direction” ) to derive a prediction of the current block. For example, prediction unit 1002 may pass the derived one or more intra prediction modes or one or more angular prediction directions to intra prediction unit 205. In one embodiment, Intra prediction unit 1004 may derive a prediction of the current block by fusing one or more predictions corresponding to the derived intra prediction modes. In one embodiment, Intra prediction unit 1004 may derive a prediction of the current block by fusing one or more predictions determined according to the derived intra prediction modes and one or more predictions determined according to one or more preset mode (for example, Planar mode, DC mode and cross-component prediction mode and etc. ) . One example may be decoder side intra mode Mode Derivation (DIMD) .
[0409] In an embodiment, a DIMD_flag may be determined to enable or disable the Mode Derivation (DIMD) .
[0410] If the DIMD_flag is determined to be 1 (enable) , whether to use the Occurrence-based intra coding (OBIC) mode may be determined. For example, an OBIC_flag may be determined to enable or disable the OBIC mode.
[0411] The DIMD_flag may be signaled in the sequence level, picture level, slice level or block level and so on.
[0412] The OBIC_flag may be signaled in the sequence level, picture level, slice level or block level and so on.
[0413] In an embodiment, the OBIC mode may derive the intra prediction modes of the current block based on the sample-wise occurrence of the intra modes in the spatial neighborhood of the block. For this, adjacent and non-adjacent spatial neighboring blocks are checked and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradient (HoG) as in DIMD, the OBIC method uses the Histogram of occurrence (HoC) , which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an IPM mode, the occurrence of the mode in that particular block is calculated as: HoC [IPM] += uiWidth * uiHeight, where uiWidth and uiHeight are the width and height of a spatial neighboring block.
[0414] The occurrences of the existing modes from the spatial neighborhood blocks are accumulated into the histogram, adjacent and non-adjacent spatial neighboring blocks are checked and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradient (HoG) as in DIMD, the OBIC method uses the Histogram of occurrence (HoC) , which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an Intra Predication Mode (IPM) mode, the occurrence of the mode in that particular block is calculated as: HoC [IPM] += uiWidth * uiHeight, where uiWidth and uiHeight are the width and height of a spatial neighboring block.
[0415] The occurrences of the existing modes from the spatial neighborhood blocks are accumulated into the histogram.
[0416] FIG. 17 shows the non-adjacent spatial neighboring blocks that are used in OBIC mode’s HoC generation.
[0417] One or multiple (e.g., up to five angular modes) with the highest occurrence along with the planar mode or block vector based prediction (same as in DIMD) are selected from the HoC and used for final prediction by blending the prediction of the selected modes.
[0418] Some blocks, mentioned below, use more than one intra mode for prediction. In such cases, all the intra modes of such blocks are selected and used when creating the OBIC histogram: DIMD may use up to up to 5 angular modes, TIMD may use up to up to 2 modes, SGPM may use up to 2 modes, and OBIC may use up to up to 5 angular modes.
[0419] Moreover, the virtual intra prediction modes (VIPMs) of the following blocks are considered only in inter slices when creating the histogram of OBIC mode: MIP block, IntraTMP block, IBC block, and EIP block.
[0420] The blending weights are calculated similarly to the DIMD mode, but instead of using gradient values from the template, the occurrence values are used for OBIC. Moreover, the planar mode’s weight is also decided similarly to the DIMD mode.
[0421] As mentioned above, FIG. 18 shows an example of a histogram of occurrences (HoC) of intra predication modes in the spatial neighborhood of a CU.
[0422] In an embodiment, in order to decrease the buffer memory, the OBIC method may use non-sample-wise occurrences, such as block-wise occurrences. For example, if a uiWidth×uiHeight block is coded with an IPM mode, the block-wise occurrence of the mode in that particular block is calculated as: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where the variable shift1 and shift2 are both positive integers greater than or equal to 1. When the variable shift1 and shift2 are set to 1, the occurrence type 2×2 block-wise occurrence is used. The variable shift 1 or shift 2 can also be set to 2, 3, 4, 5 and so on. It is not limited that the shift1 is equal to shift2.
[0423] In an embodiment, the shift1 or shift 2 may be determined according to the size (width or height) of the current block. For example, if the size of the current block is 4×4, the 2×2 block-wise occurrence may be used.
[0424] In an embodiment, the shift1 or shift2 may be determined by using the flag sps_log2_min_luma_coding_block_size_minus2. For example, parse a bitstream and obtain the value of sps_log2_min_luma_coding_block_size_minus2. If the value of sps_log2_min_luma_coding_block_size_minus2 is equal to 2, the size of the current block is determined to 4×4. Then the occurrence type may be obtained through the determined size of the current block.
[0425] In an embodiment, the block-wise occurrence may be determined through a look-up table. Some examples of look-up tables are shown in Tables 3A-3E, without limitation.
[0426] In an embodiment, a confidence level can also be used to calculate the occurrence. If there is a very large size block neighboring a small size block, the IPM of the large size block is considered as a low confidence level block and the large size block is not used to calculate the occurrence of the current block. For example, if a 64×64 block neighbors a 2×2 block, the 64×64 block may not be used to calculate the occurrence of the 2×2 block for the reason that the 2×2 block is too small than the 64×64 block. If the 64×64 block is used to calculate, it will negatively interfere with the histogram statistics. The size of current block is a factor the determined the confidence level of a neighbor block.
[0427] In an embodiment, the OBIC mode may be used to luma blocks.
[0428] In an embodiment, the OBIC mode may be used to chroma blocks.
[0429] To determine the intra prediction mode of the current block according to parameter from parsing unit 1001, Intra prediction unit 1004 may derive a most probable mode (MPM) list containing one or more intra prediction modes. If the parameter from parsing unit 1001 indicates that intra prediction mode of the current block is one candidate mode in MPM list, intra prediction unit 1004 may determine the intra prediction mode according to an index from parsing unit 1001 which indicates this intra prediction mode in MPM list. Otherwise, if the parameter from parsing unit 1001 indicates that the intra prediction mode of the current block is not in the MPM list, intra prediction unit 1004 will derive an index or indication parameter of this intra prediction mode according to further parameter from parsing unit 1001.
[0430] As mentioned above, FIG. 22 illustrates examples of adjacent reference samples of the current block 2201. The numbers on adjacent reference samples show an example of scanning order of the adjacent blocks by intra prediction unit 1004. In one example, intra prediction unit 1004 scans prediction mode of the adjacent reference sample to derive one or more MPMs. If the mode of an adjacent block is not a BV based intra prediction mode (e.g., planar mode, DC mode, angular prediction mode, etc. ) , intra prediction unit 1004 may include this intra mode as a candidate mode in MPM list. In an example, spatial reference samples can also be one or more non-adjacent blocks shown in FIG. 17.
[0431] As mentioned above, FIG. 23 illustrates a first example of an adjacent block of a BV-based prediction mode, according to some embodiments of the present disclosure. In FIG. 23, 2300 is a current picture 2300 and 2301 is a current block 2301.
[0432] For example, current block 2301 is a block in an intra coded slice. For example, current block 2301 is a block in an intra coded picture (e.g., a picture that may be employed as an access point, such as Instantaneous Decoding Refresh picture, Clean Random Access picture, Broken Link Access picture, etc. ) .
[0433] Block 2302 is an adjacent block of current block 2301, and block 2302 is coded using a BV based coding mode, e.g., IBC or IntraTMP. BV (2302) is a BV of block 2302 and indicates a reference block 2303. If block 2303 is an intra coded block which only references to reconstructed samples in the current picture 2300 and block 2303 is not using a BV based coding mode, for example, block 2303 is of an angular mode, DC mode or planar mode, the width and / or height of block 2303 may be used to derive an HoC of the current block 2301. As block 2302 is of a same size as that of 2303, equivalently, the width and / or height of block 2302 may be used to derive an HoC of the current block 2301. In one example, coding information of the block 2303 may be used to derive an HoC of the current block 2301. In one example, the coding information of the block 2303 may include one or more of the width of block 2303, the height of block 2303, or the coding mode of block 2303. In one example, the width and / or height of block 2303 as well as the coding mode of block 2303 may be used to derive an HoC of the current block 2301.
[0434] If block 2303 is also coded using a BV based coding mode, prediction unit 1002 will check a coding mode of block 2304, which is indicated by BV (2303) of block 2303. If block 2304 is an intra coded block which only references to reconstructed samples in the current picture 2300 and block 2304 is not using a BV based coding mode, for example, block 2304 is of an angular mode, DC mode or planar mode, the width and / or height of block 2304 may be used to derive an HoC of the current block 2301. As block 2302 is of a same size as that of 2303 and 2304, equivalently, the width and / or height of block 2302 may be used to derive an HoC of the current block 2301. In one example, coding information of the block 2304 may be used to derive an HoC of the current block 2301. In one example, the coding information of the block 2304 may include one or more of the width of block 2304, the height of block 2304, or the coding mode of block 2304. In one example, the width and / or height of block 2304 as well as the coding mode of block 2304 may be used to derive an HoC of the current block 2301.
[0435] Optionally, if prediction unit 1002 cannot determine an intra coding mode that is not BV based mode after recursively searching using “BV links, ” for example using N linked BVs, prediction unit 1002 will stopped and does not use any information from block 2302 to derive HoC of the current block 2301. As an example, N is a non-negative integer. When N is equal to 0, only block 2302 is checked by prediction unit 1002. As an example, when N is equal to 1, block 2302 and block 2303 may be checked if block 2302 is coded using BV-based intra mode. As an example, when N is equal to 2, blocks 2302, 2303 and 2304 may be checked by prediction unit 1002 if blocks 2302 and 2303 is coded using BV-based intra mode.
[0436] As mentioned above, FIG. 24 illustrates a second example of an adjacent block of BV-based prediction mode, according to some embodiments of the present disclosure. For example, current block 2401 may be a block in an intra coded slice. For example, current block 2401 is a block in an intra coded picture (e.g., a picture that may be employed as an access point, such as Instantaneous Decoding Refresh picture, Clean Random Access picture, Broken Link Access picture, etc. ) .
[0437] Block 2402 is an adjacent block of current block 2401, and block 2402 is coded using a BV based coding mode, e.g. IBC or IntraTMP. BV (2402) is a BV of block 2402. Block 2403 is a block which contains a reference sample indicated by BV (2402) . In an example, (x0, y0) denotes a sample in current block 2401, and the reference sample is derived as (x0 + x (2402) , y0 + y (2402) ) , wherein BV (2402) has a horizontal component equal to x (2402) and a vertical component equal to y (2402) . For example, (x0, y0) may be a top-left sample in the current block 2401. For example, (x0, y0) may be a sample located at any corner of the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-right position (e.g., (width (2401) / 2, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-right position (e.g., (width (2401) / 2, height (2401) / 2 - 1) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-left position (e.g., (width (2401) / 2 - 1, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2 - 1) ) in the current block 2401. width (2401) and height (2401) are width and height, respectively, of the current block 2401.
[0438] If block 2403 is an intra coded block which only references to reconstructed samples in the current picture 2400 and block 2403 is not using a BV based coding mode, for example, block 2403 is of an angular mode, DC mode or planar mode, in one example, the width and / or height of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 may be used to derive a HoC of the current block 2401.
[0439] In one example, coding information of the block 2403 may be used to derive a HoC of the current block 2401. In one example, the coding information of the block 2403 may include one or more of the width of block 2403, the height of block 2403, or the coding mode of block 2403. In one example, coding information of the block 2403 and one of coding information of the block 2402, coding information of the current block 2401 may be used to derive a HoC of the current block 2401. In one example, the width and / or height of block 2403 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 as well as the coding mode of block 2403 may be used to derive a HoC of the current block 2401.
[0440] If block 2403 is also coded using a BV based coding mode, prediction unit 1002 will check a coding mode of block 2404, which is indicated by BV (2403) of block 2403. In one example, Block 2404 is a block which contains another reference sample indicated further by BV (2403) . In one example, (x0, y0) denotes a sample in current block 2401, and the another reference sample is derived as (x0 + x (2402) + x (2403) , y0 + y (2402) + x (2403) ) , wherein BV (2402) has a horizontal component equal to x (2402) and a vertical component equal to y (2402) , and BV (2403) has a horizontal component equal to x (2403) and a vertical component equal to y (2403) . For example, (x0, y0) may be a top-left sample in the current block 2401. For example, (x0, y0) may be a sample located at any corner of the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-right position (i.e. (width (2401) / 2, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-right position (i.e. (width (2401) / 2, height (2401) / 2 - 1) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle bottom-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2) ) in the current block 2401. For example, (x0, y0) may be a sample at a middle top-left position (i.e. (width (2401) / 2 - 1, height (2401) / 2 - 1) ) in the current block 2401. width (2401) and height (2401) are width and height, respectively, of the current block 2401.
[0441] If block 2404 is an intra coded block which only references to reconstructed samples in the current picture 2400 and block 2404 is not using a BV based coding mode, for example, block 2404 is of an angular mode, DC mode or planar mode, in one example, the width and / or height of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2403 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 may be used to derive a HoC of the current block 2401.
[0442] In one example, coding information of the block 2404 may be used to derive a HoC of the current block 2401. In one example, the coding information of the block 2404 may include one or more of the width of block 2404, the height of block 2404, or the coding mode of block 2404. In one example, coding information of the block 2404 and one of coding information of the block 2403, coding information of the block 2402, or coding information of the current block 2401 may be used to derive a HoC of the current block 2401. In one example, the width and / or height of block 2404 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2403 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2402 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401. In one example, width and / or height of block 2401 as well as the coding mode of block 2404 may be used to derive a HoC of the current block 2401.
[0443] Optionally, if prediction unit 1002 cannot determine an intra coding mode that is not BV based mode after recursively searching using “BV links, ” for example using N linked BVs, prediction unit 1002 will stopped and does not use any information from block 2402 to derive HoC of the current block 2401. As an example, N is a non-negative integer. When N is equal to 0, only block 2402 is checked by prediction unit 1002. As an example, when N is equal to 1, block 2402 and block 2403 may be checked if block 2402 is coded using BV-based intra mode. As an example, when N is equal to 2, blocks 2402, 2403 and 2404 may be checked by prediction unit 1002 if blocks 2402 and 2403 is coded using BV-based intra mode.
[0444] As mentioned above, FIG. 25 illustrates a flowchart of an example method 2500 of predicting a current block using BV-based technique, according to some implementations. The BV-based technique may be used in intra prediction unit 1004 and inter prediction unit 1003. For example, BV-based intra prediction technology may include intra template matching prediction (IntraTMP) , intra block copy (IBC) , spatial geometric partitioning mode (SGPM) , and so on. Additionally, the BV-based technique may be used in inter prediction technology, e.g., such as the Geometric partitioning mode (GPM) . It should be noted that the following BV-based prediction may be applied to either intra or inter prediction.
[0445] Referring to FIG. 25, at block 2502, a block vector of a current block may be determined as follows.
[0446] Methods for determining the BV of the current block include, but are not limited to, 1) performing motion search in the reconstructed area of the current picture to obtain the best BV for the current block and 2) constructing a BV Merge list for the current block by using the BV of the reconstructed block in the spatial and temporal domains. In general, BV has two precision options: integer-pel precision and sub-pel precision. The sub-pel precision may include 1 / 2-pel precision, 1 / 4-pel precision, or 1 / 16-pel precision, etc.
[0447] In some implementations, the BV of the current block may be determined based on a motion search. To perform a motion search, one type of template shown in FIGs. 5A-5C may be determined for the current block. Then, a search may be performed to identify a matching template with the smallest matching cost with the template of the current block in the reconstructed area of the current picture. An area with the same size as the current block corresponding to the matching template may be determined as the reference block. The BV of the current block may be determined based on the current block and the reference block.
[0448] In some implementations, the BV of the current block may be determined based on a BV merge list. The BV merge list may include at least one of spatial merge candidate, temporal merge candidate, history-based merge candidate, pairwise average merge candidate, and default merge candidate.
[0449] In some implementations, the BV merge list may be constructed based on spatial merge candidates. As shown in FIG. 26, the spatial merge candidates 2600 include merge candidates A1, B1, B0, A0, and B2. Spatial merge candidates A1, B1, B0, A0, and B2 are sequentially checked. If one of the spatial merge candidates A1, B1, B0, A0 and B2 is available and the available spatial merge candidate applies a BV-based prediction mode, then the BV of that spatial merge candidate may be added to the BV merge list. If the number of BV merge candidates added in the BV merge list is smaller than the allowable maximum value of the BV merge list after checking the spatial merge candidates, then one or more temporal merge candidates may be checked. One or more collocated positions in the collocated picture may be determined and temporal merge candidates corresponding to the collocated positions may be checked. If a temporal merge candidate is available and applies BV-based prediction mode, then the BV of the temporal merge candidate may be added to the BV merge list. If the number of BV merge candidates added in the BV merge list is smaller than the allowable maximum value of the BV merge list after checking the temporal merge candidates, the history-based merge candidates, pairwise average merge candidate, or default merge candidate may be checked. The BV for the current block may be determined from the BV merge list based on corresponding cost values. In some implementations, the BV merge list may be reordered, and the BV for the current block may be determined based on cost values after the reordering.
[0450] It should be noted that, at operation 2502, other methods may be used to determine the BV for the current block. The following embodiments may be applied as long as a BV is used in the process of prediction, regardless of how the BV is obtained. In some implementations, the BV may be determined having a integer-pel precision. In some implementations, BV may be determined having a sub-pel precision, including 1 / 2-pel precision, 1 / 4-pel precision, or 1 / 16-pel precision, etc. A BV with a sub-pel precision means the BV may refer to a fractional-pel position of the reconstruction area of the current picture. A BV with an integer-pel precision means the BV may refer to an integer-pel position of the reconstruction area of the current picture.
[0451] If an integer-pel precision BV is determined, to improve the precision of the BV, or considering the unity of BV precision, the integer-pel precision BV may be converted to sub-pel precision. In one example, the integer-pel precision BV may be converted to the sub-pel precision BV according to the following functions: xfrac=xint<< (FRAC_BITS-INT_BITS) , and yfrac=yint<< (FRAC_BITS-INT_BITS) .
[0452] The sub-pel precision BV is represented by (xfrac, yfrac) , and integer-pel precision BV is represented by (xint, yint) , where xfrac or xint indicates a horizontal value of the BV and yfrac or yint indicates a vertical value of the BV.
[0453] The sub-pel precision may be represented by FRAC_BITS and the integer-pel precision may be represented by INT_BITS. INT_BITS indicates a number of bits representing a BV with integer-pel precision, and FRAC_BITS indicates a number of bits representing a BV with sub-pel precision. In an example, the value of INT_BITS is equal to 2. The value of FRAC_BITS of 1 / 2-pel precision is equal to 3. The value of FRAC_BITS of 1 / 4-pel precision is equal to 4. The value of FRAC_BITS of 1 / 16-pel precision is equal to 6. Of course, the values of INT_BITS and values of FRAC_BITS of different sub-pel precisions may be determined as different preset values; or the values of INT_BITS and values of FRAC_BITS of different sub-pel precisions may be determined adaptively based on the network conditions.
[0454] At operation 2504, the block vector of the current block may be adjusted.
[0455] Different prediction modes may utilize BVs with different pel precisions. Therefore, for the BV of the current block, an adjustment of the BV may be determined based on the prediction mode. In some implementations, if the determined BV of the current block has a sub-pel precision and prediction mode utilizes a BV with integer-pel precision, a rounding operation may be performed on the determined BV of the current block to obtain an adjusted BV with integer-pel precision. The rounding operation may include, but is not limited to, rounding, rounding up and rounding down.
[0456] In one example, a rounding process may be performed on the determined BV (xfrac, yfrac) of the current block to determine an adjusted BV. An adjusted sub-pel precision BV (xint, yint) of the current block may be determined according to the following function (s) : If xfrac≥0, xint=(xfrac+nOffset-1)>>rightShift; If xfrac<0, xint=(xfrac+nOffset)>>rightShift; If yfrac≥0, yint=(yfrac+nOffset-1)>>rightShift; and If yfrac<0, yint=(yfrac+nOffset)>>rightShift, where rightShift and nOffset may be determined as follows: rightShift=FRAC_BITS-INT_BITS; and nOffset=1<< (rightShift-1) .
[0457] If BV (xfrac, yfrac) has a 1 / 2-pel precision, FRAC_BITS is equal to 3, and INT_BITS is equal to 2; thus, rightShift=1 and nOffset=1. If BV (xfrac, yfrac) has a 1 / 4-pel precision, FRAC_BITS is equal to 4, and INT_BITS is equal to 2; thus, rightShift=2 and nOffset=2. If BV (xfrac, yfrac) has a 1 / 16-pel precision, FRAC_BITS is equal to 6, and INT_BITS is equal to 2; thus, rightShift=4 and nOffset=8.
[0458] In one example, a rounding down process may be performed on the determined BV (xfrac, yfrac) to determine an adjusted BV (xint, yint) of the current block using the following functions: xint=xfrac>> (FRAC_BITS-INT_BITS) , and yint=yfrac>> (FRAC_BITS-INT_BITS) .
[0459] In one example, a rounding up process may be performed on the determined BV (xfrac, yfrac) to determine an adjusted BV (xint, yint) of the current block using the following functions: xtemp= (xfrac>> (FRAC_BITS-INT_BITS) ) << (FRAC_BITS-INT_BITS) ; ytemp= (yfrac>> (FRAC_BITS-INT_BITS) ) << (FRAC_BITS-INT_BITS) ; If xfrac=xtemp, xint=xfrac>> (FRAC_BITS-INT_BITS) ; If xfrac≠xtemp, xint=xfrac>> (FRAC_BITS-INT_BITS) +1; If yfrac=ytemp, yint=yfrac>> (FRAC_BITS-INT_BITS) ; and If yfrac≠ytemp, yint=yfrac>> (FRAC_BITS-INT_BITS) +1.
[0460] If a sub-pel precision BV is obtained for the current block, the sub-pel precision BV may be converted to an integer-pel precision BV for use in predicting the current block. In this way, a reference block may be determined in the reference area based on the integer-pel precision BV with no sample interpolation, thereby improving coding efficiency.
[0461] In some implementations, the BV of the current block may be adjusted based on refinement. In one example, if the determined BV of the current block has a sub-pel precision, the determined BV may be adjusted based on a rounding operation to obtain an integer-pel precision BV, and the integer-pel precision BV may be further refined. The rounding operation may include, but is not limited to, rounding, rounding up, or rounding down. Using a refined BV for the prediction process may improve prediction accuracy.
[0462] FIG. 27 illustrates example sub-pel and integer-pel positions 2700, according to some implementations. Sub-pel positions include 1 / 4-pel positions, 1 / 2-pel positions and 3 / 4-pel positions. The sub-pel position may be represented by FracPre. The sub-pel direction may include eight directions: left (LEFT_POS) , above left (ABOVE_LEFT_POS) , left bottom (LEFT_BOTTOM_POS) , right (RIGHT_POS) , above right (ABOVE_RIGHT_POS) , right bottom (RIGHT_BOTTOM_POS) , above (ABOVE_POS) and bottom (BOTTOM_POS) . The sub-pel direction may be represented by FracDir. The sub-pel position may also be referred to as “fractional-pel position” or “fractional sample position. ” The integer-pel position may also be referred to as “integer sample position. ”
[0463] An integer-pel precision BV may be adjusted to a sub-pel precision BV. For example, a sub-pel precision BV, BV′int (x′int, y′int) , may be determined based on an integer-pel precision BV, BVint (xint, yint) , according to the following functions: x′int=xint<< (FRAC_BITS-INT_BITS) , and y′int=yint<< (FRAC_BITS-INT_BITS) .
[0464] A sub-pel precision BV (arefined BV) , BVfrac, may be determined based on the sub-pel precision BV, BV′int, the FracPre (e.g., sub-pel position) , and FracDir (e.g., the sub-pel direction) .
[0465] A sub-pel step absDistance may be determined based on sub-pel position FracPre. The sub-pel step absDistance may also be referred to as “fractional sub step absDistance” or “fractional sample step absDistance. ”
[0466] If FracPre is 1 / 4-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>2.
[0467] If FracPre is 1 / 2-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>1.
[0468] If FracPre is 3 / 4-pel position, absDistance= (1<< (FRAC_BITS-INT_BITS) ) >>2×3.
[0469] For example, if sub-pel precision is 1 / 16-pel precision, FRAC_BITS is equal to 6 and INT_BITS is equal to 2. If FracPre is 1 / 4-pel position, sub-pel step absDistance is equal to 4. If FracPre is 1 / 2-pel position, sub-pel step absDistance is equal to 8. If FracPre is 3 / 4-pel position, sub-pel step absDistance is equal to 12.
[0470] A horizontal offset xDistance and a vertical offset yDistance may be determined based on the sub-pel step absDistance and sub-pel direction FracDir.
[0471] If the FracDir is LEFT_POS, ABOVE_LEFT_POS, or LEFT_BOTTOM_POS, xDistance=-absDistance.
[0472] If FracDir is RIGHT_POS, ABOVE_RIGHT_POS, or RIGHT_BOTTOM_POS, xDistance=absDistance.
[0473] If FracDir is ABOVE_POS, ABOVE_LEFT_POS or ABOVE_RIGHT_POS, yDistance=-absDistance.
[0474] If FracDir is BOTTOM_POS, LEFT_BOTTOM_POS or RIGHT_BOTTOM_POS, yDistance=absDistance.
[0475] A refined BV, BVfrac (xfrac, yfrac) , may be determined based on the sub-pel precision BV, BV′int (x′int, y′int) , the horizontal offset xDistance, and the vertical offset yDistance, where xfrac=x′int+xDistance and yfrac=y′int+yDistance.
[0476] As shown in FIG. 27, for example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2702) is a 1 / 4-pel position, absDistance=4, FracDir is ABOVE_LEFT_POS, xDistance=-4, and yDistance=-4, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int-4 and yfrac=y′int-4.
[0477] In another example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2703) is a 1 / 2-pel position, absDistance=8, FracDir is ABOV_RIGHT_POS, xDistance=8, and yDistance=-8, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int+8 and yfrac=y′int-8.
[0478] In a further example, assuming BV (2701) is the sub-pel precision BV, BV′int (x′int, y′int) , BV (2704) is a 3 / 4-pel position, absDistance=12, FracDir is LEFT_BOTTOM_POS, xDistance=-12, and yDistance=12, the refined BV, BVfrac (xfrac, yfrac) , may be determined as: xfrac=x′int-12 and yfrac=y′int+12.
[0479] All or a preset number of sub-pel positions may be traversed in the reconstructed area of the current frame based on template matching cost to determine a refined BV of the current block. For example, the matching cost of each sub-pel precision BV, BVfrac, may be computed, and the matching cost of the sub-pel precision BV, BV′int , corresponding to integer-pel precision BV, BVint, may be determined, and determine a BVfrac with the smallest matching cost in the reconstructed area may be used as the refined BV used for predicting the current block. The type of template may be determined based on the availability of neighboring reference samples.
[0480] Referring to FIG. 5C, when the above left, above and left reference samples are all available, the template shape may be the template shown in (a) , e.g., refTemplateType=1.
[0481] When only the left reference samples are available, the template shape may be the template shown in (b) , e.g., refTemplateType=2.
[0482] When only the above reference samples are available, the template shape may be the template shown in (c) , e.g., refTemplateType=3.
[0483] When the left and above left reference samples are available, the template shape may be the template shown in (d) , e.g., refTemplateType=4.
[0484] When the left and left bottom reference samples are available, the template shape may be the template shown in (e) , e.g., refTemplateType=5.
[0485] When the above and above right reference samples are available, the template shape may be the template shown in (f) , e.g., refTemplateType=6.
[0486] When the above and above left reference samples are available, the template shape may be the template shown in (g) , e.g., refTemplateType=7.
[0487] When the above and left reference samples are available, the template shape may be the template shown in (h) , e.g., refTemplateType=8.
[0488] A cost between two templates may be represented as an error between a template and a reference template. As an example, the cost may be a SAD, which may be calculated using equation (1) , shown above. In another example, the cost may be a SATD, which may be calculated according to equation (2) , shown above. Coding accuracy may be improved using BV refinement.
[0489] In another embodiment, predicting current block may consider BV information of a neighboring reconstructed block. Additional information corresponding to the adjacent block is comprehensively considered and may improve the prediction accuracy. In this process, a BV flip operation 2800, 2801 may be applied, as shown in FIGs. 28A and 28B.
[0490] In some implementations, BV of current block may be determined based on a neighboring reconstructed block. The determined BV of the neighboring reconstructed block may be represented by, and the flip indication, which indicates whether the BV flip operation is a horizontal flip (see FIG. 28A) or a vertical flip (see FIG. 28B) , may be represented by, rribcFlipTypenbt. The adjusted BV determined based on the flip operation may be represented by, and the flip indication may be represented by, rribcFlipTypecur.
[0491] If rribcFlipTypenbr is equal to 0, no flip may be performed. In this example, and rribcFlipTypecur=rribcFlipTypenbr=0.
[0492] If rribcFlipTypenbr is equal to 1, a horizontal flip (see FIG. 28A) may be performed. In this example, and rribcFlipTypecur=rribcFlipTypenbr=1.
[0493] A schematic diagram of the horizontal flip operation 2800 is illustrated in FIG. 28A. Assume the current block is a chroma block and xnbr is the X coordinate of the center position of the adjacent reconstructed block. If the current BV information is determined based on the luma block, then xcur is the X coordinate of the center position of the co-located luma block of the current block. If the current BV information is determined based on the chroma block, then xcur is the X coordinate of the center position of the current block. Assuming the current block is a luma block, the method may be the same.
[0494] If rribcFlipTypenbr is equal to 2, a vertical flip may be performed. In this example, and rribcFlipTypecur=rribcFlipTypenbr=2.
[0495] A schematic diagram of the vertical flip operation 2801 is illustrated in FIG. 28B. Assume the current block is a chroma block and ynbr is the Y coordinate of the center position of the adjacent reconstructed block. If the current BV information is determined based on the luma block, then ycur is the Y coordinate of the center position of the co-located luma block of the current block. If the current BV information is determined based on the chroma block, then ycur is the Y coordinate of the center position of the current block. Assuming the current block is a luma block, the method may be the same.
[0496] Referring again to FIG. 25, at operation 2506, the current block may be predicted based on the adjusted BV.
[0497] A reference block in the current frame may be determined based on the adjusted BV, and a prediction of the current block may be determined based on the reference block.
[0498] The BV used for predicting current block may be stored for use in the prediction of a subsequent block. The BV may be stored in sub-pel precision.
[0499] In one example, the BV may be stored in 1 / 2-pel precision. In one example, the BV may be stored in 1 / 4-pel precision. In one example, the BV may be stored in 1 / 16-pel precision. In one example, the BV may be stored in at least two kinds of sub-pel precisions of 1 / 2-pel precision, 1 / 4-pel precision, and / or 1 / 16-pel precision. In one example, the encoder and decoder may use a unified default sub-pel precision. In one example, the encoder may determine one or more sub-pel precision and signal a precision indication in the bitstream. The decoder may determine to store the BV with a precision based on the precision indication. The precision indication may be sequence level, picture level, slice level and block level. The precision indication may be associated with prediction mode.
[0500] It is possible to store both the integer-pel precision BV and the sub-pel precision BV of the current block. In some implementations, integer-pel precision BV may be derived based on the sub-pel precision BV. However, storing both integer-pel precision BV and sub-pel precision BV consumes more memory than only storing one of them. Thus, to reduce memory consumption sub-pel precision BV rather than integer-pel precision BV may be stored. This improves memory usage efficiency.
[0501] In some implementations, intra prediction unit 1004 may use an extrapolation filter-based intra prediction (EIP) mode to derive a prediction of the current block. FIG. 30 illustrates a flow chart of an example method of EIP 3000, in accordance with some embodiments of the present disclosure.
[0502] Referring to FIG. 30, at operation 3002, neighboring reference samples of a current block may be determined. For example, in the EIP prediction process, the samples in the current block may be predicted from the top-left position to the bottom-right position by applying an extrapolation filter to neighboring reconstructed samples or predicted samples. In this process, the input of the EIP filter may be neighboring reference samples, which include neighboring reconstructed samples, neighboring prediction samples, or padding samples.
[0503] At operation 3004, filter parameters for the current block may be determined. Filter parameters of the current block may include one or more of EIP filter length, EIP filter shape, or EIP filter coefficients. Some candidate EIP filters are provided in this disclosure.
[0504] FIG. 31 illustrates a first example of EIP filter shapes 3100, in accordance with some embodiments of the present disclosure, For example, EIP filter shapes 3100 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) in the current block, ci is the filter coefficient, and the is the neighboring reconstructed sample or neighboring predicted sample.
[0505] FIG. 32 illustrates a second example of EIP filter shapes 3200, according to some embodiments of the present disclosure. The EIP filter shapes 3200 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 32) , ci is the filter coefficient, and is the neighboring reconstructed sample or neighboring predicted sample (shown at position X in FIG. 32) , and c14 is the offset term.
[0506] In all three EIP filter shapes, their support areas cover samples above and to the left of the predicting sample, e.g., offsetXi and offsetYi are greater than or equal to zero.
[0507] FIG. 33 demonstrates a third example of EIP filter shapes 3300, according to some embodiments of the present disclosure. The EIP filter shapes 3300 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. Samples above-right and below-left of the predicting sample are considered. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 33) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample or padding sample, and c14 is the offset term.
[0508] Still referring to FIG. 33, samples at position A may be combined with a same coefficient as one input of the EIP filter, and samples at position X may use an independent coefficient. In some implementations, the average of samples at position A may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A may be determined as an input of the EIP filter.
[0509] FIG. 34 illustrates a fourth example of EIP filter shapes 3400, according to some embodiments of the present disclosure. The EIP filter shapes 3400 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 15-tap filters, and the filter length is equal to 15. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. Samples above-right and below-left of the predicting sample may be considered. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 34) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample or padding sample, and c14 is an offset term.
[0510] FIG. 35 illustrates a fifth example of EIP filter shapes 3500, according to some embodiments of the present disclosure. The EIP filter shapes 3500 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 9-tap filters, and the filter length is 9. The prediction samples of the current block are determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , is computed using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 35) , ci is the filter coefficient, the is the neighboring reconstructed sample or neighboring predicted sample, and c8 is an offset term.
[0511] Still referring to FIG. 35, samples at position A may be combined with a same coefficient as one input of the EIP filter, and samples at position X may use an independent coefficient. In some implementations, the average of samples at position A may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A may be determined as an input of the EIP filter.
[0512] FIG. 36 illustrates a sixth example of EIP filter shapes 3600, according to some embodiments of the present disclosure. The EIP filter shapes 3600 may include a square shape, a horizontal shape, and a vertical shape. These three filters are 4-tap filters, and the filter length is 4. The prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring predicted samples and an offset term. The EIP filter output, pred (x, y) , may be determined using the following formula. where pred (x, y) is the predicted value at position (x, y) of the current block (shown at position O in FIG. 36) , ci is the filter coefficient, is the neighboring reconstructed sample or neighboring predicted sample, and c4 is an offset term.
[0513] Still referring to FIG. 36, samples at positions A may be combined with a same coefficient. Samples at positions B may be combined with a same coefficient. Samples at positions C may be combined with a same coefficient. Samples at positions D may combined with a same coefficient. In some implementations, the average of samples at position A / B / C / D may be determined as an input of the EIP filter. In some implementations, the sum of samples at position A / B / C / D may be determined as an input of the EIP filter.
[0514] The EIP filter length and / or EIP filter shape may be a default configuration for a current block with a size within a predefined range or for a current block in a selected EIP prediction mode. In one example, for a selected EIP prediction mode, one of the above types of EIP filter shapes may be the default configuration. For example, for regularEIP mode and mergeEIP mode, the EIP filter shapes shown in FIG. 32 may be the default configuration. For bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be the default configuration. In one example, for a current block with a size of 8*8 and the EIP prediction mode being bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be the default configuration.
[0515] The EIP filter length and / or EIP filter shape may be determined based on one or more of the size parameter (s) of the current block, the EIP prediction mode, and / or the type of EIP models. The size parameter of the current block may include, but is not limited to, the number of samples included in the current block, the width and / or the height of the current block, the ratio of the width and height of the current block, the ratio of the height and width of the current block, and so on. The EIP prediction mode may include, but is not limited to, regularEIP, mergeEIP, or bvEIP. The type of EIP model may be a multi-model EIP or a single-model EIP.
[0516] The EIP filter length may be determined based on the size parameter of the current block, where the EIP filter length refers to the number of EIP filter coefficients included in the EIP filter. For example, the length of 15-tap EIP filter is 15, and the length of 9-tap EIP filter is 9. The size parameter of the current block may include, but is not limited to, the number of samples included in the current block, the width and / or the height of the current block, the ratio of the width and height of the current block, the ratio of the height and width of the current block, and so on.
[0517] For example, if the number of samples included in the current block is greater than a predefined threshold, a filter with a first length may be used; if the number of samples included in the current block is less than or equal to the predefined threshold, a filter with a second length may be used, where the first length is less than the second length. For a coding block with a larger number of samples, a filter with a shorter length may be used to limit the computational complexity of the EIP mode.
[0518] The EIP filter length may be determined based on the type of EIP model. If the current block utilizes a multi-model EIP, a filter with a third length may be used; if the current block utilizes a single-model EIP, a filter with a fourth length may be used, where the third length is less than the fourth length. For a coding block using a multi-model EIP, a filter with shorter length may be used to limit the computational complexity of the EIP mode.
[0519] The EIP filter length may be determined based on the type of EIP prediction mode, where EIP prediction mode may include, but is not limited to, regularEIP mode, mergeEIP mode or bvEIP mode. In an example, if the EIP mode is bvEIP mode, a filter with a fifth length may be used; if the current block uses regularEIP mode or mergeEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, if the EIP prediction mode is regularEIP mode, a filter with a fifth length may be used; if the current block uses bvEIP mode or mergeEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, if the EIP prediction mode is mergeEIP mode, a filter with a fifth length may be used; if the current block uses bvEIP mode or regularEIP mode, a filter with a sixth length may be used, where the fifth length is less than the sixth length. In an example, two EIP modes selected from regularEIP mode, mergeEIP mode, and bvEIP mode may use a filter with a fifth length and the remaining EIP mode may use a filter with a sixth length, where the fifth length is less than the sixth length. The length of the EIP filter may be adaptively determined based on the type of EIP prediction mode to limit the computational complexity of EIP mode.
[0520] It should be noted that the first length, the third length, and the fifth length may be the same value, such as 9; and the second length, the fourth length, and the sixth length may be the same value, such as 15. In some implementations, the first length, the third length, and the fifth length may be different values; and the second length, the fourth length and the sixth length may be the different values.
[0521] The length of EIP filter may include multiple candidate lengths including 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more. The length of EIP filter may be adaptively determined based on one or more size parameters of the current block, the EIP prediction mode, and / or the type of EIP model (s) to achieve a balance between prediction accuracy and coding complexity.
[0522] The type of EIP filter shape (s) may be determined based on one or more the size parameters of the current block, the type of EIP models, the type of EIP prediction mode, and so on.
[0523] In some implementations, the type of EIP filter shape (s) may be determined based on the type of EIP prediction mode. For example, if the EIP prediction mode of the current block is bvEIP mode, the EIP filter shapes shown in FIG. 33 or FIG. 34 may be used. Otherwise, the EIP filter shapes shown in FIG. 32 may be used.
[0524] The EIP prediction mode includes one of regularEIP mode, mergeEIP mode, or bvEIP mode.
[0525] The regularEIP mode refers to deriving EIP filter coefficients based on neighboring reconstructed samples in a reconstructed area. The mergeEIP mode refers to deriving EIP filter parameters by inheriting EIP filter parameters from the reconstructed blocks. The bvEIP mode refers to using a block vector to determine the reference area for calculating the EIP filter parameters instead of directly using the adjacent spatial reference area.
[0526] FIG. 37 illustrates example templates 3700 of a reconstructed area used to derive EIP filter parameters, according to some embodiments of the present disclosure. In FIG. 37, three example templates of the reconstructed area, which are used to derive EIP filter parameters, are shown.
[0527] Referring to the example template shown on the left-hand side of FIG. 37, the reconstructed area includes the left and the above neighboring reconstructed area, which may be called the L-shaped reconstructed area. The reconstructed area of this shape has the largest area and may cover the surrounding adjacent above, above-left, above-right, left, and left-bottom samples.
[0528] Referring to the example template shown in the middle of FIG. 37, the reconstructed area includes the above neighboring reconstructed area, which may be called the T-shaped reconstructed area. The reconstructed area of this shape may cover the above, above-left, and above-right samples.
[0529] Referring to the example template shown on the right-hand side of FIG. 37, the reconstructed area includes the left neighboring reconstructed area, which may be called the L-shaped reconstructed area. The reconstructed area of this shape may cover the above-left, left, and left-bottom samples.
[0530] As shown in FIG. 37, taking the square filter as an example, fWidth represents the width of the square filter, and fHeight represents the height of the square filter.
[0531] In the example template shown on the left-hand side of FIG. 37, leftSize represents the width of the left template, that is, the number of columns on the left side of the current block; aboveSize represents the height of the upper template, that is, the number of rows on the upper side of the current block; the size of the template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0532] In the example template shown in the middle of FIG. 37 (e.g., the above template) , leftSize represents the number of columns on the left side of the current block and aboveSize represents the number of rows on the above side of the current block; the size of the upper template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0533] In the example template shown on the right-hand side of FIG. 37 (e.g., the left template) , leftSize represents the number of columns on the left side of the current block and aboveSize represents the number of rows on the above side of the current block; the size of the left template area may be determined according to leftSize and aboveSize, and the square filter scans in the template area to obtain the corresponding input samples and output samples.
[0534] RegularEIP modes allows the use of three reconstructed areas and three filter shapes; when constructing the regularEIP candidate list, the maximum allowable length is 9. The length of the regularEIP candidate list and the EIP candidate information constructed by the coding units of different shapes or different width and height sizes may be different. For example, a 32×32 coding unit allows the use of all combinations of EIP candidate information; that is, the length of the regularEIP candidate list is 9; while a 4×4 coding unit allows the use of an L-shaped reconstructed area to combine 3 filter shapes, which means the length of the regular EIP candidate list of the coding unit is 3.
[0535] The selected filter moves in the selected reconstructed area either horizontally or vertically with a one-pixel step to construct the auto-correlation matrix and the cross-correlation vector. The calculation of coefficients from the auto-correlation matrix and the cross-correlation vector is the same as that in CCCM.
[0536] To decrease coding complexity, a subsampling process may be optionally performed on the reconstructed area used to derive EIP filter parameters.
[0537] In some implementations, the reconstructed area may be downsampled in the horizontal direction or vertical direction with a step size to obtain a downsampled reconstructed area. In some implementations, the step size may be a predefined value, such as 2. In some implementations, the step size may be determined based on size parameters of the reconstructed area, or size parameters of the current block. The size parameters may include, but are not limited to, the number of samples, the width and / or height, the ratio of width and height, or the ratio of height and width. In some implementations, when obtaining the input of the EIP model training, the reference area samples may be subsampled with a step size, such as 2.
[0538] In some implementations, the reconstructed area may be downsampled by a low-pass filter.
[0539] FIG. 38 illustrates an example of a bvEIP mode 3800, according to some embodiments of the present disclosure. Referring to FIG. 38, a BV may be determined for the current block, and the reference block pointed to by the BV may be used as the reconstructed area for deriving EIP filter parameters for the current block.
[0540] Non-limiting examples for determining the BV of the current block may include the following: 1) performing a motion search in the reconstructed area of the current picture to obtain the BV for the current block, and 2) constructing a BV merge list for the current block by inheriting the BV of the reconstructed block from the spatial and temporal domains.
[0541] In an embodiment, the BV of the current block can be determined based on a motion search. First, one of the templates shown in FIG. 5 may be determined for the current block. Then, a search may be performed to determine the matching template with the smallest template cost corresponding to the template of the current block in the reconstructed area of the current picture.
[0542] FIG. 39 illustrates an example of a reconstructed area 3900, according to some embodiments of the present disclosure.
[0543] Referring to FIG. 39, H represents the height of the current block, and W represents the width of the current block. A template search in the search area from R1 to R6 may be performed, the distortion cost value of the template area may be obtained, and a candidate list of block vector information may be updated and maintained. For example, when encoding the current block, a coarse search in the R1 to R6 area may be performed with a step of 3 or 4 samples using the template shown in FIG. 39; the distortion cost value of each search template area may be determined using, e.g., SAD, MR-SAD, SATD, etc. Based on the principle of ascending cost value, a coarse search block vector list with a length of 33 or other natural value length may be maintained and updated such that the smallest value is located at the top of the list and the largest cost value is located at the bottom of the list.
[0544] Taking each block vector information in the coarse search block vector information list obtained by the coarse search process as the starting point, a fine search with a step of 1 sample point in the search area (R1~R6) corresponding to each block vector information may be performed, and a fine search block vector list with a length of 19 or other natural number may be updated and maintained. The fine search process is similar to the coarse search process in that evaluation indicators such as SAD, MR-SAD, SATD, or MR-SATD may be used to determine the distortion cost, and the fine search list may be maintained with the same sorting principle.
[0545] A matching template with the smallest template cost may be determined, and the area with the same size as the current block corresponding to the matching template as the reference block may be determined. The BV of the current block may be determined based on the current block and the reference block.
[0546] In some implementations, the BV of the current block may be determined by constructing a BV merge list. The BV merge list may include at least one of a spatial merge candidate, a temporal merge candidate, a history-based merge candidate, a pairwise average merge candidate, and / or a default merge candidate. Additional details for constructing the BV merge list are provided above in connection with FIG. 26.
[0547] A reference block as the reconstructed area may be used to derive EIP parameters of current block. To decrease the coding complexity, a subsampling process may be optionally performed on the reconstructed area used for deriving EIP filter parameters.
[0548] In some implementations, the reconstructed area may be downsampled in the horizontal direction or vertical direction with a step size to obtain a downsampled reconstructed area. In some implementations, the step size may be a predefined value, such as 2. In some implementations, the step size may be determined based on size parameters of the reconstructed area or size parameters of the current block. The size parameters may include, but are not limited, to the number of samples, the width and / or height, the ratio of width and height, or the ratio of height and width. In some implementations, while obtaining the input of the EIP model training, the reference area samples may be subsampled with a step size, such as 2.
[0549] In some implementations the reconstructed area is downsampled by a low-pass filter.
[0550] For a current block coded in the EIP mode, an EIP merge flag may be parsed to determine whether the EIP filter is inherited from previous blocks coded in EIP mode. An EIP merge flag may be signaled to indicate whether the EIP filter is inherited from previous blocks coded in EIP mode. When the EIP merge flag is true, an EIP merge list may be constructed based on one or more of spatial adjacent candidates, spatial non-adjacent candidates, temporal candidates, and / or historical candidates. The position and inclusion order of these candidates may be the same as those used in a cross-component prediction (CCP) merge candidate list. An EIP merge index may be parsed to determine an EIP merge candidate from the EIP merge list. The filter parameters, which include the filter shape and the filter coefficients of the selected candidate, may be inherited to code the current block.
[0551] When the EIP merge flag is false, a bvEIP flag may be parsed to determine whether bvEIP prediction mode is used; if the bvEIP flag is false, regularEIP prediction mode is used, and the relevant syntax element may be parsed to determine which one of the three types of reconstructed area and which one of the three filter shapes are used for the current block.
[0552] Referring again to FIG. 30, at operation 3006, prediction samples of the current block may be determined based on the neighboring reference samples and the filter parameters.
[0553] After determining a EIP filter for the current block, prediction samples of the current block may be determined based on the neighboring reconstructed samples or neighboring prediction samples and the EIP filter.
[0554] After generating the prediction samples of the current block using the EIP filter, an intra prediction mode may be derived by applying the DIMD process to the prediction samples. For instance, a horizontal gradient and a vertical gradient may be calculated for each predicted sample to build a histogram of gradients (HoG) . The intra prediction mode corresponding to the largest histogram count may be used to determine the low-frequency non-separable transform (LFNST) , non-separable primary transform (NSPT) , or multiple transform selection (MTS) transform set.
[0555] In one example, intra prediction unit 1004 may use an NN-based intra prediction mode to derive a prediction of the current block. FIG. 29 illustrates an example of NN-based intra prediction 2900, according to some embodiments.
[0556] Block 2901 (block Y) is a current block having w×h samples. The samples 2902 adjacent to block 2901 are reference samples. Let “X” be an input of the NN-based intra prediction process. In one example, “X” may be one or more samples among the samples 2902. In one example, “X” may be derived by filtering one or more samples among the samples 2902. Block 2903 is an output of the NN-based intra prediction process. In one example, block 2903 is an intra prediction of block 2901.
[0557] In one example, inter prediction unit 1003 may also derive motion parameters and / or prediction samples using an NN-based method or process.
[0558] In some implementations, prediction unit 1002 may conduct a subblock-based prediction process. A subblock-based merge candidate list may be constructed that includes one or more of subblock-based temporal motion vector prediction (SbTMVP) candidate (s) , subblock-based spatial motion vector prediction (SbSMVP) candidate (s) , and / or affine merge candidate (s) .
[0559] FIG. 40 illustrates an example of subblock template generation of subblock-based temporal motion vector prediction (SbTMVP) 4000, according to some embodiments of the present disclosure.
[0560] The temporal motion vector prediction (TMVP) for advanced motion vector prediction (AMVP) and merge mode may be derived by obtaining the motion information from the center or the bottom-right of the collocated block in a signaled collocated picture. For the subblock-based temporal motion vector prediction (SbTMVP) mode, the motion information from the left neighboring position may be used as a motion shift, which may be used to obtain TMVPs at sub-CU level. In some implementations, two collocated pictures, which are the two reference pictures with the least picture-order-count (POC) distance relative to the to-be-coded picture, may be utilized. The motion shift to locate TMVP may be adaptively determined from multiple locations according to template costs. Two motion shift candidate lists may be constructed respectively for the two collocated frames. The motion shifts with the minimum template matching cost may be used to derive SbTMVP or TMVP candidates. In a non-limiting example, at most 4 SbTMVP candidates may be included in the subblock merge candidate list. The SbTMVP candidate with the least template matching cost derived from the first collocated frame may be placed in the first entry without reordering, while other SbTMVP candidates may be sorted together with other candidates. Further, the prediction direction of each subblock template may be determined based on the center subblock. As illustrated in FIG. 40, if the center subblock is uni-predicted, then all the subblock templates are uni-predicted, and vice versa. If the motion vector of the corresponding adjacent subblock at the determined reference list is not available for a subblock template, zero MV may be used for that subblock template. In some implementations, if the motion vector of a corresponding adjacent subblock at the determined reference list is not available for a subblock template, a derived MV may be used for that subblock template, where the derived MV may be derived based on other available subblocks in the determined reference list. Derivation of an MV may be performed based on a weighted average or other combination technique (s) .
[0561] FIG. 41 illustrates an example of subblock-based spatial motion vector prediction (SbSMVP) candidate types 4100, according to some embodiments of the present disclosure.
[0562] Referring to FIG. 41, the subblock motion field of the current block may be inherited based on the motion of the spatial neighboring blocks. FIG. 41 shows an example of derivation of SbSMVP candidates from spatial neighboring blocks. Examples of different SbSMVP candidate types, in which MVs of subblocks are inherited in a directional way, are shown in FIG. 41.
[0563] As shown in diagram (a) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along a horizontal direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the horizontal direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the horizontal direction.
[0564] As shown in diagram (b) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along a vertical direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the vertical direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the vertical direction.
[0565] As shown in diagrams (c) - (e) of FIG. 41, SbSMVP candidates may be derived from spatial neighboring subblocks along an angular direction. In some implementations, the SbSMVP candidates may be obtained by copying motion information of the spatial neighboring subblocks along the angular direction. In some implementations, the SbSMVP candidates may be obtained by scaling or refining motion information of the spatial neighboring subblocks along the angular direction. It should be noted that angular directions may include, but are not limited to, the example angular directions shown in (c) - (e) of FIG. 41. Additional angular directions are available than those shown, and an angular direction may be adaptively determined.
[0566] In some implementations, if a spatial neighboring block in which a spatial neighboring subblock is located is coded based on an intra prediction mode or is not coded based on an inter prediction mode, or if one of spatial neighboring subblocks has no available MV, an MV of the spatial neighboring subblock may be determined to be a default value, such as zero.
[0567] In some implementations, if a spatial neighboring block in which a spatial neighboring subblock is located is coded based on an intra prediction mode or is not coded based on an inter prediction mode, or if one of spatial neighboring subblocks has no available MV, an MV of the spatial neighboring subblock may be determined based on one or more MVs of other spatial neighboring subblocks. In some implementations, the MV of the spatial neighboring subblock may be determined by copying one of the adjacent subblocks of the spatial neighboring subblock. In some implementations, the MV of the spatial neighboring subblock may be determined by a weighted average of two or more adjacent subblocks of the spatial neighboring subblock.
[0568] Using diagram (a) of FIG. 41 as a non-limiting example, if the spatial neighboring block in which the spatial neighboring subblock tagged with MV6 is coded based on intra prediction, MV6 of the spatial neighboring subblock is not available. MV6 may be determined to be a zero MV. Otherwise, MV6 may be determined by copying one of MV5, MV7, MV0, or MV8, or MV6 may be determined to be an average MV of two or more of MV5, MV7, MV0, and MV8.
[0569] In some implementations, if an SbSMVP candidate is selected, the motion information, such as motion vector, reference index, and prediction direction of the corresponding neighboring subblock may be copied to the current subblock along a predefined direction, as depicted by the arrows in FIG. 41. In some implementations, if an SbSMVP candidate is selected, the motion information, such as motion vector, reference index, and prediction direction of the current block may be determined based on the motion information of the corresponding neighboring subblock along a predefined direction.
[0570] Up to five SbSMVP candidates may be added to the subblock merge candidate list. SbSMVP candidates may be added between SbTMVP candidates and affine merge candidates.
[0571] All or part of the subblock merge candidates included in the subblock merge candidate list may be reordered. Twenty candidates may be sorted, in some implementations. A candidate may be determined based on a merge index signaled to indicate the merge candidate.
[0572] Prediction unit 1002 may pass the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction also may be called as “intra prediction direction” ) to transform unit 1006. In one embodiment, transform unit 1006 may use such information to determine transform kernel or a set of transform kernels in the primary transform and / or secondary transform.
[0573] When parameter from parsing unit 1001 indicates that region transform is applied to decode the current block, transform unit 1006 may derive residual sample of the current block as following. When transform unit 1006 uses region transform to code the current block, transform unit 1006 determines parameters indicating a position and size of a region in the current block, and performs transform to obtain the reconstructed samples in the region. The reconstructed samples are residual samples. The transform unit 1006 sets a value of a sample in the remaining region of the current block to be equal to 0, wherein the said sample is a residual sample.
[0574] Transform unit 1006 in decoder 1000 may determine, according to one or more parameters from parsing unit 1001, that a region or a sub-block in a current block and performs transform on the samples in this region or sub-block. FIG. 19 demonstrates an example of region transform of a current block. The current block 1900 may be a coding block, a coding unit, a transform unit or a sub-block. Region 1901 is the said region in the current block 1900. Transform unit 1006 will perform a transform on the coefficients obtained by parsing unit 1001 in region 1901 to determine a sample in region 1901, and set a value of a sample or coefficient in the remaining region 1902 in the current block 1900 to be equal to 0. The sample in region 1901 may be residual sample of the current block. Transform unit 1006 may obtain a region parameter from parsing unit 1001 indicating the region 1901 including at least one of the following: [Parameter 1] : position of region 1901 and / or [Parameter 2] : size of region 1901.
[0575] As an option, transform unit 1006 may determine, according to the parameter from parsing unit 1001, that more than one region in a current block may be transformed. FIG. 19 also shows an example in which two regions are in a current block. Transform unit 1006 will perform transform on coefficients in regions 1911 and 1913 in a current block 1910, and set a value of a sample or coefficient in the remaining region 1912 in the current block 1910 to be equal to 0. The sample in regions 1911 and 1912 may be residual sample of the current block. Transform unit 1006 may obtain one or more region parameters from parsing unit 1001 indicating region 1911 and 1913 including one or both of [Parameter 1] and [Parameter 2] .
[0576] In the following descriptions, current block 1900 may be taken as an example. The implementation with multiple regions containing coefficients (e.g., current block 1911) is carried out using similar method to indicate the regions.
[0577] As mentioned above, FIGs. 20A and 20B illustrates examples of region transform. Transform unit 1006 performs transform on a coefficient in a gray region in a current block and sets a value of a sample or coefficient in the remaining region in a current block to be equal to 0, wherein the sample may be a residual sample after prediction. The “arrays” of each gray region demonstrates the transform directions and transform kernel of each direction.
[0578] In one example, the gray region in FIGs. 20A and 20B is at a pre-defined position with pre-defined size. For example in FIGs. 20A and 20B, [Parameter 2] may be one or more parameters indicating a split type of a current block (e.g., quad, triple, horizontal or vertical) , and [Parameter 1] may be one or more parameters indicating which one of the regions, according to the split type of a current block as indicated by [Parameter 2] , is the region on a sample of which transform unit 1006 performs a transform. Transform unit 1006 may derive a width and a height (i.e. a size) of a gray region according to the abovementioned parameters obtained from parsing unit 1001. For example, given that a size (e.g., width x height) of the current block is 4Wx4H. The size (e.g., represented width x height of a region in the current block) and position (e.g., represented by a location of top-left sample of a region in the current block) of a gray region in FIGs. 20A and 20B are shown above in Tables 5A-5D.
[0579] As mentioned above, FIG. 21 illustrates examples of region transform. Transform unit 1006 performs transform on a coefficient in a gray region in a current block and sets a value of a sample or coefficient in the remaining region in a current block to be equal to 0, wherein the sample may be a residual sample. The “arrays” of each gray region demonstrates the transform directions and transform kernel of each direction.
[0580] A size of a gray region 2101, 2111 or 2121 may be represented as gW x gH, wherein gW is a width of the gray region, and gH is a height of the gray region, and a position of a gray region may be represented by a location of a top-left sample in the gray region in the current block, e.g., (dx, dy) .
[0581] Transform unit 1006 may determines values of dx and dy of [Parameter 1] , and gW and gH of [Parameter 2] .
[0582] Optionally, transform unit 1006 may first determine a split type of a gray region according to parameter from parsing unit 1001. For example, a split type of gray region 2101 is “arbitrary type, ” which indicates that gW and gH are smaller than a with and a height of the current block 2100, respectively. In this case, transform unit 1006 determines values of dx and dy of [Parameter 1] , and gW and gH of [Parameter 2] for gray region 2101. For example, a split type of gray region 2111 is “vertical type, ” transform unit 207 determines values of dx of [Parameter 1] , and gW of [Parameter 2] for gray region 2111, as dy may be inferred to be 0 and gH may be inferred to be equal to the height of the current block 2110. For example, a split type of gray region 2121 is “horizontal type, ” transform unit 207 determines values of dy of [Parameter 1] , and gH of [Parameter 2] for gray region 2121, as dx may be inferred to be 0 and gW may be inferred to be equal to the width of the current block 2120.
[0583] In an example, dx and dy are represented in a precision of integral sample in a bitstream.
[0584] In another example, dx and dy are represented in a precision of multiple samples. For example, dx is represented as dx>>shift in a bitstream, wherein shift is an non-negative integer, and “dx>>shift” is arithmetic right shift of a two's complement integer representation of dx by shift binary digits. When obtaining a corresponding parameter (e.g. denoted as “Offset” here) from parsing unit 1001, transform unit 1006 sets a value of dx to be equal to Offset<<shift, wherein “Offset<<shift” is arithmetic left shift of a two's complement integer representation of Offset by shift binary digits. “shift” may be a fixed value, for example 1, 2, 3, 4, …, Log2 (MaxCuSize) - 1, wherein MaxCuSize is the maximum value of a width or height of a coding unit and Log2 (MaxCuSize ) is a base-2 logarithm of MaxCuSize. Examples of a representation of dy in a bitstream and a derivation of dy value according to parameter from parsing unit 1001 is the same as that of dx.
[0585] Adder 1007 performs addition operation with its inputs of prediction block from prediction unit 1002 and reconstructed residual from 1006 to get reconstructed block of the current decoding block. The reconstructed block is also sent to prediction unit 1002 to be used as reference for other blocks coded in intra prediction mode.
[0586] In one embodiment, after the CUs in a picture or a sub-picture have been reconstructed, filtering unit 1008 performs in-loop filtering on the reconstructed picture or sub-picture. Filtering unit 1008 contains one or more filters, for example, deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF) , luma mapping with chroma scaling (LMCS) filter and neural network based filters. Alternatively, when filtering unit 1008 determines that the reconstructed block is not used as reference for decoding other blocks, filtering unit 1008 performs in-loop filtering on one or more target pixels in the reconstructed block.
[0587] In one embodiment, filtering unit 1008 would process filtering on the reconstructed samples of one or more color components of the current block (e.g., a CU) . The decoder 1000 stores the filtered reconstructed samples of one or more color components of the current block in a picture buffer for a picture in which the current block locates. Thus, the prediction unit 1002 can use the filtered samples of the current block in decoding the succeeding block of the current block in decoding order. For example, the prediction unit 1002 can use the filtered samples of the current block to derive a prediction of succeeding block of the current block in decoding order. For example, the prediction unit 1002 as well as other units in decoder 1000, would include the filtered samples of the current block in a template and derive of a prediction, reordering candidate modes or parameters, and / or decoding parameters using template matching approach. Since the filtering unit 1008 suppresses reconstruction distortion of the current block introduced by the lossy source coding of encoder 200, when the filtered sample of the current block is used to decode the succeeding block, the prediction efficiency of the succeeding block has been improved, and thus the coding efficiency may be greatly improved.
[0588] In one embodiment, the filtering unit 1008 uses one or more fixed 1D or 2D filters to process the reconstruct sample of the current block. For example, the 1D filter may be a symmetry filter. For example, the 1D filter may be an asymmetry filter. For example, the 2D filter may be a symmetry filter. For example, the 2D filter may be an asymmetry filter. For example, the 2D filter may be a separable filter. For example, the 2D filter may be a non-separable filter.
[0589] In one embodiment, the filtering unit 1008 uses one or more adaptive 1D or 2D filters to process the reconstruct sample of the current block. For example, the 1D filter may be a symmetry filter. For example, the 1D filter may be an asymmetry filter. For example, the 2D filter may be a symmetry filter. For example, the 2D filter may be an asymmetry filter. For example, the 2D filter may be a separable filter. For example, the 2D filter may be a non-separable filter.
[0590] In one embodiment, the filtering unit 1008 uses one or more neural-network based filters to process the reconstruct sample of the current block.
[0591] In one embodiment, the filtering unit 1008 can use one or more filters of the spatial and / or temporal neighboring blocks of the current block. In one example, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks before filtering which is invoked after reconstructing a picture where the neighboring block locates. In one example, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks after reconstructing a picture where the neighboring block locates. One example is that filtering unit 1008 may use the adaptive loop filter (ALF) which is used to filter a temporal neighboring block of the current block. In one example, the filtering unit 1008 may select one or more existing filters which are available before filtering the current block. One example is that the filters with parameters are obtained, by parsing unit 1001, from block layer (e.g. coding tree unit or coding unit) or a layer higher than a block layer of the current block (e.g. video parameter set, sequence parameter set, picture parameter set, adaption parameter set, picture header and / or slice header) of the bitstream.
[0592] In one embodiment, filtering unit 212 may obtain an indication parameter from parsing unit 1001, which is to indicate whether the reconstructed sample in the current block is needed to be filtered or not. For example, the indication parameter may be a 1 bit flag. For example, the indication parameter may be a variable with a number of values indicating not only whether the reconstruct sample is needed to be filtered but also which filter is used. When the variable is equal to 0, the reconstructed sample of the current block will not be filtered; otherwise, the reconstructed sample of the current block is filtered with a filter with an index equal to the value of this variable.
[0593] In one embodiment, filtering unit 1008 may also obtain indication parameter from parsing unit 1001, which is to indicate which color component may be filtered. Filtering unit 1008 can choose to filter one or more of the luma and two chroma components.
[0594] Output of filtering unit 1008 is a decoded picture or sub-picture, which is forwarded to DPB (decoded picture buffer) 1009. DPB 1009 outputs decoded pictures according to timing and controlling information. Pictures stored in DPB 1009 may also be employed as reference for performing inter or intra prediction by prediction unit 1002.
[0595] Decoder 1000 could be a computing device with a processor and a storage medium recording a decoding program. When the processor reads and executes the decoding program, the decoder 1000 reads an input video bitstream and generates corresponding decoded video.
[0596] Decoder 1000 could be a computing device with one or more chips. The units, implemented as integrated circuits, on the chip are of similar functionalities with similar connections as well as data exchangings as the corresponding ones in FIG. 10.
[0597] FIG. 11 illustrates an example source device 1100. Acquisition unit 1101 acquires a video signal and forwards the video signal to encoder 1102. Acquisition unit 1101 may be a device containing one or more cameras (including depth cameras) . Acquisition unit 1101 may be a device that partially or completely decodes a bitstream to get a video. Acquisition unit 1101 may also contain one or more elements to capture audio signal. An embodiment of encoder 1102 is the encoder 200 that codes the video signal from acquisition unit 1101 as its input video and generates a video bitstream. Encoder 1102 may also contains one or more audio encoder to code the audio signal to generate an audio bitstream. Storage / sending unit 1103 receives the video bitstream from encoder 1102. Storage / sending unit 1103 may also receive the audio bitstream from encoder 1102 and encapsulate the video bitstream together with the audio bitstream to form a media file (e.g. ISO based media file format) or transport stream. Optionally, storage / sending unit 1103 writes the media file or transport stream in a storage unit. e.g. hard disc, DVD disc, cloud, portable memory devices. Optionally, storage / sending unit 1103 sends the bitstream to a transport network, for example, Internet, wireline networks, cellular networks, wireless local area networks, etc.
[0598] FIG. 12 illustrates an example destination device 1200. Receiving unit 1201 receives the media file or transport stream from networks or reads the media file or transport stream from a storage device. Receiving unit 1201 separates the video bitstream and the audio bitstream from the media file or transport stream. Receiving unit 1201 can also generate a new video bitstream by extracting the video bitstream. Receiving unit 1201 may also generate a new audio bitstream by extracting the audio bitstream. Decoder 1202 includes one or more video decoders, e.g. the decoder 1000. Decoder 1202 may also contains one or more audio decoders. Decoder 1202 decodes the video bitstream and the audio bitstream from receiving unit 1201 to get a decoded video and one or more decoded audio corresponding to one or multiple channels. Rendering unit 1203 performs operations on the reconstructed video to make it suitable for displaying. Such operations may include one or more of the following operations to improve perceptual quality: denoising, synthesis, conversion of color space, upsampling, downsampling, etc. Rendering unit 1203 may also performs operations on the decoded audio to improve the perceptual quality of the audio signal for displaying.
[0599] FIG. 13 illustrates a communication system 1300. Source device 1301 is a source device 1100. Output of the storage / sending unit 1103 is processed by storage medium / transport networks 1302 for storage or transport the bitstream. Destination Device 1303 is a destination device 1200. Receiving unit 1201 gets the bitstream from storage medium / transport networks 1302. Receiving unit 1201 may extract a new video bitstream from the media file or transport stream. Receiving unit 1201 may also extract a new audio bitstream from the media file or transport stream.
[0600] In an example of a session negotiation between Source Device 1301 and Destination Device 1303, Source Device 1301 may send its NN-ability information to Destination Device 1303. For example, Source Device 1301 may indicate that it may provide a bitstream that may be decoded using an NN-based decoding process, and it may also provide a bitstream that may be decoded without using an NN-based decoding process. When receiving the NN-ability information of Source Device 1301, Destination Device 1303 will check its own NN-ability information and feedback to Source Device 1301. One example is that Destination Device 1303 informs Source Device 1301 that it supports an NN-based decoding process and a non-NN-based decoding process, and the Source Device 1301 may determine to send Destination Device 1303 either a bitstream that may be decoded using an NN-based decoding process or another bitstream that may be decoded without using an NN-based decoding process. One example is that Destination Device 1303 informs Source Device 1301 that it only supports NN-based decoding process, and the Source Device 1301 may only determine to send Destination Device 1303 a bitstream that may be decoded using an NN-based decoding process. One example is that Destination Device 1303 informs Source Device 1301 that it only supports a non-NN-based decoding process, and the Source Device 1301 can determine to send Destination Device 1303 a bitstream that may be decoded without using an NN-based decoding process. In the above examples, if the NN-ability information further includes a quantization-error bounds for performing NN-based process, Source device 1301 and Destination Device 1303 may also exchange their quantization-error bounds for NN-based process parameters, and if the error-bounded parameters can secure reproducibility or interoperability between Source device 1301 and Destination Device 1303, the Source Device 1301 may determine to send Destination Device 1303 a bitstream that may be decoded using an NN-based decoding process.
[0601] In an example of a session negotiation between Source Device 1301 and Destination Device 1303, Destination Device 1303 may send its NN-ability information to Source Device 1301 to request data or a bitstream from Source Device 1301. For example, Destination Device 1303 may indicate that it can decode a bitstream using an NN-based decoding process, and it can also decode a bitstream independent of an NN-based decoding process. When receiving the NN-ability information of Destination Device 1303, Source Device 1301 may check NN-ability information for decoding a bitstream and feedback to Destination Device 1303. One example is that Destination Device 1303 informs Source Device 1301 that it supports NN-based decoding process and non-NN-based decoding process, and the Source Device 1301 may determine to send Destination Device 1303 either a bitstream that may be decoded using an NN-based decoding process, or another bitstream that may be decoded without using an NN-based decoding process. One example is that Destination Device 1303 informs Source Device 1301 that it only supports NN-based decoding process, and the Source Device 1301 will only determine to send Destination Device 1303 a bitstream that may be decoded using an NN-based decoding process. One example is that Destination Device 1303 informs Source Device 1301 that it only supports non-NN-based decoding...
Claims
1.A method of decoding, comprising:determining, by a processor, a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determining, by the processor, a subblock-based merge candidate list for the current block based on the SbSMVP candidate; anddecoding, by the processor, the current block based on the subblock-based merge candidate list.2.The method of claim 1, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determining, by the processor, a motion vector (MV) of the SbSMVP candidate based on a default value.3.The method of claim 2, wherein the default value is zero.4.The method of claim 1, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determining, by the processor, the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.5.The method of claim 4, wherein the default value is zero.6.The method of claim 1, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determining, by the processor, a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.7.The method of claim 6, wherein the determining, by the processor, the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate comprises:determining, by the processor, the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.8.The method of claim 1, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determining, by the processor, the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.9.The method of claim 8, wherein the determining, by the processor, the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate comprises:determining, by the processor, the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.10.The method of claim 1, further comprising:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.11.The method of claim 10, further comprising:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determining, by the processor, the subblock-based merge candidate list for the current block based on one or more default merge candidates.12.A decoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; anddecode the current block based on the subblock-based merge candidate list.13.The decoder of claim 12, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on a default value.14.The decoder of claim 13, wherein the default value is zero.15.The decoder of claim 12, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.16.The decoder of claim 15, wherein the default value is zero.17.The decoder of claim 12, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.18.The decoder of claim 17, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the memory storing instructions, which when executed by the processor, cause the processor to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.19.The decoder of claim 12, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.20.The decoder of claim 19, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the memory storing instructions, which when executed by the processor, cause the processor to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.21.The decoder of claim 12, wherein the memory storing instructions, which when executed by the processor, cause the processor to:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.22.The decoder of claim 21, wherein the memory storing instructions, which when executed by the processor, cause the processor to:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determine the subblock-based merge candidate list for the current block based on one or more default merge candidates.23.An apparatus for decoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; anddecode the current block based on the subblock-based merge candidate list.24.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a decoder, cause the processor of the decoder to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; anddecode the current block based on the subblock-based merge candidate list.25.The non-transitory computer-readable medium of claim 24, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on a default value.26.The non-transitory computer-readable medium of claim 25, wherein the default value is zero.27.The non-transitory computer-readable medium of claim 24, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.28.The non-transitory computer-readable medium of claim 27, wherein the default value is zero.29.The non-transitory computer-readable medium of claim 24, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.30.The non-transitory computer-readable medium of claim 29, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.31.The non-transitory computer-readable medium of claim 24, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.32.The non-transitory computer-readable medium of claim 31, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.33.The non-transitory computer-readable medium of claim 24, wherein the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.34.The non-transitory computer-readable medium of claim 33, wherein the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determine the subblock-based merge candidate list for the current block based on one or more default merge candidates.35.A method of encoding, comprising:determining, by a processor, a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determining, by the processor, a subblock-based merge candidate list for the current block based on the SbSMVP candidate; andencoding, by the processor, the current block based on the subblock-based merge candidate list.36.The method of claim 35, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determining, by the processor, a motion vector (MV) of the SbSMVP candidate based on a default value.37.The method of claim 36, wherein the default value is zero.38.The method of claim 35, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determining, by the processor, the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.39.The method of claim 38, wherein the default value is zero.40.The method of claim 35, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determining, by the processor, a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.41.The method of claim 40, wherein the determining, by the processor, the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate comprises:determining, by the processor, the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.42.The method of claim 35, wherein the determining, by the processor, the SbSMVP candidate based on the spatial neighboring block of the current block comprises:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determining, by the processor, the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.43.The method of claim 42, wherein the determining, by the processor, the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate comprises:determining, by the processor, the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.44.The method of claim 35, further comprising:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.45.The method of claim 44, further comprising:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determining, by the processor, the subblock-based merge candidate list for the current block based on one or more default merge candidates.46.A encoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; andencode the current block based on the subblock-based merge candidate list.47.The encoder of claim 46, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on a default value.48.The encoder of claim 47, wherein the default value is zero.49.The encoder of claim 46, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.50.The encoder of claim 49, wherein the default value is zero.51.The encoder of claim 46, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.52.The encoder of claim 51, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the memory storing instructions, which when executed by the processor, cause the processor to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.53.The encoder of claim 46, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the memory storing instructions, which when executed by the processor, cause the processor to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.54.The encoder of claim 53, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the memory storing instructions, which when executed by the processor, cause the processor to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.55.The encoder of claim 46, wherein the memory storing instructions, which when executed by the processor, cause the processor to:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.56.The encoder of claim 55, wherein the memory storing instructions, which when executed by the processor, cause the processor to:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determine the subblock-based merge candidate list for the current block based on one or more default merge candidates.57.An apparatus for encoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; andencode the current block based on the subblock-based merge candidate list.58.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a encoder, cause the processor of the encoder to:determine a subblock-based spatial motion vector prediction (SbSMVP) candidate based on a spatial neighboring block of a current block;determine a subblock-based merge candidate list for the current block based on the SbSMVP candidate; andencode the current block based on the subblock-based merge candidate list.59.The non-transitory computer-readable medium of claim 58, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on a default value.60.The non-transitory computer-readable medium of claim 59, wherein the default value is zero.61.The non-transitory computer-readable medium of claim 58, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on a default value, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.62.The non-transitory computer-readable medium of claim 61, wherein the default value is zero.63.The non-transitory computer-readable medium of claim 58, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:in response to the spatial neighboring block corresponding to the SbSMVP candidate being coded based on an intra prediction mode or being coded independent of an inter prediction mode, determine a motion vector (MV) of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.64.The non-transitory computer-readable medium of claim 63, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.65.The non-transitory computer-readable medium of claim 58, wherein, to determine the SbSMVP candidate based on the spatial neighboring block of the current block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:in response to a motion vector (MV) of a spatial neighboring subblock corresponding to the SbSMVP candidate being unavailable, determine the MV of the SbSMVP candidate based on at least one other MV corresponding to at least one adjacent subblock of a spatial neighboring subblock corresponding to the SbSMVP candidate, wherein the spatial neighboring subblock is a subblock of the spatial neighboring block of the current block.66.The non-transitory computer-readable medium of claim 65, wherein, to determine the MV of the SbSMVP candidate based on the at least one other MV corresponding to the at least one adjacent subblock of the spatial neighboring subblock corresponding to the SbSMVP candidate, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:determine the MV of the SbSMVP candidate based on a weighted average of a plurality of other motion vectors (MVs) corresponding to a plurality of adjacent subblocks of the spatial neighboring subblock corresponding to the SbSMVP candidate.67.The non-transitory computer-readable medium of claim 58, wherein the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:determining, by the processor, the subblock-based merge candidate list for the current block based on one or more of a subblock-based temporal motion vector prediction (SbTMVP) candidate or an affine merge candidate.68.The non-transitory computer-readable medium of claim 67, wherein the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:in response to a number of candidates in the subblock-based merge candidate list being less than a threshold number after the SbSMVP candidate and the one or more of the SbTMVP candidate or the affine merge candidate are included in the subblock-based merge candidate list, determine the subblock-based merge candidate list for the current block based on one or more default merge candidates.69.A method of transmitting a bitstream, comprising:executing the method of encoding of one or more of claims 35-45 to generate a bitstream; andtransmitting the bitstream.70.A non-transitory computer-readable storage medium, having a computer program and a bitstream stored thereon, wherein the computer program, when executed by a processor, enables the processor to perform the method of encoding of one or more of claims 35-45 to generate the bitstream.