System and method of histogram of occurrence-based intra coding (OBIC) mode enhancement
Adaptive control of template-based prediction modes in video coding improves efficiency and quality by reducing memory usage, addressing the inefficiencies of current techniques.
Patent Information
- Application Number
- PCT/CN2025/108443
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-14
- Publication Date
- 2026-01-22
AI Technical Summary
Current video coding techniques using template-based prediction modes consume excessive memory and buffer space, negatively impacting coding efficiency.
A technique that signals instructions or controlling parameters in the bitstream to enable or disable template-based prediction modes adaptively, based on video characteristics, improving coding quality without increasing computational complexity.
Enhances coding efficiency by optimizing template-based prediction modes, leading to improved perceptual quality and flexible complexity management for video codecs.
Smart Images

Figure CN2025108443_22012026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD OF HISTOGRAM OF OCCURRENCE-BASED INTRA CODING (OBIC) MODE ENHANCEMENTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 672,726, filed July 18, 2024, entitled “ENCODING METHOD AND DECODING METHOD, ENCODER AND DECODER, AND STORAGE MEDIUM, ” which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Embodiments of the present disclosure relate to video coding.
[0003] Digital video has become mainstream and is being used in a wide range of applications including digital television, video telephony, and teleconferencing. These digital video applications are feasible because of the advances in computing and communication technologies as well as efficient video coding techniques. Various video coding techniques may be used to compress video data, such that coding on the video data can be performed using one or more video coding standards. Exemplary video coding standards may include, but not limited to, versatile video coding (H. 266 / VVC) , high-efficiency video coding (H. 265 / HEVC) , advanced video coding (H. 264 / AVC) , moving picture expert group (MPEG) coding, enhanced video coding model (ECM) , to name a few.SUMMARY
[0004] According to one aspect of the present disclosure, a method of decoding is provided. The method may include, in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The method may include, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0005] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0006] According to another aspect of the present disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0007] According to a further aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a decoder is provided. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0008] According to one aspect of the present disclosure, a method of encoding is provided. The method may include, in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The method may include, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0009] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0010] According to another aspect of the present disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0011] According to a further aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for an encoder is provided. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0012] According to yet another aspect of the present disclosure, a method of transmitting a bitstream is provided. The method may include generating, by a processor, the bitstream according to one or more of the operations described herein. The method may include transmitting, by the processor, the bitstream.
[0013] According to still another aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated using one or more operations described herein.
[0014] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding thereof. Additional embodiments are described in the Detailed Description, and further description is provided there.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the pertinent art to make and use the present disclosure.
[0016] FIG. 1A illustrates a block diagram of an exemplary encoding system, according to some embodiments of the present disclosure.
[0017] FIG. 1B illustrates a block diagram of an exemplary decoding system, according to some embodiments of the present disclosure.
[0018] FIG. 2 illustrates a block diagram of an exemplary encoder, according to some embodiments of the present disclosure.
[0019] FIG. 3A illustrates an exemplary technique of quadtree splitting of a coding unit, according to some embodiments of the present disclosure.
[0020] FIG. 3B illustrates an exemplary technique of binary splitting and ternary splitting of a coding unit, according to some embodiments of the present disclosure.
[0021] FIG. 3C illustrates an exemplary technique of splitting of a coding unit into various split types, according to some embodiments of the present disclosure.
[0022] FIG. 4 illustrates an exemplary technique of inter prediction based on template matching, according to some embodiments of the present disclosure.
[0023] FIG. 5A illustrates first exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0024] FIG. 5B illustrates second exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0025] FIG. 5C illustrates third exemplary templates used in template matching, according to some embodiments of the present disclosure.
[0026] FIG. 6 illustrates an exemplary technique of intra prediction based on template matching, according to some embodiments of the present disclosure.
[0027] FIG. 7 illustrates a block diagram of an exemplary decoder, according to some embodiments of the present disclosure.
[0028] FIG. 8 illustrates a block diagram of an exemplary source device, according to some embodiments of the present disclosure.
[0029] FIG. 9 illustrates a block diagram of an exemplary receiving device, according to some embodiments of the present disclosure.
[0030] FIG. 10 illustrates a block diagram of a first exemplary communication system, according to some embodiments of the present disclosure.
[0031] FIG. 11 illustrates a block diagram of a second exemplary communication system, according to some embodiments of the present disclosure.
[0032] FIG. 12 illustrates a block diagram of a third exemplary communication system, according to some embodiments of the present disclosure.
[0033] FIG. 13 illustrates an example visualization of a current block and reference samples, according to some embodiments of the present disclosure.
[0034] FIG. 14A illustrates a first example Sobel filter, according to some embodiments of the present disclosure.
[0035] FIG. 14B illustrates a second example Sobel filter, according to some embodiments of the present disclosure.
[0036] FIG. 14C illustrates a first example Edge filter, according to some embodiments of the present disclosure.
[0037] FIG. 14D illustrates a second example Edge filter, according to some embodiments of the present disclosure.
[0038] FIG. 15 illustrates a diagram of non-adjacent spatial neighboring candidates for occurrence-based intra coding (OBIC) mode, according to some embodiments of the present disclosure.
[0039] FIG. 16 illustrates an example of the histogram-of-occurrences (HoC) of intra prediction mode, according to some embodiments of the present disclosure.
[0040] FIG. 17 illustrates a flow chart of an exemplary method of decoding, according to some embodiments of the present disclosure.
[0041] FIG. 18 illustrates a flow chart of an exemplary method of encoding, according to some embodiments of the present disclosure.
[0042] Embodiments of the present disclosure will be described with reference to the accompanying drawings.DETAILED DESCRIPTION
[0043] Although some configurations and arrangements are discussed, it should be understood that this is done for illustrative purposes only. A person skilled in the pertinent art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. It will be apparent to a person skilled in the pertinent art that the present disclosure can also be employed in a variety of other applications.
[0044] It is noted that references in the specification to “one embodiment, ” “an embodiment, ” “an example embodiment, ” “some embodiments, ” “certain embodiments, ” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of a person skilled in the pertinent art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0045] In general, terminology may be understood at least in part from usage in context. For example, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a, ” “an, ” or “the, ” again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0046] Various aspects of point cloud coding systems will now be described with reference to various apparatus and methods. These apparatus and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as “elements” ) . These elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends upon the particular application and design constraints imposed on the overall system. The techniques described herein may be used for various point cloud coding applications. As described herein, point cloud coding includes both encoding and decoding a point cloud.
[0047] In August 2020, the telecommunication branch of the International Telecommunication Unit (ITU-T) finalized a standardization project namely H. 266 / VVC (Versatile Video Coding) and published the first version of the ITU-T H. 266 standard. Then, the standardization committee started exploration work aiming to achieve a performance superior to the latest H. 266 / VVC standard in coding high quality video with one or more features of high resolution, high frame rate, high bit depth, high dynamic range, wide color gamut, and omnidirectional video. JVET (Joint Video Expert Group of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29) is in charge of this exploration work. Various intra and inter prediction modes have been verified to achieve high compression efficiency in coding high quality video and thus adopted in a software platform for this exploration work.
[0048] Prediction modes based on template matching are used in this software platform. A video encoder or decoder may employ such modes to derive an inter prediction or intra prediction block of a current block.
[0049] To perform template matching, the codec derives a reference template or matching template based on a cost function that measures the difference between the template for the current block and the candidate reference template or the candidate matching template. Then, the reference template or matching template is used to construct the prediction of the current block. Currently, the template-based prediction mode derivation uses an undesirable amount of memory or buffer space, which has a negative impact on coding efficiency.
[0050] To overcome these and other challenges, the present disclosure provides a technique that signals instruction or controlling parameters in the bitstream to enable or disable template-based prediction modes (e.g., in a collectively way) , which enables the encoder to adaptively apply template-based prediction modes based on characteristics of input video or picture. Using this technique, template-based prediction modes may be selected based on the characteristics of video content during encoding. Therefore, by adaptively configuring the instruction or controlling parameters, coding quality is improved, without an undue increase in computational complexity. This improves the overall coding efficiency of the video codec, especially measured in perceptual quality. Moreover, the proposed technique that signals instructions or controls parameters in the bitstream also provides a flexible mechanism to control the complexity of both encoder and decoder, for example, by enabling or disabling template-based prediction modes. Additional details of the exemplary technique are provided below in connection with FIGs. 1A-16.
[0051] FIG. 1A illustrates a block diagram of an exemplary encoding system 100, according to some embodiments of the present disclosure. FIG. 1B illustrates a block diagram of an exemplary decoding system 150, according to some embodiments of the present disclosure. Each system 100 or 150 may be applied or integrated into various systems and apparatuses capable of data processing, such as computers and wireless communication devices. For example, system 100 or 150 may be the entirety or part of a mobile phone, a desktop computer, a laptop computer, a tablet, a vehicle computer, a gaming console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an argument reality (AR) device, or any other suitable electronic devices having data processing capability. As shown in FIGs. 1A and 1B, system 100 or 150 may include a processor 102, a memory 104, and an interface 106. These components are shown as connected one to another by a bus, but other connection types are also permitted. It is understood that system 100 or 150 may include any other suitable components for performing functions described here.
[0052] Processor 102 may include microprocessors, such as graphic processing unit (GPU) , image signal processor (ISP) , central processing unit (CPU) , digital signal processor (DSP) , tensor processing unit (TPU) , vision processing unit (VPU) , neural processing unit (NPU) , synergistic processing unit (SPU) , or physics processing unit (PPU) , microcontroller units (MCUs) , application-specific integrated circuits (ASICs) , field-programmable gate arrays (FPGAs) , programmable logic devices (PLDs) , state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout the present disclosure. Although only one processor is shown in FIGs. 1A and 1B, it is understood that multiple processors can be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Software can include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for instructing hardware are also permitted under the broad category of software.
[0053] Memory 104 can broadly include both memory (a.k.a, primary / system memory) and storage (a.k.a. secondary memory) . For example, memory 104 may include random-access memory (RAM) , read-only memory (ROM) , static RAM (SRAM) , dynamic RAM (DRAM) , ferro-electric RAM (FRAM) , electrically erasable programmable ROM (EEPROM) , compact disc read-only memory (CD-ROM) or other optical disk storage, hard disk drive (HDD) , such as magnetic disk storage or other magnetic storage devices, Flash drive, solid-state drive (SSD) , or any other medium that can be used to carry or store desired program code in the form of instructions that can be accessed and executed by processor 102. Broadly, memory 104 may be embodied by any computer-readable medium, such as a non-transitory computer-readable medium. Although only one memory is shown in FIGs. 1A and 1B, it is understood that multiple memories can be included.
[0054] Interface 106 can broadly include a data interface and a communication interface that is configured to receive and transmit a signal in a process of receiving and transmitting information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although only one memory is shown in FIGs. 1A and 1B, it is understood that multiple interfaces can be included.
[0055] Processor 102, memory 104, and interface 106 may be implemented in various forms in system 100 or 150 for performing point cloud coding functions. In some embodiments, processor 102, memory 104, and interface 106 of system 100 or 150 are implemented (e.g., integrated) on one or more system-on-chips (SoCs) . In one example, processor 102, memory 104, and interface 106 may be integrated on an application processor (AP) SoC that handles application processing in an operating system (OS) environment, including running point cloud encoding and decoding applications. In another example, processor 102, memory 104, and interface 106 may be integrated on a specialized processor chip for point cloud coding, such as a GPU or ISP chip dedicated to graphic processing in a real-time operating system (RTOS) .
[0056] As shown in FIG. 1A, in encoding system 100, processor 102 may include one or more modules, such as an encoder 101. Although FIG. 1A shows that encoder 101 is within one processor 102, it is understood that encoder 101 may include one or more sub-modules that can be implemented on different processors located closely or remotely with each other. Encoder 101 (and any corresponding sub-modules or sub-units) can be hardware units (e.g., portions of an integrated circuit) of processor 102 designed for use with other components or software units implemented by processor 102 through executing at least part of a program, i.e., instructions. The instructions of the program may be stored on a computer-readable medium, such as memory 104, and when executed by processor 102, it may perform a process having one or more functions related to point cloud encoding, such as voxelization, transformation, quantization, arithmetic encoding, etc., as described below in detail.
[0057] Similarly, as shown in FIG. 1B, in decoding system 150, processor 102 may include one or more modules, such as a decoder 120. Although FIG. 2 shows that decoder 120 is within one processor 102, it is understood that decoder 120 may include one or more sub-modules that can be implemented on different processors located closely or remotely with each other. Decoder 120 (and any corresponding sub-modules or sub-units) can be hardware units (e.g., portions of an integrated circuit) of processor 102 designed for use with other components or software units implemented by processor 102 through executing at least part of a program, i.e., instructions. The instructions of the program may be stored on a computer-readable medium, such as memory 104, and when executed by processor 102, it may perform a process having one or more functions related to point cloud decoding, such as arithmetic decoding, dequantization, inverse transformation, reconstruction, synthesis, as described below in detail.
[0058] FIG. 2 illustrates a block diagram of an exemplary encoder 200, according to some embodiments of the present disclosure. As shown, the input to the encoder 200 may be a video including a sequence of pictures or a still picture, while the output of the encoder 200 may be a bitstream representing a compressed version of the input video. As also shown, encoder 200 may include, e.g., a partition unit 201, a prediction unit 202, a block partition unit 203, an inter prediction unit 204, an intra prediction unit 205, a first adder 206, a transform unit 207, a quantization unit 208, an inverse quantization unit 209, an inverse transform unit 210, a second adder 211, a filtering unit 212, and a decoded picture buffer (DPB) 213. The various operations of encoder 200 will now be described.
[0059] For example, referring to FIG. 2, partition unit 201 divides a picture in an input video into one or more coding tree units (CTUs) . Partition unit 201 divides the picture into tiles, and optionally may further divide a tile into one or more bricks. A tile or a brick may contain one or more integral and / or partial CTUs. Partition unit 201 forms one or more slices, where a slice may contain one or more tiles in a raster order of tiles in the picture, or one or more tiles covering a rectangular region in the picture. Partition unit 201 may also form one or more sub-pictures, which may contain one or more slices, tiles, or bricks.
[0060] During the encoding process, partition unit 201 passes CTUs to prediction unit 202. Generally, prediction unit 202 is composed of block partition unit 203, inter prediction unit 204, and intra prediction unit 205. Block partition unit 203 further divides an input CTU into smaller coding units (CUs) using various split or partition types, such as quadtree split, binary split, and ternary split iteratively. Examples of quadtree split, binary split, and ternary split of a CU or a coding block are described below in connection with FIGs. 3A, 3B, and 3C.
[0061] FIG. 3A illustrates an exemplary technique of quadtree splitting 300 of a coding unit, according to some embodiments of the present disclosure. As shown in FIG. 3A, quadtree split is applied to CU or coding block 301. Blocks 3010, 3011, 3012 and 3013, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. The processing order of the four CUs obtained by partitioning coding block 301 is 3010, 3011, 3012, and 3013.
[0062] FIG. 3B illustrates an exemplary technique of binary splitting and ternary splitting 325 of a coding unit, according to some embodiments of the present disclosure. As shown in FIG. 3B, various examples of binary splitting and / or ternary splitting is applied to a CU are depicted. For instance, CU 302 is partitioned using vertical binary split. Blocks 3020 and 3021, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 303 is partitioned using horizontal binary split. Blocks 3030 and 3031, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 304 is partitioned using vertical ternary split. Blocks 3040, 3041, and 3042, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc. CU 305 is partitioned using horizontal ternary split. Blocks 3050, 3051, and 3052, as CU, can also be further partitioned iteratively using various split or partition types, such as quadtree split, binary split, ternary split, etc.
[0063] FIG. 3C illustrates an exemplary technique of splitting 350 of a coding unit into various split types, according to some embodiments of the present disclosure. As shown in FIG. 3C, an example of partitioning or splitting a CU 306 iteratively using various split or partition types, e.g., such as quadtree split, binary split, ternary split, etc. is depicted.
[0064] Referring again to FIG. 2, prediction unit 202 may derive inter prediction block of a CU using inter prediction unit 204, and may derive intra prediction block of a CU using intra prediction unit 205. In an example, prediction unit 202 may use a rate-distortion mode decision process to determine a prediction mode of a CU.
[0065] Generally, inter prediction unit 204 performs motion estimation to derive motion parameters of a CU. The motion parameters include motion vector (MV) and reference index (refIdx) . MV indicates a relative location of a matching block in a reference picture indicated by refIdx in a specific reference list (e.g., list 0 and list 1) . Generally, list 0 mainly includes reference pictures that is ahead of the current picture in an output order or a displaying order, while list 1 mainly includes reference pictures that is behind the current picture in an output order or a displaying order. Inter prediction unit 204 may derive the MV of the CU using the samples in the CU, and find the block in the reference with least cost according to a rate-distortion motion estimation method. Inter prediction unit 204 may derive the MV of the CU using spatial reference samples of the CU.
[0066] FIG. 4 illustrates an exemplary technique of inter prediction 400 based on template matching, according to some embodiments of the present disclosure.
[0067] As shown in FIG. 4, CU 403 is the current CU in the current picture 401. A reference picture of the current CU 403 is a reconstructed or decoded reference picture 402. Template 404 includes the above neighboring samples and / or left samples. Instead of deriving a matching block of the current CU 403 in a search range in a reference picture 402, inter prediction unit 204 derives a matching template 405 of template 404 in reference picture 402. An initial MV is derived based on the displacement between template 404 and matching template 405. In an example, the block 406 in a same size of the current block 403 can be used an inter prediction block of the current CU 403. FIGs. 5A-5C shows examples of templates.
[0068] FIG. 5A illustrates first exemplary templates 500 used in template matching, according to some embodiments of the present disclosure. Referring to FIG. 5A, examples of candidate templates using neighboring samples of the current CU (Curr CU) are depicted. As an example, a template of the current CU can be one of AboveLeft, Above, AboveRight, Left and BottomLeft. As another example, a template of the current CU can be a region of a combination of two or more of AboveLeft, Above, AboveRight, Left and BottomLeft. Take template 404 for example. Template 404 in FIG. 4 is a combination of Above and Left templates in FIG. 5A.
[0069] FIG. 5B illustrates second exemplary templates 525 used in template matching, according to some embodiments of the present disclosure. As shown in FIG. 5B, examples of candidate templates using samples that are not directly neighboring the current CU (Curr CU) are depicted. As an example, a template of the current CU can be one of AboveLeft, Above-Left, Above, AboveRight, Left-Above, Left and BottomLeft. As another example, a template of the current CU can be a region of a combination of two or more of AboveLeft, Above-Left, Above, AboveRight, Left-Above, Left and BottomLeft. As another example, a template of the current CU can be a combination of two or more of the example candidate templates in FIGs. 5A and 5B.
[0070] FIG. 5C illustrates third exemplary templates 550 used in template matching, according to some embodiments of the present disclosure. As shown in FIG. 5C examples of templates from the candidate templates in FIG. 5A are shown. In some other implementations, templates using non-neighboring samples may be derived in a similar way in Figure 5C using examples of candidate templates in FIG. 5B.
[0071] Referring again to FIG. 2, inter prediction unit 204 may refine an existing MV based on a displacement between a template of a current block and a matching block. Additionally and / or alternatively, inter prediction unit 204 may refine an existing MV based on a displacement between two template blocks. In some embodiments, inter prediction unit 204 may also use the matching template to derive certain compensation to the current block.
[0072] Generally, inter prediction unit 204 may determine a template shape. Then, inter prediction unit 204 may calculate a cost between various reference templates and the template. Finally, inter prediction unit 204 may determine a reference template that results in an optimal cost function as a matching template of the template.
[0073] A cost between two templates can be represented as an error between a template and a reference template. As an example, the cost can be a Sum of Absolute Differences (SAD) calculated according to equation (1) shown below. where Ti, m and Tm are samples in two templates, respectively; and M is a number of samples in a template.
[0074] As an example, the cost can be a Sum of Absolute Transformed Differences (SATD) calculated according to equation (2) shown below. where X represents a matrix of a difference between two template samples, M is the size of the matrix, and H is a normalized MxM Hadamard matrix.
[0075] As an example, the cost can be a Mean Reduced Sum of Absolute Differences (MR-SAD) calculated according to equation (3) shown below. where Ti, m and Tm are samples in two templates, respectively; M is the number of samples in a template; Avgi is the average value of a first template containing Ti, m; and Avg is the average of a second template containing Tm.
[0076] In some implementations, other cost measures that may be used by inter prediction unit 204 may include one or more of, e.g., Mean Squared Error (MSE) , Sum of Squared Differences (SSD) , Mean Absolute Difference (MAD) , Mean Squared Difference (MSD) , Normalized Cross-Correlation (NCC) , Structural Similarity Index Measure (SSIM) , or Multi-Scale Structural Similarity Index Measure (MS-SSIM) , just to name a few.
[0077] Still referring to FIG. 2, besides directly using the matching block 406 as a prediction of the current CU 403, inter prediction unit 204 may be enabled with a number of prediction modes that utilizes template matching to derive a matching template. Table 1 below shows example prediction modes, which may be carried out by the inter prediction unit 204, with default cost functions of the modes. As one example, the intra block copying (IBC) mode can be carried out by inter prediction unit 204 by setting available reconstructed area of the current picture as “reference picture” to derive a matching block of the current block. Table 1: Example Prediction Modes for Inter Prediction
[0078] Inter prediction unit 204 may determine a cost different from the default cost for one or more prediction modes in Table 1. That is, inter prediction unit 204 may choose a cost function from several candidate cost functions. In some implementations, inter prediction unit 204 may determine whether default costs are used for one or more of the inter prediction modes. In some implementations, inter prediction unit 204 can determine that for an inter prediction mode that utilizes template matching, a first cost function (e.g., SAD) can be used for coding all pictures in a video sequence, and / or pictures within a certain period in a video sequence, and / or a picture, and / or a tile, and / or a slice, and / or a CTU, and / or a CU. In an embodiment, inter prediction unit 204 may use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 1, inter prediction unit 204 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, inter prediction unit 204 can use SAD or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In an embodiment, inter prediction unit 204 can use a single cost function for all prediction modes.
[0079] Generally, intra prediction unit 205 may derive an intra prediction block of a CU using various intra prediction modes including, e.g., direct copying (DC) mode, planar mode, angular prediction mode, Matrix-based Intra Prediction (MIP) mode, cross-component linear model intra prediction (CCLM) mode, intra block copying (IBC) mode, intra template matching prediction (IntraTMP) mode, etc. In an example, rate-distortion optimized motion estimation can be invoked by intra prediction unit 205 to derive the intra prediction mode for a current block; and according to the intra prediction mode, intra prediction unit 205 determines the intra prediction block for the current block.
[0080] FIG. 6 illustrates an exemplary technique of intra prediction 600 based on template matching, according to some embodiments of the present disclosure.
[0081] Referring to FIG. 6, an example of IntraTMP mode is shown. Picture 601 is a current picture, and block 603 is current block or current CU. Region 602 in picture 601 is an already reconstructed area. Current block 603 is within region 608, which is an area to be reconstructed or encoded. Intra prediction unit 205 first determines a template of the current block 603, where a template can be one of the template types as shown in one or more of FIGs. 5A-5C. In the non-limiting example shown in FIG. 6, intra prediction unit 205 uses a template that is a combination of Above and Left neighboring samples of the current block 603. Intra prediction unit 205 may then search in a search range within region 602 to find a matching template that results in an optimal cost as a matching template of the template.
[0082] The cost between two templates can be represented as an error between a template and a reference template. As an example, the cost can be the SAD calculated according to equation (1) , shown above. As another example, the cost can be an SATD calculated according to equation (2) , shown above. As a further example, the cost can be a MR-SAD calculated according to equation (3) shown above. In some implementations, other cost measures that can be used by intra prediction unit 205 may include one or more of, e.g., MSE, SSD, MAD, MSD, NCC, SSIM, or MS-SSIM, just to name a few.
[0083] By way of example and not limitation, intra prediction unit 205 may use the SAD function to determine a matching reference template 605 for the template 604. The reference block 607, which may be in a same size as that of the current block 603, is used by the intra prediction unit 205 to derive a prediction of the current block 603. A block vector (BV) 606 represents a displacement between the template 604 and its matching reference template 605, and / or represents a displacement between the current block 603 and the reference block 607.
[0084] In some non-limiting examples, referring to FIGs. 2 and 6, intra prediction unit 205 can use the reference block 607 as the prediction of the current block 603. In some non-limiting examples, intra prediction unit 205 can use a result of filtering the reference block 607 to be the prediction of the current block 603. In some non-limiting examples, intra prediction unit 205 can use a result of the reference block 607 adjusted with weighting factor to be the prediction of the current block 603.
[0085] Still referring to FIGs. 2 and 6, intra prediction unit 205 can derive more than one reference block, in some implementations. For example, an optimal matching block and a sub-optimal matching block can be derived by the intra prediction unit 205, and two reference blocks will be available. Intra prediction unit 205 may perform a fusion operation on the multiple reference blocks derived in template matching process to derive a prediction of the current block 603. An example fusion operation is a weighted combination of the multiple reference blocks, where the weights may be determined based on a matching error between the two reference templates and the template 604.
[0086] Still referring to FIGs. 2 and 6, in one example, besides deriving the prediction of the current block, intra prediction unit 205 may also derive coding parameters for the current block based on the reference template 605. In one embodiment, intra prediction unit 205 derives model parameters for performing cross-component prediction on the current block 603 based on one or more reference templates 605.
[0087] Table 2 shows example prediction modes, which may be carried out by the intra prediction unit 205, with default cost functions of the modes. In some implementations, the intra block copying (IBC) mode can be carried out by intra prediction unit 205 as the prediction block is derived only based on the samples in the current picture and no temporal reference picture or inter-layer reference picture is used. Table 2: Example Prediction Modes for Intra Prediction
[0088] Referring to FIG. 2, intra prediction unit 205 can determine a cost different from the default cost for one or more prediction modes in Table 2. That is, intra prediction unit 205 can choose a cost function from several candidate cost functions. In some implementations, intra prediction unit 205 can determine whether default costs are used for one or more of the intra prediction modes. In some implementations, intra prediction unit 205 can determine that for an intra prediction mode that utilizes template matching, a first cost function (e.g., SAD) can be used for coding all pictures in a video sequence, and / or pictures within a certain period in a video sequence, and / or a picture, and / or a tile, and / or a slice, and / or a CTU, and / or a CU. In some implementations, intra prediction unit 205 may use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 2, intra prediction unit 205 may use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, intra prediction unit 205 can use SAD, or any one of the abovementioned cost functions for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In some implementations, intra prediction unit 205 can use a single cost function for all prediction modes.
[0089] Prediction unit 203 may also derive intra prediction mode or angular prediction direction of the current CU based on one or more reference samples, as described below in connection with FIG. 13.
[0090] FIG. 13 illustrates an example visualization 1300 of a current block 1301 and reference samples 1302, 1303, according to some embodiments of the present disclosure.
[0091] Referring to FIG. 13, the reference sample 1302, 1303, which are marked as black dot outside the current block 1301, may be one or more samples in a template ( “L-shape” template consisting of the black dots in FIG. 13) as shown in FIG. 5. For example, a gradient of reference sample 1302, 1303 can be derived by applying one or more filters to process one or more samples in a template as shown in FIG. 5. Prediction unit 203 first derives a gradient of a reference sample using an operator. Generally, the operator can be used to detect an edge or a gradient in a picture. The operator can be a 2-dimentional (2D) M x N filter, where M and N are positive integers, and M may be equal to or different from N.
[0092] One example of the operator is a Sobel filter. A first example of a 3x3 Sobel filter 1400 is depicted in FIG. 14A, and a second example of a 3x3 Sobel filter 1425 is depicted in FIG. 14B.
[0093] Another example of the operator is an Edge filter. A first example of an Edge filter 1450 is depicted in FIG. 14C, and a second example of an Edge filter 1475 is depicted in FIG. 14D.
[0094] In some implementations, prediction unit 203 can choose different operators according to a width and / or a height of the current block. For example, prediction unit 203 uses smaller operator for smaller block, and uses larger operator for larger block. One example would be that prediction unit 203 uses the abovementioned Edge filter when a size (e.g., width x height) of the current block is 4x4, 4x8 or 8x4, and uses the abovementioned Sobel filter for other sizes of the current block.
[0095] Prediction unit 203 derives a Histogram of Gradients (HoG) by analyzing one or more reference samples 1302, 1303 marked in black dots in FIG. 13. The template in FIG. 13 may include three reference sample lines above and three reference sample columns on the left of the current block 1301. The HoG is derived by accumulating the magnitudes of one or more gradients at one or more given directions, for one, a part of, or all of the reference samples 1302, 1303 as shown in FIG. 13. One or more directions indicated by the gradients with the highest or higher cumulative magnitudes may be identified as the angular prediction direction or intra prediction mode of the current block 1301.
[0096] As an example, when prediction unit 203 uses one reference sample (e.g., any black dot above and / or to the left of the current block 1301) as shown in FIG. 13 to derive a direction, the prediction unit 203 can determine the direction as the one indicated by a gradient derived at this reference sample.
[0097] As an example, when prediction unit 203 uses a part of reference samples (e.g., one or more of the black dots above and / or to the left of the current block 1301) as shown in FIG. 13 to derive directions, prediction unit 203 can choose a preset number of samples from the reference samples. For example, prediction unit 203 choose J reference samples above the current block and K reference samples left to the current block, wherein J and K are integers greater than or equal to 0. For example, both J and K are equal to 2, J equal to 4 and K equal to 8, or J plus K equal to 4. Prediction unit 203 derives the gradients at the selected reference samples using the operator and derives the HoG. One or more directions indicated by the gradients with highest or higher cumulative magnitudes may be identified as the angular prediction direction or intra prediction mode of the current block 1301.
[0098] As an example, prediction unit 203 can adaptively determine one or more reference samples used for deriving a HoG. Prediction unit 203 uses a preset scanning order of the reference samples. When scanning a reference sample, prediction unit 203 derives a gradient at this reference sample, and updates the accumulation of the magnitudes according to this gradient at one or more given directions in the HoG. When prediction unit 203 determines that the total cumulative amplitude is greater or equal than a given threshold, prediction unit 203 will stop scanning the remaining reference sample and deriving new gradient. The resulted HoG at the termination of scanning by prediction unit 203 is determined as the HoG to derive intra prediction mode or angular prediction direction.
[0099] Examples of the abovementioned preset scanning order of the reference samples can be the following.
[0100] For example, a scanning order (A) may be scanning the left column of reference samples as shown in FIG. 13 from bottom to top; a scanning order (B) may be scanning the left column of reference samples as shown in FIG. 13 from top to bottom; a scanning order (C) may be scanning the above line of reference samples as shown in FIG. 13 from left to right; a scanning order (D) can be scanning the above line of reference samples as shown in FIG. 13 from left to right.
[0101] An example of the preset scanning order can be one or more of “first order (A) then order (C) , ” “first order (A) then order (D) , ” “first order (B) then order (C) , ” “first order (B) then order (D) , ” “first order (C) then order (A) , ” “first order (D) then order (A) , ” “first order (C) then order (B) , ” and / or “first order (D) then order (B) . ”
[0102] In some implementations, the preset scanning order may be performed in an interleaving manner. In one example, one or multiple reference samples may be scanned from the left column and then one or multiple second reference samples from the above line and then one or multiple third reference sample from left column. In another example, one or multiple reference samples may be scanned from the above line and then one or multiple second reference samples from the left column and then one or multiple third reference sample from above line. Additionally, as an example, the scanning order of samples in the left column can be one or more of order (A) and (B) , and the scanning order of samples in the left column can be one or more of order (C) and (D) .
[0103] Prediction unit 203 may use the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction also can be called as “intra prediction direction” ) to derive a prediction of the current block. For example, prediction unit 203 may pass the derived one or more intra prediction modes or one or more angular prediction directions to intra prediction unit 205. In some implementations, intra prediction unit 205 may derive a prediction of the current block by fusing one or more predictions corresponding to the derived intra prediction modes. In some implementations, intra prediction unit 205 may derive a prediction of the current block by fusing one or more predictions determined according to the derived intra prediction modes and one or more predictions determined according to one or more preset mode (e.g., Planar mode, DC mode and cross-component prediction mode, etc. ) .
[0104] Prediction unit 203 may pass the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction may also be referred to as an “intra prediction direction” ) to transform unit 207. In one embodiment, transform unit 207 may use such information to determine transform kernel or a set of transform kernels in the primary transform and / or secondary transform.
[0105] Stil referring to FIG. 2, transform unit 207 performs a first transform (e.g., an integer transform) , which is originally designed based on a discrete cosine transform (DCT) , on the residual block. Transform unit 207 may determine whether a secondary transform is enabled to be applied on a block or not. When enabled, transform unit 207 further determines whether to apply the secondary transform to the coefficients obtained after performing the first transform. Transform unit 207 may generate transform coefficients by applying a transform technique to the residual signal, which is derived by the first adder 206 as a difference between the original samples of the current block and the prediction of the current block. For example, the transform technique may include at least one of a DCT, a discrete sine transform (DST) , a karhunen-loève transform (KLT) , a graph-based transform (GBT) , or a conditionally non-linear transform (CNT) . Here, GBT is a transform obtained from a graph when relationship information between pixels is represented by the graph. NT refers to the transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks having the same size or may be applied to blocks having a variable size rather than square.
[0106] Quantization unit 208 quantizes the coefficients from the transform unit 207.
[0107] Inverse quantization unit 209 performs scaling operations on the quantized coefficients to output reconstructed coefficients. Inverse transform unit 210 performs one or more inverse transforms corresponding to the transforms in transform unit 207 and outputs reconstructed residuals.
[0108] Second adder 211 calculates a reconstructed CU by adding the reconstructed residual and the prediction block of the CU output by prediction unit 202. Second adder 211 also forwards its output to prediction unit 202 to be used as an intra prediction reference. After all the CUs in a picture or a sub-picture have been reconstructed, filtering unit 212 performs in-loop filtering on the reconstructed picture or sub-picture. Filtering unit 212 contains one or more filters, for example, deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF) , luma mapping with chroma scaling (LMCS) filter, and neural network based filters, just to name a few. Alternatively, when filtering unit 212 determines that the CU is not used as reference for encoding other CUs, filtering unit 212 performs in-loop filtering on one or more target pixels in the CU.
[0109] In some implementations, filtering unit 212 may filter the reconstructed samples of one or more color components of the current block (e.g., a CU) . The encoder 200 stores the filtered reconstructed samples of one or more color components of the current block in a decoded picture buffer (DPB) 213 for a picture in which the current block is located. Thus, the prediction unit 202 may use the filtered samples of the current block in encoding the succeeding block of the current block in encoding order. For example, the prediction unit 202 may use the filtered samples of the current block to derive a prediction of succeeding block of the current block in an encoding order. For example, the prediction unit 202 (as well as other units in encoder 200) may include the filtered samples of the current block in a template and derive of a prediction, reorder candidate modes or parameters, and / or code parameters using the template-matching approach. Since the filtering unit 212 suppresses reconstruction distortion of the current block introduced by the lossy source coding, when the filtered sample of the current block is used to encode the succeeding block, the prediction efficiency of the succeeding coding block, and thus the coding performance of encoder 200, is improved.
[0110] In some implementations, the filtering unit 212 may use one or more fixed 1D or 2D filters to process the reconstruct sample of the current block. In some implementations, the 1D filter may be a symmetry filter. In some implementations, the 1D filter may be an asymmetry filter. In some implementations, the 2D filter may be a symmetry filter. In some implementations, the 2D filter may be an asymmetry filter. In some implementations, the 2D filter may be a separable filter. In some implementations, the 2D filter may be a non-separable filter.
[0111] In some implementations, the filtering unit 212 may use one or more adaptive 1D or 2D filters to process the reconstructed sample of the current block. In some implementations, the 1D filter may be a symmetry filter. In some implementations, the 1D filter may be an asymmetry filter. In some implementations, the 2D filter may be a symmetry filter. In some implementations, the 2D filter may be an asymmetry filter. In some implementations, the 2D filter may be a separable filter. In some implementations, the 2D filter may be a non-separable filter.
[0112] In some implementations, the filtering unit 212 may use one or more neural-network based filters to process the reconstruct sample of the current block.
[0113] In some implementations, the filtering unit 212 may use one or more filters of the spatial and / or temporal neighboring blocks of the current block. In some examples, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks before filtering and is invoked after reconstructing a picture where the neighboring block is located. In some examples, the filters from neighboring blocks may include the filter used to filter reconstructed sample of the neighboring blocks after reconstructing a picture where the neighboring block is located. One example is that filtering unit 212 may use the adaptive loop filter (ALF) which is used to filter a temporal neighboring block of the current block. In one example, the filtering unit 212 may select one or more existing filters, which are available before filtering the current block. One example is that the filters with parameters are signaled at block layer (e.g., coding tree unit or coding unit) or a layer higher than a block layer of the current block (e.g., video parameter set, sequence parameter set, picture parameter set, adaption parameter set, picture header and / or slice header) . In one example, filtering unit 212 may adaptively derive the parameter (s) of one or more filters to process the sample in the current block using spatial and / or temporal samples in one or more templates. In one example, filtering unit 212 may adaptively derive parameter (s) of one or more filters to process the sample in the current block based on the reconstructed samples and the original samples, and filtering unit 212 may pass the filter parameters to entropy coding unit 214, which signals the filter parameters in the bitstream.
[0114] In some implementations, filtering unit 212 may derive an indication parameter to indicate whether the reconstructed sample in the current block will be filtered. For example, the indication parameter can be a 1-bit flag. For example, the indication parameter can be a variable with a number of values indicating not only whether the reconstruct sample will be filtered but also which filter will be used. When the variable is equal to 0, the reconstructed sample of the current block will not be filtered; otherwise, when the variable is equal to 1, the reconstructed sample of the current block is filtered with a filter corresponding to an index equal to the value of this variable. Filter unit 212 will pass this indication parameter to entropy coding unit 214, which signals the parameter value in the bitstream.
[0115] In some implementations, filtering unit 212 may also set the indication parameter to indicate which color component will be filtered. Filtering unit 212 may choose to filter one or more of the luma and two chroma components. Filter unit 212 may pass this indication parameter to entropy coding unit 214 to signal the parameter value in the bitstream.
[0116] The output of filtering unit 212 is a decoded picture or sub-picture, which is into to decoded picture buffer (DPB) 213. DPB 213 outputs decoded pictures according to timing and controlling information. Pictures stored in DPB 213 may also be used as reference for performing inter or intra prediction by prediction unit 202.
[0117] Entropy coding unit 214 converts parameters from units in encoder 200 that are used for deriving the decoded picture. Entropy coding unit 214 may also derive control parameters and form supplemental information into binary representations. Moreover, entropy coding unit 214 may write such binary representations according to syntax structure of each data unit into a generated video bitstream.
[0118] Encoder 200 may provide control parameters to prediction unit 202 to instruct inter prediction unit 204 and intra prediction unit 205 whether one or more template-based prediction modes are enabled in the encoding process, or equivalently whether one or more template-based prediction modes are disabled in the encoding process. In the present disclosure, the embodiments are described from a perspective of “enabling” a template-based prediction mode. The embodiments from the perspective of “disabling” a template-based prediction mode can be directly derived based on the following descriptions.
[0119] In some implementations, encoder 200 may perform pre-analysis processing on the input video or picture to determine whether template-based prediction modes improve coding efficiency, e.g., such as perceptual quality. For example, template-based prediction modes may reduce perceptual quality of a picture or video containing complex texture and / or motion. One example of such picture or video would be a waterfront with random waves, and the reflection on the surface of the waterfront is random because of the small waves. In this case, encoder 200 may determine to disable all template-based prediction modes or enable only one or several template-based prediction modes in the encoding process. In one embodiment, encoder 200 can use a rate-distortion based method to make such decisions. In one embodiment, encoder 200 can use multi-pass encoding method to make such decisions. For instance, encoder 200 may code the input video or picture will all or several template-based prediction modes enabled. Then, encoder 200 may determine which of the template-based prediction modes are used in the second pass encoding.
[0120] In some implementations, the controlling parameters are obtained from configurations for encoder 200. For example, such controlling parameters are set in an encoder configuration file according to a pre-analysis on the input video or picture. For example, controlling parameters are set in an encoder configuration file according to complexity restrictions on encoder 200 and / or a decoder. For example, controlling parameters may be set according to conformance indications, e.g., such as indications by one or more of Profile, Tier, or Level. By way of example and not limitation, a Profile may specify that all of or several of template-based prediction modes are enabled or disabled.
[0121] Encoder 200 may send the controlling parameters to the entropy coding unit 214. Entropy coding unit 214 may perform entropy coding on the value of the controlling parameters, and then writes the coding bits into the bitstream. In another non-limiting example, if the controlling parameters are set according to conformance indications, such as indications by one or more of Profile, Tier, or Level, encoder 200 may not pass the controlling parameters to the entropy coding unit 214 to write the controlling parameters into the output bitstream because the entropy coding unit 214 has written the conformance indications such as Profile, Tier, or Level into the bitstream.
[0122] Tables 3A-3E, 4A-4G, and 5A-5E depict exemplary syntax elements for signaling the controlling parameters for template-based prediction modes. Table 3A: First Exemplary Syntax Element Table 3B: Second Exemplary Syntax Element Table 3C: Third Exemplary Syntax Element Table 3D: Fourth Exemplary Syntax Element Table 3E: Fifth Exemplary Syntax Element
[0123] Tables 3A-3E show exemplary syntax elements of “high level” or “collective” indications. Entropy coding unit 214 can write the exemplary syntax elements into the output bitstream.
[0124] In one embodiment as shown in Table 3A, encoder 200 can set an indication information by syntax element “tm_prediction_enable_flag” to indicate whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, prediction unit 202 may use the template-based prediction modes in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_flag is equal to 0, prediction unit 202 will not use the template-based prediction modes in encoding one or more blocks in the input video or picture.
[0125] In one embodiment as shown in Table 3B, encoder 200 can set an indication information for inter prediction. As shown in Table 3B, the indication information can be syntax element “tm_prediction_enable_for_inter_flag” to indicate whether template-based prediction modes can be used for inter predictions. When tm_prediction_enable_for_inter_flag is equal to 1, inter prediction unit 204 within prediction unit 202 may use the template-based inter prediction modes (as shown as example in Table 1) in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_for_inter_flag is equal to 0, inter prediction unit 204 within prediction unit 202 will not use the template-based inter prediction modes (as shown as example in Table 1) in encoding one or more blocks in the input video or picture.
[0126] In one embodiment as shown in Table 3C, encoder 200 can set an indication information for intra prediction. As shown in Table 3C, the indication information can be syntax element “tm_prediction_enable_for_intra_flag” to indicate whether template-based prediction modes can be used for intra predictions. When tm_prediction_enable_for_intra_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template-based intra prediction modes (as shown as example in Table 2) in encoding one or more blocks in the input video or picture. Otherwise, when tm_prediction_enable_for_intra_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template-based intra prediction modes (as shown as example in Table 2) in encoding one or more blocks in the input video or picture.
[0127] In one embodiment, as shown in Table 3D, encoder 200 can set indication information for a set of prediction modes. As shown in Table 3D, the indication information can be syntax element “tm_mode_setA_enable_flag, ” which indicates whether a set of template-based prediction modes can be used for intra and / or inter predictions. The set of template-based prediction modes can be referred to as “setA. ” The template-based prediction modes in “setA” may include one or more prediction modes. One example would be that “setA” includes IntraTMP and DMVR. When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template-based intra prediction modes in “setA” (e.g., IntraTMP in this example) , and inter prediction unit 204 within prediction unit 202 may use the template-based inter prediction modes in “setA” (e.g., DMVR in this example) to encode one or more blocks in the input video or picture. Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template-based intra prediction modes in “setA” (e.g., IntraTMP in this example) and inter prediction unit 204 within prediction unit 202 will not use the template-based inter prediction modes in “setA” (e.g., DMVR in this example) . By way of example and not limitation, assume that IntraTMP and TIMD (e.g., two intra prediction modes) are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 205 within prediction unit 202 may use the template-based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in this example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 205 within prediction unit 202 will not use the template-based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in this example) . By way of example and not limitation, assume that DMVR and BDOF (e.g., two inter prediction modes) are within “setA. ” Here, when tm_mode_setA_enable_flag is equal to 1, inter prediction unit 204 within prediction unit 202 may use the template-based intra prediction modes in “setA” (e.g., DMVR and BDOF in this example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, inter prediction unit 204 within prediction unit 202 will not use the template-based intra prediction modes in “setA” (e.g., DMVR and BDOF in this example) .
[0128] In one embodiment as shown in Table 3E, encoder 200 can set indication information for one prediction mode in Table 1 and / or Table 2 (e.g., referred to as “modeA” in this description) . As shown in Table 3E, the indication information can be syntax element “tm_modeA_enable_flag, ” which indicates whether the template-based prediction mode ( “modeA” ) can be used for intra and / or inter predictions. When tm_modeA_enable_flag is equal to 1, prediction unit 202 may use the template-based prediction mode ( “modeA” ) in encoding one or more blocks in the input video or picture. Otherwise, when tm_modeA_enable_flag is equal to 0, prediction unit 202 will not use the template-based prediction mode ( “modeA” ) in encoding one or more blocks in the input video or picture.
[0129] Tables 4A-4G show exemplary syntax elements, which may be implemented as a combination of the exemplary syntax elements shown in Tables 3A-3E to enable precise controlling options. Entropy coding unit 214 can write the example syntax elements into the output bitstream. Table 4A: Sixth Exemplary Syntax Element Table 4B: Seventh Exemplary Syntax Element Table 4C: Eighth Exemplary Syntax Element Table 4D: Nineth Exemplary Syntax Element Table 4E: Tenth Exemplary Syntax Element Table 4F: Eleventh Exemplary Syntax Element Table 4G: Twelfth Exemplary Syntax Element
[0130] In one embodiment, as shown in Table 4A, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for one or more template-based prediction modes. Table 4A provides an adaption of different template-based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those described above in connection with Tables 3A and 3E.
[0131] In one embodiment, as shown in Table 4B, encoder 200 indicates whether template-based prediction modes can be used for inter predictions (as shown as example in Table 1) based on the “tm_prediction_enable_for_inter_flag” syntax element. When tm_prediction_enable_for_inter_flag is equal to 1, encoder 200 can further set separate indications for one or more template-based inter prediction modes. Table 4B provides an adaption of different template-based inter prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Table 3B and Table 3E.
[0132] In one embodiment, as shown in Table 4C, encoder 200 indicates whether template-based prediction modes can be used for intra predictions (as shown as example in Table 2) based on the “tm_prediction_enable_for_intra_flag” syntax element. When tm_prediction_enable_for_intra_flag is equal to 1, encoder 200 can further set separate indications for one or more template-based intra prediction modes. Table 4C provides an adaption of different template-based intra prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Table 3C and Table 3E.
[0133] In one embodiment as shown in Table 4D, encoder 200 indicates whether template-based prediction modes can be used for inter and / or intra predictions in “setA” based on the “tm_mode_setA_enable_flag” syntax element. When tm_mode_setA_enable_flag is equal to 1, encoder 200 can further set separate indications for one or more template-based inter and / or intra prediction modes in “setA. ” Table 4D provides an adaption of different template-based inter and / or intra prediction modes in “setA” based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Table 3D and Table 3E.
[0134] In one embodiment, as shown in Table 4E, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template-based inter prediction modes and template-based intra prediction modes. Table 4E provides an adaption of different template-based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Tables 3A-3C.
[0135] In one embodiment as shown in Table 4F, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template-based inter prediction modes using “tm_prediction_enable_for_inter_flag” and several template-based prediction modes in “setA. ” For example, a template-based inter prediction mode can be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” while “setA” indicates a set of one or more template-based intra prediction modes are enabled. Table 4E provides an adaption of different template-based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Tables 3A, 3B and 3E.
[0136] In one embodiment as shown in Table 4G, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 can further set separate indications for template-based intra prediction modes using “tm_prediction_enable_for_intra_flag” and several template-based prediction modes in “setA. ” For example, a template-based intra prediction mode can be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” while “setA” may indicate a set of one or more template-based inter prediction modes are enabled. Table 4F provides an adaption of different template-based prediction modes based on characteristics of input video or picture. The controlling or indication by the example syntax elements are identical or similar to those in Tables 3A, 3B, and 3E.
[0137] Tables 5A-5E shows example syntax elements, which can be implemented as a combination of the example syntax elements shown in Tables 3A-3E and / or Tables 4A-4G to enable sophisticated controlling options. Entropy coding unit 214 can write the example syntax elements into the output bitstream. Table 5A: Thirteenth Exemplary Syntax Element Table 5B: Fourteenth Exemplary Syntax Element Table 5C: Fifteenth Exemplary Syntax Element Table 5D: Sixteenth Exemplary Syntax Element Table 5E: Seventeenth Exemplary Syntax Element
[0138] In one embodiment as shown in Table 5A, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in FIG. 3A based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of Table 3E for one template-based prediction mode (with the single mode being represented as “modeC” ) , Table 3D for one or more template-based prediction modes in “setA, ” and / or Table 4D for separate controlling or indication for the one or more template-based prediction modes in “set A. ” Table 5A provides an adaption of different template-based prediction modes based on characteristics of input video or picture.
[0139] In one embodiment as shown in Table 5B, encoder 200 indicates whether template-based inter prediction modes can be used for inter predictions (as shown as example in Table 1) as in Table 3B based on the “tm_prediction_enable_for_inter_flag” syntax element. When tm_prediction_enable_for_inter_flag is equal to 1, encoder 200 may use a combination of Table 3E for one template-based inter prediction mode (with the single mode being represented as “modeC” ) , Table 3D for one or more template-based inter prediction modes in “setA, ” and / or Table 4D for separate controlling or indication for the one or more template-based inter prediction modes in “set A. ” Table 5B provides an adaption of different template-based prediction modes based on characteristics of input video or picture.
[0140] In one embodiment as shown in Table 5C, encoder 200 indicates whether template-based intra prediction modes can be used for intra predictions (as shown as example in Table 2) as in Table 3C based on the “tm_prediction_enable_for_intra_flag” syntax element. When tm_prediction_enable_for_intra_flag is equal to 1, encoder 200 may use a combination of Table 3E for one template-based intra prediction mode (with the single mode being represented as “modeC” ) , Table 3D for one or more template-based intra prediction modes in “setA, ” and / or Table 4D for separate controlling or indication for the one or more template-based intra prediction modes in “set A. ” Table 5C provides an adaption of different template-based prediction modes based on characteristics of input video or picture.
[0141] In one embodiment as shown in Table 5D, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in Table 3A based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of Table 3B for template-based inter prediction modes, and template-based inter prediction modes can be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” Table 3D for one or more template-based intra prediction modes in “setA, ” Table 4D for separate controlling or indication for the one or more template-based intra prediction modes in “setA, ” and Table 3E for one template-based intra prediction mode (with the single mode being represented as “modeC” ) .
[0142] In one embodiment as shown in Tables 5E, encoder 200 indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in Table 3A based on the “tm_prediction_enable_flag” syntax element. When tm_prediction_enable_flag is equal to 1, encoder 200 may use a combination of Table 3B for template-based intra prediction modes, and template-based intra prediction modes can be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” Table 3D for one or more template-based inter prediction modes in “setA, ” Table 4D for separate controlling or indication for the one or more template-based inter prediction modes in “set A, ” and Table 3E for one template-based inter prediction mode (with the single mode being represented as “modeC” ) .
[0143] In Tables 3A-3E, Tables 4A-4G, and Tables 5A-5E, the descriptor may refer to an entropy coding method for the corresponding syntax element. The definition and algorithms of u (1) , u (n) , ue (v) and ae (v) are the same as those in the H. 265 / HEVC standard.
[0144] In some implementations, entropy coding unit 214 in encoder 200 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in one or more of the following data units in the output bitstream. In one embodiment, the level order from high to low is sequence level, picture level, slice level, and block level. The indications or controlling parameters in lower levels may overwrite or override the counterparts in higher levels. Prediction unit 202 in encoder 200 may follow the final valid instruction or controlling parameters in deriving the prediction of the current coding block using the template-based prediction modes.
[0145] A sequence level data unit may be valid for all pictures in a coded video sequence. An example of sequence level data unit can be one or more of video parameter set (VPS) , sequence parameter set (SPS) , picture parameter set (PPS) with parameters for all pictures in a coded video sequence, adaptation parameter set (APS) with parameters for all pictures in a coded video sequence, picture header with parameters for all pictures in a coded video sequence, supplemental enhancement information (SEI) with parameters for all pictures in a coded video sequence. Entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in one or more sequence level data units. For example, entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in the sequence level data unit. For example, entropy coding unit 214 may write syntax elements in Tables 3A-3E in SPS, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in PPS and / or APS directly or indirectly referring to the said SPS.
[0146] A picture level data unit may be valid for one picture. An example of picture level data unit can be one or more of PPS, APS with parameters for one picture, picture header, slice header (s) of one picture with parameters, SEI with parameters for one picture. Entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in one or more picture level data units, e.g., in picture header, PPS and / or APS. For example, entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in a picture level data unit. For example, entropy coding unit 214 may write the syntax elements of Tables 3A-3E in PPS, and the syntax elements Tables 4A-4G and / or Tables 5A-5E in picture header, slice header, and / or APS directly or indirectly referring to the said PPS. For example, entropy coding unit 214 writes the syntax elements of Tables 3A-3E in picture header, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in slice header and / or APS directly or indirectly referred to by the said picture header. For example, entropy coding unit 214 may write the syntax elements of Tables 3A-3E in APS, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in picture header and / or slice header directly or indirectly referring to the said APS.
[0147] A slice level data unit may be valid for one slice. An example of slice level data unit can be one or more of slice header, APS, SEI for a slice. Entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in one or more slice level data units. For example, entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in a slice level data unit. For example, entropy coding unit 214 writes the syntax elements of Tables 3A-3E in slice header, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in APS directly or indirectly referred to by the slice header. For example, entropy coding unit 214 may write syntax elements in Tables 3A-3E in slice header, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in slice header directly or indirectly referring to the said APS.
[0148] A block level data unit may be valid for one or more CTU, CU, coding block, and / or transform block. Entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in one or more block level data units. For example, entropy coding unit 214 may write the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E in a block level data unit. For example, entropy coding unit 214 writes the syntax elements of Tables 3A-3E in a CTU, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in a CU that is within the CTU. For example, entropy coding unit 214 may write the syntax elements of Tables 3A-3E in a first CU, and the syntax elements of Tables 4A-4G and / or Tables 5A-5E in the CU that is within the first CU.
[0149] Encoder 200 may be implemented as a computing device with a processor and a storage medium recording an encoding program. When the processor reads and executes the encoding program, the encoder 200 reads an input video and generates a corresponding video bitstream.
[0150] Encoder 200 may be implemented as a computing device with one or more chips. The units, implemented as integrated circuits, on the chip are of similar functionalities and with similar connections and data exchanges as the corresponding ones in FIG. 2.
[0151] FIG. 7 illustrates a block diagram of an exemplary decoder 700, according to some embodiments of the present disclosure. As shown, the input of decoder 700 may be a bitstream representing a compressed version of a video or a still picture. The output of decoder 700 may be a decoded video including a sequence of pictures or a decoded still picture.
[0152] Referring to FIG. 7, decoder 700 may receive an input bitstream generated by the encoder 200. Parsing unit 701 parses the input bitstream and obtains values of syntax elements from the input bitstream. Parsing unit 701 converts binary representations of syntax elements to numerical values and forwards the numerical values to the units in the decoder 700 to derive one or more decoded pictures. Parsing unit 701 may also parse one or more syntax elements from the input bitstream for rendering the decoded pictures.
[0153] The syntax elements shown above in Tables 3A-3E indicate “high level” or “collective” indications. Parsing unit 701 can process the example syntax elements in Tables 3A-3E in the bitstream to determine corresponding values of such syntax elements.
[0154] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 3A. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, prediction unit 702 may use the template-based prediction modes to decode one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_flag is equal to 0, prediction unit 702 may not use the template-based prediction modes to decode one or more blocks in the video or picture bitstream.
[0155] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 3B. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag, ” which indicates whether template-based prediction modes can be used for inter predictions. When tm_prediction_enable_for_inter_flag is equal to 1, inter prediction unit 703 within prediction unit 702 may use the template-based inter prediction modes (as shown as example in Table 1) to decode one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_for_inter_flag is equal to 0, inter prediction unit 703 will not use the template-based inter prediction modes (as shown as example in Table 1) to decode one or more blocks in the video or picture bitstream.
[0156] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 3C. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag, ” which indicates whether template-based prediction modes can be used for intra predictions. When tm_prediction_enable_for_intra_flag is equal to 1, intra prediction unit 704 within prediction unit 702 may use the template-based intra prediction modes (as shown as example in Table 2) to decode one or more blocks in the video or picture bitstream. Otherwise, when tm_prediction_enable_for_intra_flag is equal to 0, intra prediction unit 704 within prediction unit 702 will not use the template-based intra prediction modes (as shown as example in Table 2) to decode one or more blocks in the video or picture bitstream.
[0157] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 3D. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_mode_setA_enable_flag, ” which indicates whether several template-based prediction modes can be used for intra and / or inter predictions. These template-based prediction modes can be viewed or classified as “setA. ” “setA” can contain one or more prediction modes. One example would be that IntraTMP and DMVR are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 704 may use the template-based intra prediction modes in “setA” (e.g., IntraTMP in this example) , and inter prediction unit 703 may use the template-based inter prediction modes in “setA” (e.g., DMVR in this example) to decode one or more blocks in the video or picture bitstream. Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 704 will not use the template-based intra prediction modes in “setA” (e.g., IntraTMP in this example) , and inter prediction unit 703 will not use the template-based inter prediction modes in “setA” (e.g., DMVR in this example) in decoding one or more blocks in the video or picture bitstream. One example would be that IntraTMP and TIMD (that is, two intra prediction modes) are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, intra prediction unit 704 may use the template-based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in this example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, intra prediction unit 704 will not use the template-based intra prediction modes in “setA” (e.g., IntraTMP and TIMD in this example) . One example would be that DMVR and BDOF (that is, two inter prediction modes) are within “setA. ” When tm_mode_setA_enable_flag is equal to 1, inter prediction unit 703 may use the template-based intra prediction modes in “setA” (e.g., DMVR and BDOF in this example) . Otherwise, when tm_mode_setA_enable_flag is equal to 0, inter prediction unit 703 will not use the template-based intra prediction modes in “setA” (e.g., DMVR and BDOF in the above example) .
[0158] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 3E. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_modeA_enable_flag, ” which indicates whether the template-based prediction mode ( “modeA” ) can be used for intra and / or inter predictions. When tm_modeA_enable_flag is equal to 1, prediction unit 702 may use the template-based prediction mode ( “modeA” ) to decode one or more blocks in the video or picture bitstream. Otherwise, when tm_modeA_enable_flag is equal to 0, prediction unit 702 will not use the template-based prediction mode ( “modeA” ) to decode one or more blocks in the video or picture bitstream.
[0159] The exemplary syntax elements in Tables 4A-4G may be implemented as a combination of the example syntax elements shown in Tables 3A-3E to enable sophisticated controlling options. Parsing unit 701 can process these example syntax elements in the bitstream to determine corresponding values of such syntax elements.
[0160] For instance, in some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4A. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, parsing unit 701 may further determine separate indications for one or more template-based prediction modes according to the additional syntax elements in the bitstream. Parsing unit 701 determines the controlling or indication by the example syntax elements in the same or similar way to those described above in connection with Table 3A and Table 3E.
[0161] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4B. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag, ” which indicates whether template-based prediction modes can be used for inter predictions (as shown as example in Table 1) . When tm_prediction_enable_for_inter_flag is equal to 1, parsing unit 701 may further determine separate indications for one or more template-based inter prediction modes according to the additional syntax elements in the bitstream. Parsing unit 701 determines the controlling or indication by the example syntax elements in the same or similar way as that described above in connection with Table 3B and Table 3E.
[0162] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4C. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag, ” which indicates whether template-based prediction modes can be used for intra predictions (as shown as example in Table 2) . When tm_prediction_enable_for_intra_flag is equal to 1, parsing unit 701 may further determine separate indications for one or more template-based intra prediction modes according to the additional syntax elements in the bitstream. Parsing unit 701 determines further the controlling or indication by the example syntax elements are identical or similar to those in Table 3C and Table 3E.
[0163] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4D. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_mode_setA_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and / or intra predictions in “setA. ” When tm_mode_setA_enable_flag is equal to 1, parsing unit 701 may further determine separate indications for one or more template-based inter and / or intra prediction modes in “setA” according to the additional syntax elements in the bitstream. Parsing unit 701 determines the controlling or indication by the example syntax elements in the same or similar way as described above in connection with Table 3D and Table 3E.
[0164] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4E. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, parsing unit 701 may further determine separate indications for template-based inter prediction modes and template-based intra prediction modes according to the additional syntax elements in the bitstream. Parsing unit 701 determines further the controlling or indication by the example syntax elements are identical or similar to those in Table 3A, Table 3B, and Table 3C.
[0165] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4F. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, parsing unit 701 may further determine separate indications for template-based inter prediction modes using “tm_prediction_enable_for_inter_flag” and several template-based prediction modes in “setA” according to the additional syntax elements in the bitstream. For example, a template-based inter prediction mode can be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” “setA” may contain one or more template-based intra prediction modes. Parsing unit 701 determines the controlling or indication by the example syntax elements in the same or similar way as described above in connection with Table 3A, Table 3B and Table 3E.
[0166] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 4G. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) . When tm_prediction_enable_flag is equal to 1, parsing unit 701 may further determine separate indications for template-based intra prediction modes using “tm_prediction_enable_for_intra_flag” and several template-based prediction modes in “setA” according to the additional syntax elements in the bitstream. For example, a template-based intra prediction modes can be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” “setA” may contain one or more template-based inter prediction modes. Parsing unit 701 determines the controlling or indication by the example syntax elements in the same or similar way as described above in connection with Table 3A, Table 3B and Table 3E.
[0167] The exemplary syntax elements in Tables 5A-5E may be implemented as a combination of the example syntax elements shown in Tables 3A-3E and / or Tables 4A-4G to enable sophisticated controlling options. Parsing unit 701 can process the example syntax elements in the bitstream to determine corresponding values of such syntax elements.
[0168] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 5A. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag, ” which indicates whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in Table 3A. When tm_prediction_enable_flag is equal to 1, according to a combination of syntax elements in Table 3E for one template-based prediction mode (with the single mode being represented as “modeC” ) , in Table 3D for one or more template-based prediction modes in “setA, ” and in Table 4D for separate controlling or indication for the one or more template-based prediction modes in “set A, ” parsing unit 701 may further determine corresponding indication or controlling parameters.
[0169] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 5B. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_inter_flag, ” which indicates whether template-based inter prediction modes can be used for inter predictions (as shown as example in Table 1) as in Table 3B. When tm_prediction_enable_for_inter_flag is equal to 1, according to a combination of syntax elements in Table 3E for one template-based inter prediction mode (with the single mode being represented as “modeC” ) , in Table 3D for one or more template-based inter prediction modes in “setA, ” and in Table 4D for separate controlling or indication for the one or more template-based inter prediction modes in “set A, ” parsing unit 701 may further determine corresponding indication or controlling parameters.
[0170] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 5C. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_for_intra_flag, ” which indicates whether template-based intra prediction modes can be used for intra predictions (as shown as example in Table 2) as in Table 3C. When tm_prediction_enable_for_intra_flag is equal to 1, according to a combination of syntax elements in Table 3E for one template-based intra prediction mode (with the single mode being represented as “modeC” ) , in Table 3D for one or more template-based intra prediction modes in “setA, ” and in Table 4D for separate controlling or indication for the one or more template-based intra prediction modes in “setA, ” parsing unit 701 may further determine corresponding indication or controlling parameters.
[0171] In some implementations, parsing unit 701 can process the bitstream containing one or more of the syntax elements as shown in Table 5D. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” to indicate whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in Table 3A. When tm_prediction_enable_flag is equal to 1, according to a combination of syntax elements in Table 3B for template-based inter prediction modes, and a template-based inter prediction modes can be collectively controlled or indicated by “tm_prediction_enable_for_inter_flag, ” in Table 3D for one or more template-based intra prediction modes in “setA, ” in Table 4D for separate controlling or indication for the one or more template-based intra prediction modes in “set A, ” and in Table 3E for one template-based intra prediction mode (with the single mode being represented as “modeC” ) , parsing unit 701 may further determine corresponding indication or controlling parameters.
[0172] In some implementations, decoder 700 can process the bitstream containing one or more of the syntax elements as shown in Table 5E. Parsing unit 701 uses one of the entropy decoding methods in the “descriptor” to get a value of syntax element “tm_prediction_enable_flag” to indicate whether template-based prediction modes can be used for inter and intra predictions (as shown as example in Table 1 and Table 2) as in Table 3A. When tm_prediction_enable_flag is equal to 1, according to a combination of syntax elements in Table 3B for template-based intra prediction modes, and a template-based intra prediction modes can be collectively controlled or indicated by “tm_prediction_enable_for_intra_flag, ” in Table 3D for one or more template-based inter prediction modes in “setA, ” in Table 4D for separate controlling or indication for the one or more template-based inter prediction modes in “set A, ” and in Table 3E for one template-based inter prediction mode (with the single mode being represented as “modeC” ) , parsing unit 701 may further determine corresponding indication or controlling parameters.
[0173] In Tables 3A-3E, Tables 4A-4G, and Tables 5A-5E, the descriptor refers to an entropy decoding method for the corresponding syntax element. The definition and algorithms of u (1) , u (n) , ue (v) and ae (v) are the same as those in the H. 265 / HEVC standard.
[0174] In some implementations, parsing unit 701 determines the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E from one or more of the following data units in the input bitstream. The “parameters” mentioned below may refer to the example syntax elements in Tables 3A-3E, 4A-4G, and 5A-5E; and the parameters may indicate whether one or more template-based prediction modes are enabled in decoding one or more blocks in the input bitstream. In one embodiment, an order of the levels from high to low is sequence level, picture level, slice level, and block level. The indications or controlling parameters in lower levels may overwrite or override the counterparts in higher levels. Parsing unit 701 may pass the final valid indication or controlling parameters to prediction unit 702 and / or scaling unit 705 in decoder 700. Prediction unit 702 will follow the final valid instruction or controlling parameters in deriving the prediction of one or more blocks using the template-based prediction modes.
[0175] For instance, a sequence level data unit is valid for all pictures in a coded video sequence. An example of sequence level data unit can be one or more of video parameter set (VPS) , sequence parameter set (SPS) , picture parameter set (PPS) with consistent parameters for all pictures in a coded video sequence, adaptation parameter set (APS) with consistent parameters for all pictures in a coded video sequence, picture header with consistent parameters for all pictures in a coded video sequence, supplemental enhancement information (SEI) with consistent parameters for all pictures in a coded video sequence.
[0176] If the parsing unit 701 determines the instruction or controlling parameters according to syntax elements in a sequence level data unit, prediction unit 702 in decoder 700 may use the template-based prediction modes that can be enabled according to the instruction or controlling parameters from the parsing unit 701 in decoding one or more blocks in this coded video sequence in the input bitstream. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a sequence level data unit can also be overwritten or overridden by the instruction or controlling parameters determined by parsing lower level (e.g., one or more of picture level, slice level and block level) data units.
[0177] A picture level data unit is valid for one picture. An example of picture level data unit can be one or more of picture parameter set (PPS) , adaptation parameter set (APS) with consistent parameters for one picture, picture header, slice header (s) of one picture with consistent parameters, supplemental enhancement information (SEI) with consistent parameters for one picture.
[0178] If the parsing unit 701 determines the instruction or controlling parameters according to syntax elements in a picture level data unit, prediction unit 702 may use the template-based prediction modes. The template-based prediction modes may be enabled according to the instruction or controlling parameters from the parsing unit 701 in decoding one or more blocks in this picture in the input bitstream. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a picture level data unit can also be overwritten or overridden by the instruction or controlling parameters determined by parsing lower level (e.g., one or more of slice level and block level) data units. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a picture level data unit can also overwrite or override the instruction or controlling parameters determined by parsing higher level (e.g. sequence level) data unit.
[0179] A slice level data unit may be valid for one slice. An example of slice level data unit can be one or more of slice header, adaptation parameter set (APS) , supplemental enhancement information (SEI) for a slice. In a slice level data unit, ae (v) would not be used.
[0180] If the parsing unit 701 determines the instruction or controlling parameters according to syntax elements in a slice level data unit, prediction unit 702 may use the template-based prediction modes. The template-based prediction modes may be enabled according to the instruction or controlling parameters from the parsing unit 701 in decoding one or more blocks in this slice in the input bitstream. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a slice level data unit can also be overwritten or overridden by the cost function determined by parsing lower level (e.g., block level) data units. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a picture level data unit can also overwrite or override the instruction or controlling parameters determined by parsing higher level (e.g., one or more of sequence level, picture level) data units.
[0181] A block level data unit may be valid for one or more of a CTU, a CU, a coding block, and / or a transform block. In a block level data unit, ae (v) may be used.
[0182] If the parsing unit 701 determines the instruction or controlling parameters according to syntax elements in a block level data unit, prediction unit 702 may use the template-based prediction modes, which may be enabled according to the instruction or controlling parameters from the parsing unit 701 in decoding the block in the bitstream. In some implementations, the instruction or controlling parameters determined by the parsing unit 701 from a block level data unit can also overwrite or override the cost function determined by parsing higher level (e.g., one or more of sequence level, picture level, slice level) data units.
[0183] Parsing unit 701 forwards the instruction or controlling parameters to indicate template-based prediction modes that can be enabled in deriving prediction of a block by the prediction unit 702, the values of other syntax elements, as well as one or more variables set or determined according the values of syntax elements, for deriving one or more decoded pictures to the units in the decoder 700. Prediction unit 702 determines a prediction block of a current decoding block (e.g., a CU) . When it is indicated that an inter coding mode is used to decode the current decoding block, prediction unit 702 passes relative parameters from parsing unit 701 to inter prediction unit 703 to derive inter prediction block. When it is indicated that an intra prediction mode is used to decoding the current decoding block, prediction unit 702 passes relative parameters from parsing unit 701 to intra prediction unit 704 to derive intra prediction block.
[0184] In one embodiment, parsing unit 701 may determine controlling parameters according to conformance indications, such as indications by one or more of Profile, Tier, and Level. One example implementation would be that a Profile may specify that all of or several of template-based prediction modes are enabled or disabled.
[0185] In one embodiment, if a template-based inter prediction mode is enabled and used in decoding a block, inter prediction unit 703 derives a reference template. One example of deriving the reference template of a current decoding block is the same as that shown in FIG. 4, using the template as an example shown in FIG. 5. In one example, the prediction block of the current decoding block can be derived based on the reference template, e.g., in a same way as that described for encoder 200. In some implementations, inter prediction unit 703 can use a unified cost function for a same or similar process, using a template. For example, for “reordering” functions as listed in Table 1, inter prediction unit 703 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, inter prediction unit 703 can use SAD or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In some implementations, inter prediction unit 703 can use a single cost function for all prediction modes.
[0186] In one embodiment, if a template-based intra prediction mode is enabled and used in decoding a block, intra prediction unit 704 derives a reference template with the determined cost function. One example of deriving the reference template of a current decoding block is the same as that shown in FIG. 6, using the template as an example shown in FIG. 5. In one example, the prediction block of the current decoding block can be derived based on the reference template, e.g., in a same way as that described for encoder 200. In some implementations, intra prediction unit 704 can use a unified cost function for a same or similar process using a template. For example, for “reordering” functions as listed in Table 2, intra prediction unit 704 can use SATD for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. For example, for “reordering” functions as listed in Table 1, intra prediction unit 704 can use SAD, or any one of the abovementioned cost function for one, multiple but not all, or all of the prediction modes having “reordering” of candidates in a candidate list. In an embodiment, intra prediction unit 704 can use a single cost function for all prediction modes.
[0187] Prediction unit 702 may derive intra prediction mode or angular prediction direction of the current CU based on one or more reference samples. As mentioned above, FIG. 13 depicts an example of a current block 1301 and its reference sample 1302, 1303. For example, the reference sample 1302, 1303, which are marked as black dot outside the current block 1301, may be one or more samples in a template ( “L-shape” template consisting of the black dots in FIG. 13) as shown in FIG. 5. For example, a gradient of reference sample 1302, 1303 can be derived by applying one or more filters to process one or more samples in a template as shown in FIG. 5. Prediction unit 702 first derives a gradient of a reference sample using an operator. Generally, the operator can be used to detect an edge or a gradient in a picture. The operator can be a 2-dimentional (2D) MxN filter, wherein M and N are positive integers, and M can be equal to or different from N. An example of the operator is Sobel filter. Two different examples of 3x3 Sobel filter are shown in FIGs. 14A and 14B, respectively. Another example of the operator is an edge filter. Two different examples of a 2x2 edge filter are shown in FIGs. 14C and 14D, respectively.
[0188] In an example, prediction unit 702 can choose different operators according to a width and / or a height of the current block. For example, prediction unit 702 uses smaller operator for smaller block, and uses larger operator for larger block. One example would be that prediction unit 702 uses the abovementioned Edge filter when a size (e.g., width x height) of the current block is 4x4, 4x8 or 8x4, and uses the abovementioned Sobel filter for other sizes of the current block.
[0189] Prediction unit 702 derives a Histogram of Gradients (HoG) by analyzing one or more reference sample marked in black dots in FIG. 13. The template in FIG. 13 comprises three reference sample lines above and three reference sample columns on the left of the current block. The HoG is derived by accumulating the magnitudes of one or more gradients at one or more given directions, for one, a part of, or all of the reference samples as shown in FIG. 13. One or more directions indicated by the gradients with highest or higher cumulative magnitudes may be used as the angular prediction direction or the intra prediction mode of the current block.
[0190] As an example, when prediction unit 702 uses one reference sample to derive a direction, the prediction unit 702 can determine the direction as the one indicated by a gradient derived at this reference sample.
[0191] As an example, when prediction unit 702 uses a plurality of reference samples to derive directions, prediction unit 702 may choose a preset number of samples from the reference samples. For example, prediction unit 702 choose J reference samples above the current block and K reference samples left to the current block, where J and K are integers greater than or equal to 0. For example, J and K are equal to 2, J is equal to 4 and K is equal to 8, or J plus K is equal to 4. Prediction unit 702 derives the gradients at the selected reference samples using the operator and derives the HoG. One or more directions indicated by the gradients with the highest or higher cumulative magnitudes may be used as the angular prediction direction or the intra prediction mode of the current block.
[0192] As an example, prediction unit 702 can adaptively determine one or more reference samples used for deriving a HoG. Prediction unit 702 uses a preset scanning order of the reference samples. When scanning a reference sample, prediction unit 702 derives a gradient at this reference sample and updates the accumulation of the magnitudes according to this gradient at one or more given directions in HoG. When prediction unit 702 determines that the total cumulative amplitude is greater or equal than a given threshold, prediction unit 702 may stop scanning the remaining reference sample and stop deriving new gradient. The resulted HoG at the termination of the scanning by prediction unit 702 is determined as the HoG to derive intra prediction mode or angular prediction direction.
[0193] Examples of the abovementioned preset scanning order of the reference samples can be the following.
[0194] For example, a scanning order (A) can be scanning the left column of reference samples as shown in FIG. 13 from bottom to top; a scanning order (B) can be scanning the left column of reference samples as shown in FIG. 13 from top to bottom; a scanning order (C) can be scanning the above line of reference samples as shown in FIG. 13 from left to right; a scanning order (D) can be scanning the above line of reference samples as shown in FIG. 13 from left to right.
[0195] An example of the preset scanning order can be one or more of “first order (A) then order (C) , ” “first order (A) then order (D) , ” “first order (B) then order (C) , ” “first order (B) then order (D) , ” “first order (C) then order (A) , ” “first order (D) then order (A) , ” “first order (C) then order (B) ” and “first order (D) then order (B) . ”
[0196] An example of the preset scanning order may include scanning in an interleaving manner. For example, one or multiple reference samples from the left column and then one or multiple second reference samples from the above line and then one or multiple third reference sample from left column may be scanned. For example, one or multiple reference samples from the above line and then one or multiple second reference samples from the left column and then one or multiple third reference sample from above line may be scanned. Additionally, as an example, the scanning order of samples in the left column can be one or more of order (A) and (B) , and the scanning order of samples in the left column can be one or more of order (C) and (D) .
[0197] Prediction unit 702 may use the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction may also be referred to as “intra prediction direction” ) to derive a prediction of the current block. For example, prediction unit 702 may pass the derived one or more intra prediction modes or one or more angular prediction directions to intra prediction unit 704. In one embodiment, intra prediction unit 704 may derive a prediction of the current block by fusing one or more predictions corresponding to the derived intra prediction modes. In one embodiment, intra prediction unit 704 may derive a prediction of the current block by fusing one or more predictions determined according to the derived intra prediction modes and one or more predictions determined according to one or more preset mode (e.g., Planar mode, DC mode, cross-component prediction mode, etc. ) . One example can be Decoder Side Intra Mode Derivation (DIMD) .
[0198] In some embodiments, a DIMD_flag may indicate whether the DIMD mode is enabled or disabled.
[0199] If the DIMD_flag is determined to be 1 (enable) , whether to use the Occurrence-based intra coding (OBIC) mode may be determined. For example, an OBIC_flag may be decoded to determine whether the OBIC mode is enabled or disabled.
[0200] The DIMD_flag may be signaled in the sequence level, picture level, slice level or block level and so on.
[0201] The OBIC_flag may be signaled in the sequence level, picture level, slice level or block level and so on.
[0202] In some embodiments, the OBIC mode derives the intra prediction modes of the current block based on the sample-wise occurrence of the intra modes in the spatial neighborhood of the block. For this, adjacent and non-adjacent spatial neighboring blocks are checked and the intra prediction modes of the blocks are collected into an occurrence histogram. Instead of Histogram of Gradient (HoG) as in DIMD, the OBIC method uses the Histogram of oCcurrence (HoC) , which consists of the intra modes and their sample-wise occurrences. The occurrence values are calculated based on the number of samples that are coded in a certain intra prediction mode in that neighborhood. For example, if a uiWidth × uiHeight block is coded with an intra prediction mode (IPM) mode, the occurrence of the mode (HoC [IPM] ) in that particular block may be calculated according to equation (4) . HoC [IPM] += uiWidth *uiHeight (4) . where uiWidth and uiHeight are the width and height of a spatial neighboring block.
[0203] The occurrences of the existing modes from the spatial neighborhood blocks may be accumulated into the histogram.
[0204] FIG. 15 illustrates a diagram of non-adjacent spatial neighboring candidates 1500 for occurrence-based intra coding (OBIC) mode, according to some embodiments of the present disclosure.
[0205] Referring to FIG. 15, up to five angular modes with the highest occurrence along with the planar mode or block vector based prediction (same as in DIMD) are selected from the HoC and used for the final prediction by blending the prediction of the selected modes.
[0206] Some blocks use more than one intra mode for prediction. In such cases, all the intra modes of such blocks are selected and used when creating the OBIC histogram. For instance, for DIMD, up to 5 angular modes may be used. For TIMD, up to 2 modes may be used. For SGPM, 2 modes may be used. For OBIC, up to 5 angular modes may be used.
[0207] Moreover, the virtual intra prediction modes (VIPMs) of the following blocks are considered only in inter slices when creating the histogram of OBIC mode: MIP block, IntraTMP block, IBC block, EIP block, etc.
[0208] The blending weights are calculated as a similar manner as in DIMD mode; however, instead of using gradient values from the template, the occurrence values are used for OBIC. Moreover, the planar mode’s weight is also decided similar to the DIMD mode.
[0209] FIG. 16 illustrates an example of the histogram-of-occurrences (HoC) 1600 of intra prediction modes (IPMs) , according to some embodiments of the present disclosure.
[0210] Referring to FIG. 16, in some embodiments, in order to decrease the buffer memory, the OBIC method may use non-sample-wise occurrences, such as block-wise occurrences. For example, if a uiWidth× uiHeight block is coded with an intra prediction mode (IPM) mode, the block-wise occurrence of the mode (HoC [IPM] ) in that particular block may be calculated according to equation (5) . HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) (5) , where the variable shift1 and shift2 are both positive integers greater than or equal to 1.
[0211] When the variable shift1 and shift2 are set to 1, the occurrence type 2×2 block-wise occurrence is used. The variable shift 1 or shift 2 can also be set to values of 2, 3, 4, 5, etc. Variable shift1 may or may not be equal to shift2.
[0212] In some embodiments, variable shift1 or shift 2 may be determined according to the size (width or height) of the current block. For example, if the size of the current block is 4×4, the 2×2 block-wise occurrence may be used.
[0213] In some embodiments, variable shift1 or shift2 may be determined based on the flag sps_log2_min_luma_coding_block_size_minus2. For example, parsing unit 701 may parse a bitstream and obtain the value of sps_log2_min_luma_coding_block_size_minus2. If the value of sps_log2_min_luma_coding_block_size_minus2 is equal to 2, the size of the current block is determined to 4×4. Then, the occurrence type can be obtained through the determined size of the current block.
[0214] In an embodiment, the block-wise occurrence can be determined through a look-up table. Tables 6A-6E illustrate various non-limiting examples of look-up tables that correlate CU size to occurrence type. Table 6A: First Example Occurrence-Type Look-up Table Table 6B: Second Example Occurrence-Type Look-up Table Table 6C: Third Example Occurrence-Type Look-up Table Table 6D: Fourth Example Occurrence-Type Look-up Table Table 6E: Fifth Example Occurrence-Type Look-up Table
[0215] In an embodiment, a confidence level may also be used to calculate the occurrence. If there is a very large size block neighboring a small size block, the intra prediction mode (IPM) of the large size block is considered as a low-confidence level block and the large size block is not used to calculate the occurrence of the current block. For example, if a 64×64 block neighbors a 2×2 block, the 64×64 block may not be used to calculate the occurrence of the 2×2 block due to the size difference between the 2×2 block and the 64×64 block. This is because if the 64×64 block is used to calculate, it will negatively interfere with the histogram statistics. The size of current block is a factor the determined the confidence level of a neighbor block.
[0216] In some embodiments, the OBIC mode may be used to predict luma blocks.
[0217] In some embodiments, the OBIC mode can be used to predict chroma blocks.
[0218] Prediction unit 702 may pass the derived one or more intra prediction modes or one or more angular prediction directions (intra prediction mode or angular prediction direction may also be referred to as “intra prediction direction” ) to transform unit 706. In one embodiment, transform unit 706 may use such information to determine transform kernel or a set of transform kernels in the primary transform and / or secondary transform.
[0219] Adder 707 performs an addition operation with its inputs of prediction block from prediction unit 702 and reconstructed residual from transform unit 706 to get reconstructed block of the current decoding block. The reconstructed block is also sent to prediction unit 702 to be used as reference for other blocks coded in intra prediction mode.
[0220] In one embodiment, after the CUs in a picture or a sub-picture have been reconstructed, filtering unit 708 performs in-loop filtering on the reconstructed picture or sub-picture. Filtering unit 708 contains one or more filters, e.g., deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF) , luma mapping with chroma scaling (LMCS) filter and neural network based filters. Alternatively, when filtering unit 708 determines that the reconstructed block is not used as reference for decoding other blocks, filtering unit 708 performs in-loop filtering on one or more target pixels in the reconstructed block.
[0221] In some embodiments, filtering unit 708 may filter the reconstructed samples of one or more color components of the current block (e.g., a CU) . The decoder 700 stores the filtered reconstructed samples of one or more color components of the current block in DPB 709 for a picture in which the current block is located. Thus, the prediction unit 702 may use the filtered samples of the current block to decode the succeeding block of the current block in a decoding order. For example, the prediction unit 702 may use the filtered samples of the current block to derive a prediction of a succeeding block of the current block in a decoding order. For example, the prediction unit 702 (as well as other units in decoder 700) may include the filtered samples of the current block in a template and derive a prediction, reorder candidate modes or parameters, and / or decode parameters using a template-matching approach. Since the filtering unit 708 suppresses reconstruction distortion of the current block introduced by the lossy source coding of encoder 200, when the filtered sample of the current block is used to decode the succeeding block, the prediction efficiency of the succeeding block, and thus the coding efficiency, may be improved.
[0222] In some embodiments, the filtering unit 708 may use one or more fixed 1D or 2D filters to process the reconstruct sample of the current block. In some embodiments, the 1D filter may be a symmetry filter. In some embodiments, the 1D filter may be an asymmetry filter. In some embodiments, the 2D filter may be a symmetry filter. In some embodiments, the 2D filter may be an asymmetry filter. In some embodiments, the 2D filter may be a separable filter. In some embodiments, the 2D filter may be a non-separable filter.
[0223] In some embodiments, the filtering unit 708 may use one or more adaptive 1D or 2D filters to process the reconstruct sample of the current block. In some embodiments, the 1D filter may be a symmetry filter. In some embodiments, the 1D filter may be an asymmetry filter. In some embodiments, the 2D filter may be a symmetry filter. In some embodiments, the 2D filter may be an asymmetry filter. In some embodiments, the 2D filter may be a separable filter. In some embodiments, the 2D filter may be a non-separable filter.
[0224] In some embodiments, the filtering unit 708 may use one or more neural-network based filters to process the reconstruct sample of the current block.
[0225] In some embodiments, the filtering unit 708 may use one or more filters of the spatial and / or temporal neighboring blocks of the current block. In some examples, the filters from neighboring blocks may include the filter used to filter the reconstructed sample of the neighboring blocks before filtering and that is invoked after reconstructing a picture where the neighboring block is locates. In some examples, the filters from neighboring blocks may include the filter used to filter the reconstructed sample of the neighboring blocks after reconstructing a picture where the neighboring block is located. One example is that the filtering unit 708 may use the adaptive loop filter (ALF) , which is used to filter a temporal neighboring block of the current block. In some examples, the filtering unit 708 may select one or more existing filters, which are available before filtering the current block. One example is that the filters with parameters are obtained, by parsing unit 701, from a block layer (e.g., coding tree unit or coding unit) or a layer higher than a block layer of the current block (e.g., video parameter set, sequence parameter set, picture parameter set, adaption parameter set, picture header and / or slice header) from the bitstream.
[0226] In some embodiments, the filtering unit 708 may obtain an indication parameter from parsing unit 701. The indication parameter may indicate whether the reconstructed sample in the current block will be filtered. For example, the indication parameter may be a 1-bit flag. For example, the indication parameter may be a variable with a number of values indicating not only whether the reconstruct sample will be filtered but also which filter is used. When the variable is equal to 0, the reconstructed sample of the current block will not be filtered; otherwise, when the variable is equal to 1, the reconstructed sample of the current block is filtered using a filter with an index equal to the value of this variable.
[0227] In some embodiments, the filtering unit 708 may also obtain indication parameter from parsing unit 701, which indicates which color component will be filtered. Filtering unit 708 may choose to filter one or more of the luma and two chroma components.
[0228] The output of filtering unit 708 is a decoded picture or sub-picture, which is forwarded to DPB 709. DPB 709 outputs decoded pictures according to timing and controlling information. Pictures stored in DPB 709 may also be employed as reference for performing inter or intra prediction by prediction unit 702.
[0229] Decoder 700 may be a computing device with a processor and a storage medium recording a decoding program. When the processor reads and executes the decoding program, the decoder 700 reads an input video bitstream and generates corresponding decoded video.
[0230] Decoder 700 may be a computing device with one or more chips. The units, implemented as integrated circuits, on the chip are of similar functionalities with similar connections as well as data exchanges as the corresponding ones in FIG. 7.
[0231] FIG. 8 illustrates a block diagram of an exemplary source device 800, according to some embodiments of the present disclosure.
[0232] Referring to FIG. 8, acquisition unit 801 may acquire a video signal and forwards the video signal to encoder 802. Acquisition unit 801 can be a device containing one or more cameras (including depth cameras) . Acquisition unit 801 can be a device that partially or completely decodes a bitstream to get a video. Acquisition unit 801 may also contain one or more elements to capture an audio signal. An embodiment of encoder 802 is the encoder 200 that codes the video signal from acquisition unit 801 as its input video and generates a video bitstream. Encoder 802 may also contain one or more audio encoder to code the audio signal to generate an audio bitstream. Storage / sending unit 803 receives the video bitstream from encoder 802. Storage / sending unit 803 may also receive the audio bitstream from encoder 802 and encapsulate the video bitstream together with the audio bitstream to form a media file (e.g. ISO based media file format) or transport stream. Optionally, storage / sending unit 803 writes the media file or transport stream in a storage unit. e.g. hard disc, DVD disc, cloud, portable memory devices. Optionally, storage / sending unit 803 sends the bitstream to a transport network, for example, Internet, wireline networks, cellular networks, wireless local area networks, etc.
[0233] FIG. 9 illustrates a block diagram of an exemplary receiving device 900, according to some embodiments of the present disclosure.
[0234] Referring to FIG. 9, receiving unit 901 receives the media file or transport stream from networks or reads the media file or transport stream from a storage device. Receiving unit 901 separates the video bitstream and the audio bitstream from the media file or transport stream. Receiving unit 901 can also generate a new video bitstream by extracting the video bitstream. Receiving unit 901 may also generate a new audio bitstream by extracting the audio bitstream. Decoder 902 includes one or more video decoders, e.g. the decoder 700. Decoder 902 may also contain one or more audio decoders. Decoder 902 decodes the video bitstream and the audio bitstream from receiving unit 901 to get a decoded video and one or more decoded audio corresponding to one or multiple channels. Rendering unit 903 performs operations on the reconstructed video to make it suitable for displaying. Such operations may include one or more of the following operations to improve perceptual quality: denoising, synthesis, conversion of color space, upsampling, downsampling, etc. Rendering unit 903 may also performs operations on the decoded audio to improve the perceptual quality of the audio signal for displaying.
[0235] FIG. 10 illustrates a block diagram of a first exemplary communication system 1000, according to some embodiments of the present disclosure.
[0236] Referring to FIG. 10, source device 1001 may correspond to source device 800. The output of the storage / sending unit 803 is processed by storage medium / transport networks 1002 for storage or transport to the bitstream. Destination device 1003 may be receiving device 900. Receiving unit 901 gets the bitstream from storage medium / transport networks 1002. Receiving unit 901 may extract a new video bitstream from the media file or transport stream. Receiving unit 901 may also extract a new audio bitstream from the media file or transport stream.
[0237] FIG. 11 illustrates a block diagram of a second exemplary communication system 1100, according to some embodiments of the present disclosure.
[0238] Referring to FIG. 11, in some embodiments, second exemplary communication system 1100 may be a video codec system. The video codec system may include an encoding apparatus 1110 and a decoding apparatus 1120. The encoding apparatus 1110 may deliver encoded video and / or image information or data to the decoding apparatus 1120 in the form of a file or streaming via a digital storage medium or network.
[0239] The encoding apparatus 1110 according to an embodiment may include a video source generator 1111, an encoding unit 1112 (which can be encoder 200) and a transmitter 1113. The decoding apparatus 1120 according to an embodiment may include a receiver 1121, a decoding unit 1122 (which can be decoder 700) , and a renderer 1123. The encoding unit 1112 may be called a video / image encoding unit, and the decoding unit 1122 may be called a video / image decoding unit. The transmitter 1113 may be included in the encoding unit 1112. The receiver 1121 may be included in the decoding unit 1122. The renderer 1123 may include a display and the display may be configured as a separate device or an external component.
[0240] The video source generator 1111 may acquire a video / image through a process of capturing, synthesizing or generating the video / image. The video source generator 1111 may include a video / image capture device and / or a video / image generating device. The video / image capture device may include, for example, one or more cameras, video / image archives including previously captured video / images, and the like. The video / image generating device may include, for example, computers, tablets and smartphones, and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating related data.
[0241] The encoding unit 1112 may encode an input video / image. The encoding unit 1112 may perform a series of procedures such as prediction, transform, and quantization for compression and coding efficiency. The encoding unit 1112 may output encoded data (encoded video / image information) in the form of a bitstream.
[0242] The transmitter 1113 may transmit the encoded video / image information or data output in the form of a bitstream to the receiver 1121 of the decoding apparatus 1120 through a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage mediums such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. The transmitter 1113 may include an element for generating a media file through a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiver 1121 may extract / receive the bitstream from the storage medium or network and transmit the bitstream to the decoding unit 1122.
[0243] The decoding unit 1122 may decode the video / image by performing a series of procedures such as dequantization, inverse transform, and prediction corresponding to the operation of the encoding unit 1112.
[0244] The renderer 1123 may render the decoded video / image. The rendered video / image may be displayed through the display.
[0245] The embodiments described herein may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (ex. Information on instructions) or an algorithm may be stored in a digital storage medium.
[0246] In addition, the decoding apparatus and the encoding apparatus to which the present invention are applied may be included in a multimedia broadcasting transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video chat device, and a real time communication device such as video communication, a mobile streaming device, a storage medium, camcorder, a video-on-demand (VoD) service provider, an over the top video (OTT) device, an internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augment reality (AR) device, an image telephone video device, a vehicle terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc. ) and a medical video device, and the like, and may be used to process an image signal or data. For example, the OTT video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR) , and the like.
[0247] In addition, the processing method to which the present invention is applied may be produced in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment (s) of this document may also be stored in the computer-readable recording medium. The computer readable recording medium includes all kinds of storage devices and distributed storage devices in which computer readable data is stored. The computer readable recording medium may be, for example, a Blu-ray disc (BD) , a universal serial bus (USB) , a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes media embodied in the form of a carrier wave (ex. transmission over the Internet) . In addition, a bitstream generated by the encoding method may be stored in the computer-readable recording medium or transmitted through a wired or wireless communication network.
[0248] In addition, the embodiment of the present invention may be embodied as a computer program product based on a program code, and the program code may be executed on a computer by the embodiment (s) of the present invention. The program code may be stored on a carrier readable by a computer.
[0249] FIG. 12 illustrates a block diagram of a third exemplary communication system 1200, according to some embodiments of the present disclosure.
[0250] Referring to FIG. 12, third exemplary communication system 1200 may include a content streaming system. The content streaming system include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0251] The encoding server compresses content input from multimedia input devices such as a smartphone, a camera, a camcorder, etc. into digital data to generate a bitstream and transmit the bitstream to the streaming server. As another example, when the multimedia input devices such as smartphones, cameras, camcorders, etc. directly generate a bitstream, the encoding server may be omitted.
[0252] The bitstream may be generated by an encoding method or a bitstream generating method to which the present invention is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0253] The streaming server transmits the multimedia data to the user device based on a user’s request through the web server, and the web server serves as a medium for informing the user of a service. When the user requests a desired service from the web server, the web server delivers it to a streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server serves to control a command / response between devices in the content streaming system.
[0254] The streaming server may receive content from a media storage and / or an encoding server. For example, when the content is received from the encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a predetermined time.
[0255] Examples of the user device may include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA) , a portable multimedia player (PMP) , navigation, a slate PC, tablet PCs, ultra-books, wearable devices (ex. smartwatches, smart glasses, head mounted displays) , digital TVs, desktops computer, digital signage, and the like.
[0256] Each server in the content streaming system may be operated as a distributed server, in which case data received from each server may be distributed.
[0257] FIG. 17 illustrates a flow chart of an exemplary method 1700 of decoding, according to some embodiments of the present disclosure. Method 1700 may be performed by an apparatus, e.g., decoder 120 of decoding system 150, decoder 700, prediction unit 702, intra prediction unit 704, decoding apparatus 1120, decoding unit 1122, or any other suitable decoding systems. Method 1700 may include operations 1702-1708 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order than shown in FIG. 17.
[0258] Referring to FIG. 17, at 1702, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block may be determined based on at least one of a width or a height of the neighboring block.
[0259] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, decoding a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining the block-wise occurrence value based on the set of weights.
[0260] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0261] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may include solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0262] At 1704, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, the current block may be decoded based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0263] At 1706, in response to the neighboring block being coded with a non-fusion-based IPM, another block-wise occurrence value corresponding to the neighboring block may be determined.
[0264] In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, determining another block-wise occurrence value corresponding to the neighboring block.
[0265] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, the determining the another block-wise occurrence value corresponding to the neighboring block may include solving: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0266] In some implementations, the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0267] At 1708, in response to the neighboring block being coded with a non-fusion-based IPM, the current block based on the another block-wise occurrence value.
[0268] FIG. 18 illustrates a flow chart of an exemplary method 1800 of point cloud encoding, according to some embodiments of the present disclosure. Method 1800 may be performed by encoder 101 of encoding system 100, encoder 200, prediction unit 202, intra prediction unit 205, encoding apparatus 1110, encoding unit 1112, or any other suitable encoding systems. Method 1800 may include operations 1802-1808, as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order than shown in FIG. 18.
[0269] Referring to FIG. 18, at 1802, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block may be determined based on at least one of a width or a height of the neighboring block.
[0270] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, encoding a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining the block-wise occurrence value based on the set of weights.
[0271] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0272] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may include solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0273] At 1804, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, the current block may be encoded based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0274] At 1806, in response to the neighboring block being coded with a non-fusion-based IPM, another block-wise occurrence value corresponding to the neighboring block may be determined.
[0275] In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, determining another block-wise occurrence value corresponding to the neighboring block.
[0276] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, the determining the another block-wise occurrence value corresponding to the neighboring block may include solving: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0277] In some implementations, the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0278] At 1808, in response to the neighboring block being coded with a non-fusion-based IPM, the current block based on the another block-wise occurrence value.
[0279] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a processor, such as processor 102 in FIGs. 1A and 1B. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, HDD, such as magnetic disk storage or other magnetic storage devices, Flash drive, SSD, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system, such as a mobile device or a computer. Disk and disc, as used herein, includes CD, laser disc, optical disc, digital video disc (DVD) , and floppy disk where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0280] According to one aspect of the present disclosure, a method of decoding is provided. The method may include, in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The method may include, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0281] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, decoding, by the processor, a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, the block-wise occurrence value based on the set of weights.
[0282] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0283] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may include solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0284] In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, determining, by the processor, another block-wise occurrence value corresponding to the neighboring block. In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, decoding, by the processor, the current block based on the another block-wise occurrence value.
[0285] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, the determining, by the processor, the another block-wise occurrence value corresponding to the neighboring block may include solving: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0286] In some implementations, the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0287] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0288] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, decode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine the block-wise occurrence value based on the set of weights.
[0289] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0290] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may cause the processor to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0291] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a non-fusion-based IPM, determine another block-wise occurrence value corresponding to the neighboring block. In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a non-fusion-based IPM, decode the current block based on the another block-wise occurrence value.
[0292] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the memory storing instructions, which when executed by the processor, may cause the processor to solve: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0293] In some implementations the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0294] According to another aspect of the present disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0295] According to a further aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a decoder is provided. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, decode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0296] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a fusion-based IPM, decode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a fusion-based IPM, determine the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a fusion-based IPM, determine a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a fusion-based IPM, determine the block-wise occurrence value based on the set of weights.
[0297] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0298] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0299] In some implementations, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a non-fusion-based IPM, determine another block-wise occurrence value corresponding to the neighboring block. In some implementations, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to, in response to the neighboring block being coded with a non-fusion-based IPM, decode the current block based on the another block-wise occurrence value.
[0300] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to solve: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0301] In some implementations the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0302] According to one aspect of the present disclosure, a method of encoding is provided. The method may include, in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The method may include, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0303] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, encoding, by the processor, a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may further include, in response to the neighboring block being coded with a fusion-based IPM, determining, by the processor, the block-wise occurrence value based on the set of weights.
[0304] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0305] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block may include solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0306] In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, determining, by the processor, another block-wise occurrence value corresponding to the neighboring block. In some implementations, the method may further include, in response to the neighboring block being coded with a non-fusion-based IPM, encoding, by the processor, the current block based on the another block-wise occurrence value.
[0307] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, the determining, by the processor, the another block-wise occurrence value corresponding to the neighboring block may include solving: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0308] In some implementations, the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0309] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0310] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, encode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a fusion-based IPM, determine the block-wise occurrence value based on the set of weights.
[0311] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0312] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, may cause the processor to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0313] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a non-fusion-based IPM, determine another block-wise occurrence value corresponding to the neighboring block. In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to, in response to the neighboring block being coded with a non-fusion-based IPM, encode the current block based on the another block-wise occurrence value.
[0314] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the memory storing instructions, which when executed by the processor, may cause the processor to solve: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0315] In some implementations the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0316] According to another aspect of the present disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The memory storing instructions, which when executed by the processor, may cause the processor to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0317] According to a further aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for an encoder is provided. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to, in response to an IPM corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block, encode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.
[0318] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a fusion-based IPM, encode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a fusion-based IPM, determine the minimum block size for the video sequence based on the syntax element. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a fusion-based IPM, determine a set of weights for determining the block-wise occurrence value based on the minimum block size. In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a fusion-based IPM, determine the block-wise occurrence value based on the set of weights.
[0319] In some implementations, the fusion-based IPM may include DIMD, TIMD, SGPM, OBIC, MIP, intraTMP, IBC, or EIP.
[0320] In some implementations, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) , where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.
[0321] In some implementations, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a non-fusion-based IPM, determine another block-wise occurrence value corresponding to the neighboring block. In some implementations, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to, in response to the neighboring block being coded with a non-fusion-based IPM, encode the current block based on the another block-wise occurrence value.
[0322] In some implementations, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to solve: HoC [IPM] += uiWidth *uiHeight, where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.
[0323] In some implementations the non-fusion-based IPM may include planar mode, vertical mode, or horizontal mode.
[0324] According to yet another aspect of the present disclosure, a method of transmitting a bitstream is provided. The method may include generating, by a processor, the bitstream according to one or more of the operations described herein. The method may include transmitting, by the processor, the bitstream.
[0325] According to still another aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated using one or more operations described herein.
[0326] The foregoing description of the embodiments will so reveal the general nature of the present disclosure that others can, by applying knowledge within the skill of the art, readily modify and / or adapt for various applications such embodiments, without undue experimentation, without departing from the general concept of the present disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description and not of limitation, such that the terminology or phraseology of the present specification is to be interpreted by the skilled artisan in light of the teachings and guidance.
[0327] Embodiments of the present disclosure have been described above with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.
[0328] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor (s) , and thus, are not intended to limit the present disclosure and the appended claims in any way.
[0329] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and without limitation. Accordingly, the functional blocks, modules, and steps may be reordered or combined in different ways than in the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.
[0330] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1.A method of decoding, comprising:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; anddecoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.2.The method of claim 1, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block further comprises:in response to the neighboring block being coded with a fusion-based IPM,decoding, by the processor, a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determining, by the processor, the minimum block size for the video sequence based on the syntax element;determining, by the processor, a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermining, by the processor, the block-wise occurrence value based on the set of weights.3.The method of claim 2, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.4.The method of claim 1, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block comprises solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.5.The method of claim 1, further comprising:in response to the neighboring block being coded with a non-fusion-based IPM,determining, by the processor, another block-wise occurrence value corresponding to the neighboring block; anddecoding, by the processor, the current block based on the another block-wise occurrence value.6.The method of claim 5, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, the determining, by the processor, the another block-wise occurrence value corresponding to the neighboring block comprises solving: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.7.The method of claim 5, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.8.A decoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; anddecode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.9.The decoder of claim 8, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, further cause the processor to:in response to the neighboring block being coded with a fusion-based IPM,decode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determine the minimum block size for the video sequence based on the syntax element;determine a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermine the block-wise occurrence value based on the set of weights.10.The decoder of claim 9, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.11.The decoder of claim 8, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, cause the processor to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.12.The decoder of claim 8, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:in response to the neighboring block being coded with a non-fusion-based IPM,determine another block-wise occurrence value corresponding to the neighboring block; anddecode the current block based on the another block-wise occurrence value.13.The decoder of claim 12, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the memory storing instructions, which when executed by the processor, cause the processor to solve: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.14.The decoder of claim 12, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.15.An apparatus for decoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; anddecode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.16.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a decoder, cause the processor of the decoder to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; anddecode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.17.The non-transitory computer-readable medium of claim 16, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, further cause the processor of the decoder to:in response to the neighboring block being coded with a fusion-based IPM,decode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determine the minimum block size for the video sequence based on the syntax element;determine a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermine the block-wise occurrence value based on the set of weights.18.The non-transitory computer-readable medium of claim 17, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.19.The non-transitory computer-readable medium of claim 16, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.20.The non-transitory computer-readable medium of claim 16, wherein the instructions, which when executed by the processor of the decoder, further cause the processor of the decoder to:in response to the neighboring block being coded with a non-fusion-based IPM,determine another block-wise occurrence value corresponding to the neighboring block; anddecode the current block based on the another block-wise occurrence value.21.The non-transitory computer-readable medium of claim 20, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to solve: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.22.The non-transitory computer-readable medium of claim 20, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.23.A method of encoding, comprising:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determining, by a processor, a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; andencoding, by the processor, the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.24.The method of claim 23, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block further comprises:in response to the neighboring block being coded with a fusion-based IPM,encoding, by the processor, a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determining, by the processor, the minimum block size for the video sequence based on the syntax element;determining, by the processor, a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermining, by the processor, the block-wise occurrence value based on the set of weights.25.The method of claim 24, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.26.The method of claim 23, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, the determining, by the processor, the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block comprises solving: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.27.The method of claim 23, further comprising:in response to the neighboring block being coded with a non-fusion-based IPM,determining, by the processor, another block-wise occurrence value corresponding to the neighboring block; andencoding, by the processor, the current block based on the another block-wise occurrence value.28.The method of claim 27, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, the determining, by the processor, the another block-wise occurrence value corresponding to the neighboring block comprises solving: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.29.The method of claim 27, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.30.An encoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; andencode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.31.The encoder of claim 30, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, further cause the processor to:in response to the neighboring block being coded with a fusion-based IPM,encode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determine the minimum block size for the video sequence based on the syntax element;determine a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermine the block-wise occurrence value based on the set of weights.32.The encoder of claim 31, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.33.The encoder of claim 30, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the memory storing instructions, which when executed by the processor, cause the processor to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.34.The encoder of claim 30, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:in response to the neighboring block being coded with a non-fusion-based IPM,determine another block-wise occurrence value corresponding to the neighboring block; andencode the current block based on the another block-wise occurrence value.35.The encoder of claim 34, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the memory storing instructions, which when executed by the processor, cause the processor to solve: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.36.The encoder of claim 34, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.37.An apparatus for encoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; andencode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.38.A non-transitory computer-readable medium storing instructions, which when executed by a processor of an encoder, cause the processor of the encoder to:in response to an intra prediction mode (IPM) corresponding to a current block being determined based on an occurrence of an intra coding mode of a neighboring block,determine a block-wise occurrence value of an IPM corresponding to a neighboring block of the current block based on at least one of a width or a height of the neighboring block; andencode the current block based on the block-wise occurrence value of the IPM corresponding to the neighboring block.39.The non-transitory computer-readable medium of claim 38, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, further cause the processor of the encoder to:in response to the neighboring block being coded with a fusion-based IPM,encode a syntax element to determine a minimum block size for a video sequence corresponding to the neighboring block and the current block;determine the minimum block size for the video sequence based on the syntax element;determine a set of weights for determining the block-wise occurrence value based on the minimum block size; anddetermine the block-wise occurrence value based on the set of weights.40.The non-transitory computer-readable medium of claim 39, wherein the fusion-based IPM comprises decoder-side intra mode derivation (DIMD) , template-based intra mode derivation (TIMD) , spatial geometric partitioning mode (SGPM) , occurrence-based intra coding (OBIC) , matrix-based intra prediction (MIP) , intra template matching prediction (intraTMP) , intra block copying (IBC) , or EIP.41.The non-transitory computer-readable medium of claim 38, wherein, in response to the IPM corresponding to the current block being determined based on the occurrence of the intra coding mode of the neighboring block, to determine the block-wise occurrence value of the IPM corresponding to the neighboring block of the current block based on at least one of the width or the height of the neighboring block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to solve: HoC [IPM] += (uiWidth>>shift1) * (uiHeight>>shift2) ,where HoC [IPM] is the block-wise occurrence value, uiWidth is the width of the neighboring block, uiHeight is the height of the neighboring block, shift1 is a first weight of a set of weights, and shift2 is a second weight of the set of weights.42.The non-transitory computer-readable medium of claim 38, wherein the instructions, which when executed by the processor of the encoder, further cause the processor of the encoder to:in response to the neighboring block being coded with a non-fusion-based IPM,determine another block-wise occurrence value corresponding to the neighboring block; andencode the current block based on the another block-wise occurrence value.43.The non-transitory computer-readable medium of claim 42, wherein, in response to the neighboring block being coded with the non-fusion-based IPM, to determine the another block-wise occurrence value corresponding to the neighboring block, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to solve: HoC [IPM] += uiWidth *uiHeight,where HoC [IPM] is the another block-wise occurrence value, uiWidth is the width of the neighboring block, and uiHeight is the height of the neighboring block.44.The non-transitory computer-readable medium of claim 42, wherein the non-fusion-based IPM comprises planar mode, vertical mode, or horizontal mode.45.A non-transitory computer-readable medium storing a bitstream, the bitstream being generated according to one or more of claims 23-29.46.A method of transmitting a bitstream, comprising;generating, by a processor, a bitstream according to one or more of claims 23-29; andtransmitting, by the processor, the bitstream.
Citation Information
Patent Citations
An encoder, a decoder and corresponding methods of intra prediction
CN113748677A
Intra prediction method and intra prediction apparatus based on coding information
KR1020140081933A
Method and apparatus for deriving intra-prediction mode
US20190281290A1
Adaptive construction of most probable modes candidate list for video data encoding and decoding
WO2020056779A1