Method and device for chroma coding and decoding
By introducing the Block Vector Guided Convolution Cross-Component Intra-Prediction Model (BVG-CCCM), the problem of low chroma block prediction efficiency in existing technologies is solved, achieving more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in handling chroma block prediction, especially when using a single tree to divide co-occurring luminance blocks, as they lack effective prediction modes, resulting in low encoding efficiency.
The prediction of chroma blocks is performed using the Block Vector Guided-Convolutional Cross-Component Intra-Prediction Model (BVG-CCCM) mode. The BVG-CCCM flag is obtained by parsing the bitstream, and the appropriate prediction mode is selected according to the enabling conditions.
It improves the prediction efficiency of chroma blocks, thereby enhancing the overall coding efficiency and quality of video coding.
Smart Images

Figure CN121986478A_ABST
Abstract
Description
[0001] Cross-references This application claims priority to U.S. Provisional Application No. 63 / 542,674 entitled “CHROMA CODING”, filed on October 5, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0002] The embodiments disclosed herein relate to video encoding and decoding. Background Technology
[0003] Digital video has become mainstream and is used in a wide range of applications, including digital television, video telephony, and teleconferencing. These digital video applications are feasible due to advancements in computing and communication technologies, as well as efficient video codecs. Various video codecs can be used to compress video data, allowing the use of one or more video codec standards to perform encoding and decoding of video data. Exemplary video codec standards may include, but are not limited to, Universal Video Codec (H.266 / VVC), High Efficiency Video Codec (H.265 / HEVC), Advanced Video Codec (H.264 / AVC), Moving Picture Experts Group (MPEG) codec, Enhanced Video Codec Model (ECM), etc. Summary of the Invention
[0004] According to one aspect of this disclosure, a method for decoding by a decoder is provided. The method may include: in response to a single-tree partitioned co-occurrence luma block being encoded using intraTMP or IBC, a processor determining that the corresponding chroma block meets the enable condition for prediction using a block vector guided (BVG)-convolutional cross-component intra prediction model (CCCM) mode. The method may include the processor parsing the bitstream to obtain a BVG-CCCM flag. The method may also include: in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block.
[0005] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0006] According to another aspect of this disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided for storing instructions for a decoder. When executed by a decoder's processor, the instructions cause the decoder's processor to determine, in response to the single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. When executed by the decoder's processor, the instructions can cause the decoder's processor to parse the bitstream to obtain a BVG-CCCM flag. When executed by the decoder's processor, the instructions can cause the decoder's processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0008] According to one aspect of this disclosure, a method for encoding by an encoder is provided. The method may include: in response to a single-tree partitioning of a co-positional luma block being encoded using intraTMP or IBC, a processor determining that the corresponding chroma block meets the enable conditions for prediction using a BVG-CCCM mode. The method may include encoding a BVG-CCCM flag by the processor. The method may also include: in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block.
[0009] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable condition for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to encode a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0010] According to another aspect of this disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of a co-occurrence luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using a BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to encode a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided for storing instructions for an encoder. When executed by a processor of the encoder, the instructions can cause the encoder processor to determine, in response to the single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable condition for prediction using the BVG-CCCM mode. When executed by the encoder processor, the instructions can also cause the encoder processor to encode a BVG-CCCM flag. When executed by the encoder processor, the instructions can also cause the encoder processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0012] According to another aspect of this disclosure, a non-transitory computer-readable medium for storing a bitstream generated according to one or more of the operations described herein.
[0013] These illustrative embodiments are mentioned not to limit or restrict this disclosure, but to provide examples to aid in understanding it. Further embodiments are described in the detailed description, and further description is provided therein. Attached Figure Description
[0014] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and enable those skilled in the art to make and use the present disclosure.
[0015] Figure 1 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.
[0016] Figure 2 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.
[0017] Figure 3 Some embodiments according to this disclosure are shown. Figure 1 A detailed block diagram of an exemplary encoder in an encoding system.
[0018] Figure 4 Some embodiments according to this disclosure are shown. Figure 2 A detailed block diagram of an exemplary decoder in a decoding system.
[0019] Figure 5 Exemplary images showing a division into coding tree units (CTUs) according to some embodiments of the present disclosure are shown.
[0020] Figure 6An exemplary CTU, divided into coding units (CUs), is shown according to some embodiments of the present disclosure.
[0021] Figure 7 A schematic visualization of the current CU block and reconstructed samples that are spatially adjacent and non-adjacent to the current block is shown according to some embodiments of the present disclosure.
[0022] Figure 8 A schematic visualization of the angular pattern of a VVC according to some embodiments of the present disclosure is shown.
[0023] Figure 9A The diagram illustrates slice partitioning for intra-frame prediction according to some embodiments of the present disclosure.
[0024] Figure 9B The illustration shows a tile partitioning for intra-frame prediction according to some embodiments of the present disclosure.
[0025] Figure 9C A diagram illustrating wavefront parallel processing for intra-frame prediction according to some embodiments of the present disclosure is shown.
[0026] Figure 10A A diagram illustrating intra-block copy (IBC) according to some embodiments of the present disclosure is shown.
[0027] Figure 10B A diagram illustrating intra-template matching prediction (intraTMP) according to some embodiments of the present disclosure is shown.
[0028] Figure 10C A diagram illustrating an extended search region for intraTMP according to some embodiments of the present disclosure is shown.
[0029] Figure 11 A diagram showing subpixel positions for intraTMP according to some embodiments of the present disclosure is illustrated.
[0030] Figure 12 A diagram illustrating a direct block vector (DBV) pattern according to some embodiments of the present disclosure is shown.
[0031] Figure 13 A graph showing DBV prediction according to some embodiments of this disclosure is illustrated.
[0032] Figure 14 A diagram illustrating the spatial components of a convolutional cross-component intra-prediction model (CCCM) filter according to some embodiments of the present disclosure is provided.
[0033] Figure 15 A diagram showing a reference region for calculating a CCCM filter according to some embodiments of the present disclosure is illustrated.
[0034] Figure 16 A reference region for computing a block vector guided CCCM (BVG-CCCM) filter is shown according to some embodiments of the present disclosure.
[0035] Figure 17 A flowchart of a decoding method according to some embodiments of the present disclosure is shown.
[0036] Figure 18 A flowchart of an encoding method according to some embodiments of the present disclosure is shown.
[0037] Embodiments of this disclosure will be described with reference to the accompanying drawings. Detailed Implementation
[0038] While some configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of this disclosure. It will be apparent to those skilled in the art that this disclosure can also be used in a variety of other applications.
[0039] It should be noted that references to "one embodiment," "embodiment," "example embodiment," "some embodiments," "certain embodiments," etc., in the specification indicate that the described embodiments may include specific features, structures, or characteristics, but each embodiment may not necessarily include said specific features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in connection with an embodiment, whether explicitly described or not, implementing such a feature, structure, or characteristic in conjunction with other embodiments will be within the knowledge of those skilled in the art.
[0040] Generally, terms can be understood, at least in part, from their usage in context. For example, the term "one or more," as used herein, can be used, at least in part, to describe any feature, structure, or characteristic in a singular sense, or in a plural sense, to describe a combination of multiple features, structures, or characteristics. Similarly, terms such as "an," "a," or "the" can again be understood to convey either a singular or a plural usage, at least in part, depending on the context. Furthermore, the term "based on" can be understood not necessarily to convey an exclusive set of factors, but rather to allow for the presence of additional factors that are not necessarily explicitly described again, at least in part, depending on the context.
[0041] Various aspects of a video coding system will now be described with reference to various apparatuses and methods. These apparatuses and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively, “elements”). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.
[0042] The techniques described herein can be used in a variety of video encoding and decoding applications. As described herein, video encoding and decoding includes both encoding and decoding of video. Video encoding and decoding can be performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, reconstruction, etc., can be performed on encoded blocks, transform blocks, or prediction blocks. As described herein, the block to be encoded / decoded will be referred to as the “current block.” For example, the current block can represent an encoded block, transform block, or prediction block according to the current encoding / decoding process. Furthermore, it should be understood that the term “unit” as used in this disclosure refers to a basic unit for performing a particular encoding / decoding process, and the term “block” refers to a sample array of a predetermined size. Unless otherwise stated, “block” and “unit” are used interchangeably.
[0043] Figure 1 A block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure is shown. Figure 2 A block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure is shown. Each system 100 or 200 can be applied to or integrated into a variety of systems and devices capable of data processing, such as computers and wireless communication devices. For example, system 100 or 200 can be all or part of a mobile phone, desktop computer, laptop computer, tablet computer, in-vehicle computer, game console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. Figure 7 and Figure 8 As shown, system 100 or 200 may include processor 102, memory 104, and interface 106. These components are shown as being connected to each other via a bus, but other connection types are also permitted. It should be understood that system 100 or 200 may include any other suitable components for performing the functions described herein.
[0044] Processor 102 may include microprocessors such as graphics processing unit (GPU), image signal processor (ISP), central processing unit (CPU), digital signal processor (DSP), tensor processing unit (TPU), vision processing unit (VPU), neural processing unit (NPU), synergistic processing unit (SPU), or physics processing unit (PPU), microcontroller unit (MCU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device (PLD), state machine, gated logic, discrete hardware circuitry, and other suitable hardware configured to perform the various functions described throughout this disclosure. Figure 7 and Figure 8 Only one processor is shown, but it should be understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 can execute software. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, application programs, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description languages, or others. Software may include computer instructions written in interpreted languages, compiled languages, or machine code. Other techniques used to indicate hardware are also permitted under the broad category of software.
[0045] Memory 104 can broadly include both memory (also known as main / system memory) and storage devices (also known as auxiliary memory). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferro-electric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage devices, hard disk drive (HDD) (such as disk storage devices or other magnetic storage devices), flash memory drive, solid-state drive (SSD), or any other medium that can be used to carry or store desired program code in the form of instructions that can be accessed and executed by processor 102. More broadly, memory 104 can be embodied in any computer-readable medium (such as non-transitory computer-readable media). Although in Figure 1 and Figure 2 Only one memory is shown, but it should be understood that multiple memories may be included.
[0046] Interface 106 can broadly include data interfaces and communication interfaces, the communication interface being configured to receive and transmit signals during the reception and transmission of information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although in Figure 7 and Figure 8 Only one memory is shown, but it should be understood that it may include multiple interfaces.
[0047] Processor 102, memory 104, and interface 106 may be implemented in various forms within system 100 or 200 for performing video encoding and decoding functions. In some embodiments, processor 102, memory 104, and interface 106 of system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chip (SoCs). In one example, processor 102, memory 104, and interface 106 may be integrated on an application processor (AP) SoC that handles application processing within an operating system (OS) environment, including running video encoding and decoding applications. In another example, processor 102, memory 104, and interface 106 may be integrated on a dedicated processor chip for video encoding, such as a GPU or ISP chip dedicated to image and video processing within a real-time operating system (RTOS).
[0048] like Figure 1 As shown, in the encoding system 100, the processor 102 may include one or more modules, such as the encoder 101. Although Figure 1 Encoder 101 is shown within a processor 102; however, it should be understood that encoder 101 may include one or more submodules that may be implemented on different processors, either close to or far from each other. Encoder 101 (and any corresponding submodules or subunits) may be a hardware unit (e.g., a portion of an integrated circuit) of processor 102, designed for use with other components or software units implemented by processor 102 by executing at least a portion of a program (e.g., instructions). The instructions of the program may be stored on a computer-readable medium such as memory 104, and when executed by processor 102, may perform processes having one or more functions related to video coding, such as image segmentation, inter-frame prediction, intra-frame prediction, transform, quantization, filtering, entropy coding, etc., as described in detail below.
[0049] Similarly, such as Figure 2 As shown, in the decoding system 200, the processor 102 may include one or more modules, such as the decoder 201. Although Figure 2Decoder 201 is shown within a processor 102; however, it should be understood that decoder 201 may include one or more submodules that may be implemented on different processors, either close to or far from each other. Decoder 201 (and any corresponding submodules or subunits) may be a hardware unit (e.g., a portion of an integrated circuit) of processor 102, designed for use with other components or software units implemented by processor 102 by executing at least a portion of a program (e.g., instructions). The instructions of the program may be stored on a computer-readable medium such as memory 104, and when executed by processor 102, may perform processes having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transform, inter-frame prediction, intra-frame prediction, filtering, as described in detail below.
[0050] Figure 3 Some embodiments according to this disclosure are shown. Figure 1 A detailed block diagram of an exemplary encoder 101 in the encoding system 100. (See attached diagram.) Figure 3 As shown, encoder 101 may include a partitioning module 302, an inter-frame prediction module 304, an intra-frame prediction module 306, a transform module 308, a quantization module 310, an inverse quantization module 312, an inverse transform module 314, a filter module 316, a buffer module 318, and an encoding module 320. It should be understood that... Figure 3 Each element shown is illustrated independently to represent a distinct feature function within the video encoder, and does not imply that each component is formed by a separate hardware or software configuration unit. That is, for ease of illustration, elements are listed as components, and at least two elements can be combined to form a single element, or a single element can be divided into multiple elements to perform its function. It should also be understood that some elements are not essential for performing the functions described in this disclosure, but may be optional elements used to improve performance. It should also be understood that these elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on encoder 101.
[0051] The partitioning module 302 can be configured to partition an input image of a video into at least one processing unit. The image can be a frame or field of a video. In some embodiments, the image includes a monochrome luminance sample array, or a luminance sample array and two corresponding chrominance sample arrays. In this case, the processing unit can be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The partitioning module 302 can partition the image into a combination of multiple coding units, prediction units, and transform units, and encode the image by selecting the combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function).
[0052] Similar to H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal predictive coding scheme. Figure 5 As shown, during encoding, the input image 500 is first divided into square blocks – CTU 502 – by the partitioning module 302. For example, CTU 502 can be a block of 128 × 128 pixels. Figure 6 As shown, each CTU 502 in the input image 500 can be divided into one or more CUs 602 by the partitioning module 302, which can be used for prediction and transformation. Unlike H.265 / HEVC, in H.266 / VVC, CUs 602 can be rectangular or square and can be encoded without further partitioning into prediction units or transformation units. For example, as Figure 6 As shown, dividing CTU 502 into CU 602 can include quadtree partitioning (indicated by solid lines), binary tree partitioning (indicated by dashed lines), and ternary tree partitioning (indicated by dotted lines). According to some embodiments, each CU 602 can be the same size as its root CTU, or a subdivision of the root CTU 502, as small as a 4×4 block.
[0053] refer to Figure 4Inter-frame prediction module 304 can be configured to perform inter-frame prediction on prediction units, and intra-frame prediction module 306 can be configured to perform intra-frame prediction on prediction units. It can be determined whether to use inter-frame prediction or perform intra-frame prediction for a prediction unit, and specific information (e.g., intra-frame prediction mode, motion vectors, reference image, etc.) is determined based on each prediction method. In this case, the processing unit used to perform prediction can be different from the processing unit used to determine the prediction method and specific content. For example, the prediction method and prediction mode can be determined in the prediction unit, and prediction can be performed in the transform unit. The residual coefficients in the residual block between the generated prediction block and the original block can be input to the transform module 308. Additionally, prediction mode information, motion vector information, etc., used for prediction can be encoded into the bitstream by the encoding module 320 along with the residual coefficients or quantization level. It should be understood that in some encoding modes, the original block can be encoded as is without generating a prediction block through prediction modules 304 or 306. It should also be understood that in some encoding modes, prediction, transform, and / or quantization can be skipped.
[0054] In some embodiments, the inter-frame prediction module 304 may predict prediction units based on information about at least one image preceding or following the current image, and in some cases, the inter-frame prediction module 304 may predict prediction units based on information about the encoded partitions in the current image. The inter-frame prediction module 304 may include sub-modules such as a reference image interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference image interpolation module may receive reference image information from the buffer module 318 and generate pixel information of an integer number or fewer pixels from the reference image. In the case of luminance pixels, an 8-tap interpolation filter based on discrete cosine transform (DCT) with varying filter coefficients may be used to generate pixel information of an integer number of pixels or fewer pixels in units of 1 / 4 pixels. In the case of chrominance signals, a 4-tap interpolation filter based on DCT with varying filter coefficients may be used to generate pixel information of an integer number of pixels or fewer pixels in units of 1 / 8 pixels. The motion prediction module may perform motion prediction based on a reference image interpolated by the reference image interpolation portion. Various methods, such as the full search-based block matching algorithm (FBMA), three-step search (TSS), and the new three-step search algorithm (NTS), can be used to compute motion vectors. Based on interpolated pixels, motion vectors can have values in units of 1 / 2, 1 / 4, or 1 / 16 pixels or whole pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods, such as skip methods, merge methods, advanced motion vector prediction (AMVP) methods, and intra-block copying methods, can be used as motion prediction methods.
[0055] Still referencing Figure 3In some embodiments, the intra-prediction module 306 may generate prediction units based on information about reference pixels surrounding the current block (which is pixel information in the current image). The reference pixels may be located in reference lines that are not adjacent to the current block. When a block in the neighborhood of the current prediction unit has already undergone inter-frame prediction and therefore the reference pixel is a pixel that has undergone inter-frame prediction, the reference pixel information of the block in the neighborhood that has undergone intra-frame prediction can be replaced by a reference pixel included in the block that has undergone inter-frame prediction. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information. In intra-frame prediction, the prediction mode may have an angular prediction mode that uses reference pixel information according to the prediction direction and a non-angular prediction mode that does not use direction information when performing prediction. The mode used to predict luminance information may be different from the mode used to predict chromatic difference information, and chromatic difference information can be predicted using intra-frame prediction mode information for predicting luminance information or luminance signal information for predicting luminance information. If the size of the prediction unit is the same as the size of the transform unit when performing intra-prediction, intra-prediction can be performed based on the left, top-left, and top pixels of the prediction unit. However, if the size of the prediction unit is different from the size of the transform unit when performing intra-prediction, intra-prediction can be performed using reference pixels based on the transform unit.
[0056] Intra-prediction methods can generate prediction blocks after applying an adaptive intra-smoothing (AIS) filter to a reference pixel based on the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform intra-prediction, the intra-prediction mode of the current prediction unit can be predicted based on the intra-prediction modes of prediction units existing in the neighborhood of the current prediction unit. When using mode information predicted from neighboring prediction units to predict the prediction mode of the current prediction unit, if the intra-prediction mode of the current prediction unit is the same as that of the neighboring prediction units, predetermined flag information can be used to send information indicating that the prediction mode of the current prediction unit is the same as that of the neighboring prediction units; and if the prediction mode of the current prediction unit and that of the neighboring prediction units are different from each other, additional flag information can be used to encode the prediction mode information of the current block.
[0057] like Figure 3 As shown, a residual block can be generated, which includes prediction units that have been used to perform predictions based on prediction units generated by prediction modules 304 or 306, and residual coefficient information (also referred to herein as "residual"), which is the difference between the prediction unit and the original block. The generated residual block can be input into transform module 308. Additional details regarding the residuals and transforms used for video coding will now be provided.
[0058] In hybrid video coding systems, redundancy in the video signal is first exploited by applying inter-frame or intra-frame prediction tools to each control unit (CU). The difference between the original sample of a CU and its predicted block is often referred to as the residual. Even after prediction, the residual can still be highly spatially correlated. While conditional entropy coding can capture some spatial dependencies between adjacent samples, it is computationally impractical to formulate an entropy coding statistical model that can fully utilize the spatial correlation in the residual. In contrast, transform coding is a practical and efficient method for spatial decorrelation of the residual.
[0059] For example, the transformation module 308 can use an integer version of the two-dimensional discrete cosine transform (DCT) to transform the residuals, which can be applied separably in the horizontal and vertical directions. For an MxN residual sample block (where M is the width of the block and N is the height of the block), the transformation module 308 can obtain the transformation coefficients by applying the MxM DCT to each row to generate intermediate transformation coefficients, and then applying the NxN DCT to each column of the intermediate transformation coefficients.
[0060] For an intra-coded CU (also referred to herein as an "intra-CU"), spatially adjacent reconstructed samples are used to predict the current block, and the intra-prediction mode is signaled once for the entire CU. Each CU consists of one or more coded blocks (CBs) corresponding to the color components of the video sequence. For example, consumer video typically uses a 4:2:0 chroma format, in which case each CU consists of a luma CB and two chroma CBs with a quarter sample of the luma CB. Intra-prediction and transform coding are performed at the prediction block (PB) and transform block (TB) levels, respectively. Each CB consists of a single TB, except in the case of intra-fractional subdivision (ISP) mode and implicit splitting. For luma CBs, the maximum side length of a TB is 64, and the minimum side length is 4. Furthermore, the luma TB is further specified as a W×H rectangular block of width W and height H, where W, H ∈ {4, 8, 16, 32, 64}. For chroma CBs, the maximum TB side length is 32, and the chroma TB is a W×H rectangular block of width W and height H. Here, W, H ∈ {2, 4, 8, 16, 32}, but blocks of shape 2×H and 4×2 are excluded in order to address memory architecture and throughput requirements.
[0061] Figure 7 The illustration shows a schematic visualization 700 of the current CU block 702 and reconstructed samples that are spatially adjacent and non-adjacent to the current block, according to some aspects of this disclosure. Figure 7 In the text, the numbers 0, 1, 2, ... indicate the pixel line index associated with the current CU block 702.
[0062] In VVC, reference samples obtained from reconstructed samples of neighboring blocks are used to generate intra-prediction samples for the current block. For a W×H block, the reference sample is spatially adjacent to the current block and consists of a vertical line of 2.H reconstructed samples extending downwards to the left of the block, an upper-left reconstructed sample, and a horizontal line of 2.W reconstructed samples extending to the right above the current block. This "L"-shaped set of samples may be referred to as a "reference line" in this disclosure. The reference line directly adjacent to the current CU block 702 is shown as... Figure 7 The line has index 0.
[0063] Similar to AVC and HEVC, VVC also supports angular intra-prediction modes. Angular intra-prediction is a type of directional intra-prediction method. Compared to HEVC, VVC's angular intra-prediction is modified by increasing prediction accuracy and adapting to the new partitioning frame. Increased prediction accuracy is achieved by increasing the number of angular prediction directions and using more precise interpolation filters, while adaptation to the new partitioning frame is achieved by introducing wide-angle intra-prediction modes. In VVC, the number of directional modes available for a given block increases from 33 HEVC directions to 65 directions. Figure 8 The corner mode 800 of VVC is described in the text.
[0064] Orientations with even indices between 2 and 66 are equivalent to the orientations of the angle modes supported in HEVC. For square-shaped blocks, an equal number of angle modes are assigned to the top and left sides of the block. On the other hand, rectangular-shaped intra blocks (which do not exist in HEVC) are a core part of VVC's partitioning scheme, where additional intra-prediction orientations are assigned to the longer side of the block. These additional orientations assigned along the longer side are called Wide-Angle IntraPrediction (WAIP) modes because they correspond to prediction orientations with an angle greater than 45° relative to the horizontal or vertical orientation. Figure 8 As shown, a WAIP pattern for a given pattern index is defined by mapping the original orientation pattern to a pattern with the opposite orientation and an index offset of 1. For a given rectangular block, the aspect ratio (e.g., the ratio of width to height) is used to determine which angle patterns will be replaced by the corresponding wide-angle pattern.
[0065] For a square block in VVC, each pair of horizontally or vertically adjacent predicted samples is predicted from a pair of adjacent reference samples. In contrast, WAIP extends the angular range of directional prediction to more than 45°, so for a coded block predicted using the WAIP pattern, adjacent predicted samples can be predicted from non-adjacent reference samples.
[0066] In addition to the directly adjacent lines of adjacent samples, Figure 7One of the two non-adjacent reference lines (line 1 and line 2) depicted can include input samples for intra-frame prediction in VVC. For ECM, more non-adjacent reference lines can be used. The use of adjacent and non-adjacent reference samples is called multiple reference line (MRL) prediction.
[0067] Intra-frame modes available for MRL are DC mode and angle prediction mode. However, not all of these modes can be combined with MRL for a given block. MRL modes are always coupled with a mode from the Most Probable Mode (MPM) list in VVC. This coupling means that if non-adjacent reference lines are used, the intra-frame prediction mode is one of the MPMs. This design of MPM-based MRL prediction modes was inspired by the observation that non-adjacent reference lines primarily favor texture patterns with sharp and strongly oriented edges. In these cases, MPMs are chosen more frequently because there is often a strong correlation between the texture patterns of neighboring blocks and the current block. On the other hand, choosing a non-MPM for intra-frame prediction is an indication that edges are inconsistently distributed in neighboring blocks, and therefore, MRL prediction modes are not expected to be very useful in this case. Furthermore, it has been observed that MRL does not provide additional coding gain when the intra-frame prediction mode is a planar mode, as this mode is typically used for smooth regions. Therefore, MRL excludes planar modes, which are always one of the MPMs. The angle or DC prediction process in MRL is very similar to the case of directly adjacent reference lines. However, for angular modes with non-integer slopes, a DCT-based interpolation filter (DCTIF) is always used. This design choice is supported by both experimental results and empirical observations that MRL is most favorable for sharp and strongly directional edges, where DCTIF is more suitable because it preserves more high frequencies than some other filters.
[0068] From a hardware design perspective, applying multiple reference lines, as proposed in the initial approach, requires additional cost for line buffers to hold the additional reference lines. In typical hardware designs, line buffers are part of the on-chip memory architecture used for image and video encoding, and minimizing their on-chip area is crucial. To address this issue, the MRL is disabled, and no MRL is signaled for encoding units attached to the top boundary of the CTU. In this way, the additional buffers used to hold non-adjacent reference lines are bounded by 128, which is the width of the maximum unit size.
[0069] In some known methods, intra-prediction fusion methods have been proposed to improve the accuracy of intra-prediction. More specifically, if the current block is a luma block, and it is encoded using a non-integer slope angle mode instead of the ISP mode, and the block size (width * height) is greater than 16, then two prediction blocks generated from two different reference lines will be "fused," where the prediction fusion is calculated as a weighted sum of the two prediction blocks. More specifically, the index i is specified using the current signaling method in the bitstream. line i The first reference line at () and the prediction block generated from that reference line using the selected intra-frame prediction mode are represented as p ( line i ),in This refers to the operation of generating prediction blocks from a reference line using a given intra-frame prediction mode. In known methods, the reference line... line i+1 It is implicitly chosen as the second reference line. That is, the second reference line is an index position further away from the current block relative to the first reference line. Similarly, the predicted block generated from the second reference line is represented as... p ( line i+1 The weighted sum of the two prediction blocks is obtained as follows and used as the predictor for the current block according to equation (1).
[0070] (1), in p fusion Indicates fusion prediction, w 0 and w 1 represents two weighting factors, which were set to 3 / 4 and 1 / 4 respectively in the experiment.
[0071] In the intra-frame prediction method described above, the predictor is derived based on neighboring reference samples. However, this itself depends on the reference samples available to the current CU. The availability of samples depends on two factors: 1) whether the sample has been reconstructed, and 2) whether the sample belongs to a logical unit that the current CU is allowed to use.
[0072] To determine whether a sample has been reconstructed, we consider the VVC partitioning structure. (Reference) Figure 5 Each image is divided into tiles of a square CTU, which are processed in raster scan order. When the intra-frame prediction method is performed on the current CU 602 in the current CTU 502, samples belonging to other CTUs preceding the current CTU 502 in raster scan order are reconstructed and available for prediction. Samples belonging to CTUs following the current CTU 502 in raster scan order are not reconstructed and are therefore unavailable.
[0073] Each CTU 502 itself is divided into CUs through a hierarchical structure consisting of quadtrees, binary trees, and ternary trees. An example of this partitioning is shown in... Figure 6 As shown in the diagram, the scanning order of the CUs within the CTU 502 is determined by the partitioning structure. For single-level partitioning, the partitions are scanned in the following order: 1) from left to right for horizontal binary or ternary partitioning, 2) from top to bottom for vertical binary or ternary partitioning, and 3) from top left, top right, bottom left, and bottom right for quadtree partitioning.
[0074] If a partition contains further hierarchical splits, all CUs within that partition are scanned before proceeding to the next partition. Figure 6 An example of dividing a CTU 502 into 15 CUs is shown. Figure 6 Each CU 602 in the array is numbered from 1 to 15 to indicate its scan order. When an intra-frame prediction method is performed on the current CU in the current CTU, samples of other CUs belonging to the current CTU but preceding the current CU in the current CTU's partitioned scan order are reconstructed and available for prediction. Samples of CUs belonging to the current CU or following the current CU in the current CTU's partitioned scan order are not reconstructed and are therefore unavailable.
[0075] Samples belonging to CTUs preceding the current CTU in the raster scan sequence are considered reconstructed according to the above definition. However, they are not necessarily usable for intra-frame prediction. To be considered usable for prediction, they must also belong to logical units allowed for use by the current CU. A picture can be divided into sub-picture partitions, each containing an integer number of CTUs. Figure 9A The illustration shows a slice partitioning 900 of an image for intra-frame prediction according to some embodiments of the present disclosure. Samples belonging to slice partition 904 other than the slice containing the current CU (CTU 902) are not available for intra-frame prediction. Imposing this restriction allows slices to be decoded independently.
[0076] Figure 9B The diagram illustrates tile partitioning 901 of an image for intra-frame prediction according to some embodiments of the present disclosure. Samples belonging to tile partitioning 906 other than the tiles containing the current CU (CTU 902) are not available for intra-frame prediction. Imposing this restriction allows for independent tile decoding.
[0077] Figure 9C A diagram is shown of wavefront parallel processing 903 for intra-frame prediction of images according to some embodiments of the present disclosure.
[0078] The way intra-frame prediction methods (e.g., slice partitioning, tile partitioning, or wavefront parallel processing) handle the unavailability of the reference samples required for prediction varies depending on the method. When such samples are unavailable, the method can simply be disabled. Alternatively, some extrapolation of the unavailable samples can be performed, such as by boundary expansion.
[0079] References above Figures 9A to 9C In the described intra-prediction method, the predictor is derived only from spatially adjacent reference samples. However, greater coding gain can be achieved by expanding the region of reconstructed samples available for deriving the predictor. An example of this is the intra-block copy (IBC) mode.
[0080] Figure 10A A diagram of an IBC 1000 according to some embodiments of the present disclosure is shown.
[0081] refer to Figure 10A When predicting the current CU 1004 using intra-frame block copy mode, the block vector (BV) 1010 is signaled to indicate which block within the same image will be copied to be used as the predictor 1012 for the current block. Signaling of this block vector can be performed in the bitstream by signaling the block vector difference (BVD), allowing the block vector to be determined by adding the BVD to the block vector predictor. Alternatively, if the block vector from a previous CU is an exact match for the current block vector, it can be signaled via a merge flag. Regardless of the signaling mechanism, the block vector points to a location within the same image indicating a sample block of the same size as the current CU 1004, which is used as the predictor block for the current CU 1004. Several restrictions can be applied to the block vector. In a first restriction, the block vector 1010 can point to a sample block in the current image that can be used for intra-frame prediction. In a second restriction, the block vector 1010 can be restricted to a search area defined for the IBC tool, which can be smaller than the current image. For example, in VVC, the IBC search area is the current CTU 1002 and the previous CTU. In ECM, when the CTU size is 256×256, the IBC search area is the current CTU line 1006 and the CTU line above it 1014; or when the current CTU size is 128×128 or smaller, the IBC search area is the current CTU line 1006 and the two CTU lines above it.
[0082] To further improve coding performance, a subpixel IBC method was introduced in ECM-9.0. More specifically, in addition to the existing integer-pixel IBC, 1 / 16 pixel resolution is also supported. An 8-tap luminance filter and a chrominance filter used for fractional motion compensation in VVC are used to interpolate subpixel values. After the IBC block is encoded, the 1 / 16 pixel resolution is stored for encoding future blocks.
[0083] Figure 10B A diagram of intraTMP 1001 according to some embodiments of the present disclosure is shown.
[0084] refer to Figure 10B IntraTMP is an intra-frame prediction mode similar to IBC, because the current CU 1022 is also predicted from sample blocks from the current image. IntraTMP can be selected only as the prediction mode for CUs of size 64x64 or smaller. However, unlike IBC, in intraTMP, the block vector 1024 is not signaled in the bitstream. Instead, the decoder 201 compares a predefined L-shaped template or other shaped template of the reconstructed sample adjacent to the current CU 1022 with a template of the same shape of the candidate predictor within a predetermined search area. For the L-shaped template, the adjacent samples to the left and above of the current CU 1022 or the intraTMP predictor 1026 are used. Let the width of the left template region be TmpW, and the width of the upper template region be TmpH. Other template shapes include a left template and an upper template, where the left template includes only the template region on the left, and the upper template includes only the template region on the top.
[0085] The intraTMP predictor block is determined by finding the best candidate template that matches the current CU template. The best match can be determined by finding the template that minimizes the sum of absolute differences (SAD) or the sum of absolute transformed differences (SATD), or by comparing the hashes between templates. The search algorithm over the search region can be exhaustive (e.g., by scanning templates over the search region using a sample resolution offset) or fast (e.g., by performing a coarse search first, followed by a local refinement search around the best match from the coarse search). In any case, the search algorithm is performed identically by both the encoder and decoder, such that both encoder 101 and decoder 201 are implicitly aware of the intraTMP predictor without needing to be signaled in the bitstream. Figure 10B The example shown is of intraTMP, with the current CU template and the best matching template indicated by the shading of the shaded lines.
[0086] Still referencing Figure 10B For the selected intraTMP predictor 1026, the sample block corresponding to intraTMP predictor 1026 must be completely contained within the search region. The search region is... Figure 10B The area is shown in dashed shaded lines. Within the current CTU1020, the search area is restricted to a rectangular block of samples, with one corner restricted by the top left corner of the current CTU 1020 and the other corner restricted by the top left corner of the current CU 1022.
[0087] Outside the current CTU 1020, the search region is limited by imposing a maximum length on the intraTMP block vector of (searchRangeWidth 1028, searchRangeHeight 1030), where searchRangeWidth 1028 and searchRangeHeight 1030 are set to be proportional to the size of the current CU 1022. That is, searchRangeWidth = a * BlkW, searchRangeHeight = a * BlkH, where "a" is a constant controlling the gain / complexity tradeoff, and BlkW and BlkH are the width and height of the current CU 1022, respectively. Here, "a" is set to 5 in the ECM-7.0 testing software. searchRangeHeight 1030 limits the length of the block vector only in the negative vertical direction (i.e., the direction to the top of the image). For block vectors with a positive vertical component, the search region is limited by the bottom boundary of the current CTU row. For example, in Figure 10B In this case, the search area extends to the bottom boundary of the left CTU 1032, regardless of the value of searchRangeHeight 1030. Furthermore, these limitations on the search range do not apply to the current CTU 1020. For example, for a small CU where searchRangeWidth 1028 and searchRangeHeight 1030 may be smaller compared to the size of the current CTU 1020, the search area still extends to the top left corner of the current CTU 1020.
[0088] In addition to the constraints imposed by the search region, the intraTMP predictor 1026 and its template must consist of samples available for intra-frame prediction. For example, the boundaries of the search region are still covered by image, slice, or tile boundaries. Let the coordinates of the top-left corner of currentCU relative to the current image be (currCuX, currCuY). Then, the left boundary of the intraTMP search region is initially intraTmpLeftBound = CurrCUx - SearchRangeWidth. To account for image boundaries, the left boundary is cropped to allow a sample width of TmpW for the predictor template. intraTmpLeftBound = max(intraTmpLeftBound, TmpW).
[0089] To accelerate the template matching process, the search region is initially traversed horizontally or vertically in increments of 2 pixels each time. This is also known as the search subsampling factor of 2. This results in a 4x reduction in template matching search complexity. After finding the best match from the initial search, a refinement process is performed. Refinement is accomplished by a second template matching search around the best match with a reduced range. In ECM-7.0, the reduced range is set to BlkH / 2.
[0090] Figure 10C A diagram is shown illustrating an extended search area for intraTMP 1003 according to some embodiments of the present disclosure.
[0091] refer to Figure 10C For small CUs, the search range may be overly restricted, making it difficult to obtain a good predictor. For example, a 4×4 CU would only be allowed a maximum search range of (20,20). To improve this for small CUs, some implementations impose a minimum restriction on the intraTMP search range. For example, searchRangeWidth=max(a*BlkW,minSearchRange), searchRangeHight=max(a*BlkH,minSearchRange), where minSearchRange is set to 128.
[0092] This implementation excludes some regions in the current CTU 1020 that are available for prediction. Here, it is proposed to expand the search area in the current CTU 1020 to include the region directly above and to the left of the current CU 1022. The proposed modified search area is... Figure 10C The area marked with a crosshair is added relative to the area marked with a crosshair. Figure 10B The search area.
[0093] Figure 11A diagram is shown of subpixel position 1100 for intraTMP according to some embodiments of the present disclosure.
[0094] refer to Figure 11 ECM-9.0 employs multi-candidate intraTMP. A candidate list is constructed, where candidate BVs are ordered in ascending order of their template matching cost. In other words, BV0 has the lowest template matching cost, BV1 has the next lowest template matching cost, and so on. The index of the selected candidate is signaled in the bitstream.
[0095] Subpixel precision is enabled for intraTMP in ECM-9.0. More specifically, intraTMP blocks can have a quarter-pixel fractional resolution (BV). Three subpixel offsets (e.g., half-pixel, quarter-pixel, and three-quarter-pixel) are supported in eight directions around a whole-pixel location, resulting in... Figure 11 The subpixel position is shown. If a non-zero subpixel offset is signaled, the direction index is signaled to indicate which direction to use. The four-tap DCT-IF interpolation filter in the ECM is used for subpixel interpolation in intraTMP.
[0096] ECM-9.0 also employs model-derived intraTMP prediction blocks. Model parameters are derived using the template of the current block and the corresponding matching template. Predicted blocks are obtained by applying the model to filter the reference block.
[0097] ECM-9.0 employs a fusion method that uses a Wiener filter-based weight derivation approach to mix multiple reference blocks to derive the final prediction block. The block vectors (BVs) of these reference blocks are obtained through a template matching search process.
[0098] ECM-9.0 uses three additional intraTMP modes: left template, top template, and L-shaped fusion mode. The left template and top template modes derive template matching candidates using only the left or top template, while the L-shaped fusion mode uses both the left and top templates. The fusion mode fuses the two or five best L-shaped candidates based on a linear combination formula that minimizes the template matching cost or the mean-square error (MSE).
[0099] In the current ECM-9.0, the syntax related to intraTMP is shown in Table 1.
[0100]
[0101] Table 1: intraTMP syntax elements in ECM-9.0 Referring to Table 1, `intra_tmp_flag` indicates whether the intra-prediction type of the current block is `intraTMP`, `intra_tmp_fusion_flag` indicates whether fusion is used in the current block, and `intra_tmp_fusion_idx` specifies the candidate set used for `intraTMP` fusion. `intra_tmp_fusion_idx` ranges from 0 to 2 and is used to indicate one of three candidate sets: {BV0 to BV4}, {BV5 to BV9}, and {BV10 to BV14}. `intra_tmp_fusion_weight_type` indicates whether the SAD-based weight derivation method or the Wiener-filter-based weight derivation method is used. `intra_tmp_idx` specifies the index of the BV in the candidate list used for the current block. `intra_tmp_idx` ranges from 0 to 18. Candidates from the L-shaped template, the upper template, and the left template are included in the same candidate list. `intra_tmp_sub_pel_precision_idx` specifies the precision index of the current block. The range of `intra_tmp_sub_pel_precision_idx` is 0 to 3, indicating integer pixel precision, 1 / 2 pixel precision, 1 / 4 pixel precision, and 3 / 4 pixel precision, respectively. `intra_tmp_sub_pel_direction_idx` specifies the sub-pixel direction index of the current block. The range of `intra_tmp_sub_pel_phase_idx` is 0 to 7.
[0102] In the current ECM-9.0, an intraTMP block can only be encoded at subpixel BV resolution if the current block is not encoded using either fused intraTMP (e.g., intra_tmp_fusion_flag is 1) or filtered intraTMP (e.g., intra_tmp_filter_flag is 1). If the block is encoded using either fused or filtered intraTMP, the intraTMP block will only have an integer pixel (integer) resolution BV.
[0103] After encoding an intraTMP block, regardless of whether the current intraTMP block has an integer pixel or quarter-pixel fraction resolution (BV), only the integer pixel BV information of the current intraTMP block is stored for encoding future blocks. More specifically, if the current intraTMP block has a quarter-pixel fraction BV, it is first rounded to an integer pixel resolution. Then, the integer pixel BV is converted to a 1 / 16 pixel resolution (the current integer pixel BV is shifted left by 4). The converted 1 / 16 pixel resolution BV is stored for encoding future blocks in ECM-9.0.
[0104] Refer again Figure 6 The diagram illustrates an example of dividing a CTU 502 into CU 602s, where each CU 602 consists of three color components—one luminance (Y) component and two chromaticity (Cb and Cr) components. In this case where all color components are divided together, the partitioning structure is called a single tree.
[0105] VVC supports an alternative partitioning structure for intra-frame slicing, where the luma and chroma components are partitioned independently. This structure can be called a dual-tree partitioning structure. This dual-tree partitioning structure introduces overhead for signaling the partitioning of both the luma and chroma trees. However, the greater flexibility of the partitioning structure outweighs this additional signaling overhead, such as in cases where a larger chroma CU can be signaled independently of the luma partitioning structure.
[0106] In single-tree partitioning, each intra-prediction CU signals both the luma and chroma intra-prediction modes. Because the luma intra-prediction mode is signaled first, information from the luma prediction mode can be used to determine the chroma prediction mode. For example, a direct mode (DM) flag can be signaled to indicate the intra-prediction direction in which the chroma component replicates the luma component. That is, the chroma component will use the same intra-prediction mode as the luma component (e.g., ...). Figure 8 (As shown).
[0107] In a two-tree partitioning, the entire luma tree of the luma CU is decoded before the chroma tree. Therefore, information such as the selected luma intra-prediction mode of the co-located luma CU is available when determining the chroma intra-prediction mode of the chroma CU. Similar to the single-tree case, the Direct Mode (DM) flag can be signaled to indicate the intra-prediction direction in which the chroma CU copies the co-located luma CU. The Direct Block Vector (DBV) flag can be signaled to indicate the intraTMP or IBC block vector in which the chroma CU copies the co-located luma CU.
[0108] Figure 12 A diagram of a direct block vector (DBV) pattern 1200 according to some embodiments of the present disclosure is illustrated. Figure 13 A diagram illustrating a DBV prediction process 1300 according to some embodiments of the present disclosure is provided. (Described together) Figure 12 and Figure 13 .
[0109] refer to Figure 12 The five corresponding luma blocks 1202 of chroma block 1204 can be checked according to a predefined order (e.g., from center (C), top left (TL), top right (TR), bottom left (BL) to bottom right (BR) blocks). If any of these five luma blocks 1202 are encoded in IBC mode or intraTMP mode, then the current chroma block 1204 meets the DBV mode enabling condition. Furthermore, when DBV mode is selected, the BV of the first IBC or intraTMP encoded luma block will be used to derive the BV for encoding the current chroma block 1204.
[0110] For example, refer to Figure 13 The BV(bvL) of the first luma block encoded in IBC or IntraTMP mode will be used to derive the BV(bvC) of the current chroma block, where bvC[0] and bvC[1] represent pixel displacements along the horizontal and vertical directions relative to the top-left corner of the current chroma block, respectively. bvC can be derived by scaling bvL using the chroma sampling ratio. For example, if the chroma format is 4:2:0 (which indicates that the chroma components are subsampled by a factor of 2 in both the vertical and horizontal directions), then bvC can be obtained by 0.5*bvL. Figure 13 As shown, the reference chroma block 1304 pointed to by the block vector (BV) of the current chroma block (bvC) will be directly copied to predict the current chroma block 1302. Assume that the coordinates of the top left corner of the current chroma block 1302 are (xCb, yCb). The reference chroma block 1304 refers to the block of the same size with the coordinates of the top left corner being (xCb+bvC[0], yCb+bvC[1]).
[0111] When using a single-tree partition, the chroma components and their corresponding luma components use the same partition. Currently, DBV is not allowed for single-tree partitioning. However, when the luma blocks of a single-tree CU are encoded using IBC, the chroma direct mode (DM) is interpreted as signaling the same prediction as DBV. That is, when the prediction mode of the corresponding chroma block is signaled as direct mode (DM), then bvL is inherited by the corresponding chroma block to derive bvC and directly copy the reference chroma block indicated by bvC.
[0112] Figure 14 The illustration shows a diagram of the spatial component 1400 of a convolutional cross-component intra-prediction model (CCCM) filter according to some embodiments of the present disclosure. Figure 15A diagram is shown of a reference region 1500 for calculating a CCCM filter according to some embodiments of the present disclosure. Figure 16 Reference region 1600 for computing a block vector guided CCCM (BVG-CCCM) filter according to some embodiments of the present disclosure is shown. It will be described together. Figures 14 to 16 .
[0113] refer to Figure 14 CCCM applies a convolutional cross-component model to predict chromaticity samples from reconstructed luminance samples. The reconstructed luminance samples are downsampled to match a lower-resolution chromaticity grid used when employing chromaticity subsampling. There may be a single model available. Therefore, a single-model or multi-model variant of CCCM can be selected. The multi-model variant uses two models, one derived for samples above the average luminance reference value and the other for the remaining samples. A multi-model CCCM mode can be selected for a PU with at least 128 available reference samples.
[0114] The 7-tap convolutional filter used in CCCM consists of a 5-tap spatial component with a sign shape, a nonlinear term P, and a bias term B. The input to the 5-tap spatial component of the filter consists of a center (C) luminance sample (which is co-located with the chrominance sample to be predicted) and its adjacent samples above / north (N), below / south (S), left / west (W), and right / east (E), as shown below. Figure 14 As shown.
[0115] The nonlinear term P is represented as the square of the center brightness sample C and is scaled to the range of sample values of the content, as shown in equation (1) below.
[0116] P=(C*C+midVal)>>bitDepth (1).
[0117] That is, for a 10-bit content, the nonlinear term P is calculated as shown in equation (2) below.
[0118] P=(C*C+512)>>10 (2).
[0119] The bias term B represents the scalar offset between the input and output and is set to an intermediate chroma value (512 for 10-bit content).
[0120] The output of the filter is calculated as the filter coefficients c. i The convolution between the input value and the chromaticity sample is truncated to the range of valid chromaticity samples, as shown in equation (3) below.
[0121] predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B (3).
[0122] The filter coefficients c are calculated by minimizing the mean-square error (MSE) between the predicted and reconstructed chromaticity samples in the reference region. i . Figure 15 Reference region 1502 is shown, consisting of six rows of chroma samples above and to the left of prediction unit (PU) 1504. Reference region 1502 extends one PU width to the right of the boundary of PU 1504 and one PU height below the boundary of PU 1504. The reference region is adjusted to include only usable samples. Extension 1506 to the area shown in blue is needed to support the "side sample" of the plus shape space filter, and extension 1506 is filled when there are unusable areas.
[0123] MSE minimization is performed by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and chrominance output. The autocorrelation matrix is decomposed using LDL, and the final filter coefficients are calculated using inverse substitution. This process roughly follows the calculation of adaptive loop filter (ALF) coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. This method uses only integer arithmetic.
[0124] The CCCM mode is signaled using context adaptive binary arithmetic (CABAC) PU-level flags. A new CABAC context is included to support this. When signaling is involved, CCCM can be considered a sub-mode of the cross-component linear model (CCLM). That is, the CCCM flag is signaled only if the intra-frame prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-model CCCM).
[0125] In CCCM technology, Figure 15 The spatially adjacent reconstructed regions shown are used to calculate the CCCM filter. (Reference) Figure 16In BVG-CCCM, the block vector (BV) of the co-occurrence lumen block encoded in IBC or intraTMP mode is used to determine reference regions (reference lumen blocks), where arrows represent lumen BV and chromaticity BV. Instead of spatially adjacent regions, these reference regions are used to compute BVG-CCCM parameters. BVG-CCCM uses the computed filter parameters and co-occurrence lumen samples to form CCCM predictions.
[0126] The BVG-CCCM mode is enabled only in intra-frame slices. Additionally, a sequence parameter set (SPS) level flag is introduced to enable or disable the mode. Furthermore, the BVG-CCCM mode uses an 11-tap filter for cross-component prediction, as shown in equation (4) below.
[0127] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P(C)+c6P(N)+c7P(S)+c8P(W)+c9P(E)+c 10 B (4). The input to the spatial 5-tap component of the filter consists of the center (C) luminance sample (which is co-located with the chrominance sample to be predicted) and its adjacent samples above / north (N), below / south (S), left / west (W), and right / east (E), as shown below. Figure 14 As shown. The nonlinear term P is represented by the square power of the corresponding brightness sample, and B is the bias term.
[0128] To determine the block vector from the co-positional lumen block, BVG-CCCM checks the correlation with... Figure 12 The five positions shown are identical in the DBV mode. Using dual-tree partitioning, when the corresponding luminance is encoded using intraTMP or IBC, the chroma block can be encoded as BVG-CCCM or DBV mode.
[0129] Refer again Figure 3 Transform module 308 can transform the video signal in the residual block from the pixel domain to the transform domain (e.g., the frequency domain depending on the transform method). It should be understood that in some examples, transform module 308 can be skipped, and the video signal can be omitted from the transform domain.
[0130] Quantization module 310 can be configured to quantize the coefficients at each location in the encoded block to generate a quantization level for that location. The current block can be a residual block. That is, quantization module 310 can perform quantization processing on each residual block. A residual block can include N×M locations (samples), each location associated with a transformed or untransformed video signal / data (such as luminance and / or chrominance information), where N and M are positive integers. In this disclosure, the transformed or untransformed video signal at a particular location before quantization is referred to herein as a “coefficient”. After quantization, the quantized value of the coefficient is referred to herein as a “quantization level” or “level”.
[0131] Quantization can be used to reduce the dynamic range of transformed or untransformed video signals, allowing fewer bits to be used to represent the signal. Quantization typically involves dividing by the quantization step size and subsequent rounding, while inverse quantization (also known as dequantization) involves multiplying by the quantization step size. The quantization step size can be indicated by the quantization parameter (QP). This type of quantization is called scalar quantization. Quantization of all coefficients within a coded block can be performed independently, and this method is used in some existing video compression standards such as H.264 / AVC and H.265 / HEVC. The QP in quantization can affect the bitrate of the frames used to encode / decode the video. For example, a higher QP can result in a lower bitrate, while a lower QP can result in a higher bitrate.
[0132] For an N×M coded block, a specific coding scan order can be used to convert the two-dimensional (2D) coefficients of the block into a one-dimensional (1D) sequence for coefficient quantization and encoding. Typically, the coding scan begins at the top left corner of the coded block and stops at the bottom right corner or the last non-zero coefficient / level in the lower right direction. It should be understood that the coding scan order can include any suitable order, such as a zigzag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. The quantization of coefficients within the coded block can utilize coding scan order information. For example, it can depend on the state of previous quantization levels along the coding scan order. To further improve coding efficiency, the quantization module 310 can use more than one quantizer (e.g., two scalar quantizers). Which quantizer will be used to quantize the current coefficient can depend on information preceding the current coefficient in the coding scan order. Such quantization is called dependent quantization.
[0133] refer to Figure 3The encoding module 320 can be configured to encode the quantization level at each position in the coded block into the bitstream. In some embodiments, the encoding module 320 can perform entropy coding on the coded block. Entropy coding can use various binarization methods (such as Golomb-Rice binarization) to convert each quantization level into a corresponding binary representation (such as binary bin). The binary representation can then be further compressed using an entropy coding algorithm. The compressed data can be added to the bitstream. In addition to the quantization levels, the encoding module 320 can encode various other information, such as block type information of the coding unit, prediction mode information, partitioning unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information input from, for example, prediction modules 304 and 306. In some embodiments, the encoding module 320 can perform residual coding on the coded block to convert the quantization levels into the bitstream. For example, after quantization, there can be N×M quantization levels for an N×M block. These N×M levels can be zero or non-zero values. If the non-zero level is not binary, it can be further binarized to binary bin, for example, using combined truncated Rice (TR) and finite EGk binarization.
[0134] Non-binary syntax elements can be mapped to binary codewords. The bijective mapping between symbols and codewords (typically using simple structured code) is called binarization. Binary arithmetic coding can be used to encode both binary syntax elements and the binary symbols (also called bins) used for non-binary data. The core coding engine of context-adaptive binary arithmetic coding (CABAC) supports two operating modes: context coding mode, where bins are encoded using an adaptive probability model, and a less complex bypass mode, which uses a fixed probability of 1 / 2. The adaptive probability model is also called the context, and assigning the probability model to each bin is called context modeling.
[0135] like Figure 3 As shown, the dequantization module 312 can be configured to dequantize the quantization level, and the inverse transform module 314 can be configured to perform an inverse transform on the coefficients transformed by the transform module 308. The reconstructed residual block generated by the dequantization module 312 and the inverse transform module 314 can be combined with the prediction unit predicted by the prediction module 304 or 306 to generate a reconstructed block.
[0136] Filter module 316 may include at least one of a deblocking filter, an sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter removes block distortion caused by boundaries between blocks in the reconstructed image. The SAO module corrects the offset relative to the original video on a pixel-by-pixel basis for the video that has been deblocked. ALF can be performed based on values obtained by comparing the reconstructed and filtered video with the original video. Buffer module 318 can be configured to store the reconstructed blocks or images calculated by filter module 316, and can provide the reconstructed and stored blocks or images to inter-frame prediction module 304 when inter-frame prediction is performed.
[0137] Figure 4 Some embodiments according to this disclosure are shown. Figure 2 A detailed block diagram of an exemplary decoder 201 in the decoding system 200. (See attached diagram.) Figure 4 As shown, decoder 201 may include decoding module 402, inverse quantization module 404, inverse transform module 406, inter-frame prediction module 408, intra-frame prediction module 410, filter module 412, and buffer module 414. It should be understood that... Figure 4 Each element shown is illustrated independently to represent a distinct feature function within the video decoder, and does not imply that each component is formed by a separate hardware or software configuration unit. That is, for ease of illustration, elements are listed as components, and at least two elements can be combined to form a single element, or a single element can be divided into multiple elements to perform its function. It should also be understood that some elements are not essential for performing the functions described in this disclosure, but may be optional elements used to improve performance. It should also be understood that these elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the decoder 201.
[0138] When a video stream is input from a video encoder (e.g., encoder 101), the input stream can be decoded by decoder 201 in the reverse process of the video encoder. Therefore, for ease of description, some details of the decoding described above regarding encoding can be skipped. Decoding module 402 can be configured to decode the stream to obtain various information encoded into it, such as the quantization level at each position in the encoded block. In some embodiments, decoding module 402 can perform entropy decoding (decompression) corresponding to entropy coding (compression) performed by the encoder, such as, for example, VideoLAN coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), PIPE coding, etc., to obtain a binary representation (e.g., a binary binary). The decoding module 402 can also convert the binary representation to a quantization level using Golomb-Rice binarization (including, for example, EGk binarization and combined TR and finite EGk binarization). In addition to the quantization level of the position in the transform unit, the decoding module 402 can decode various other information, such as parameters used for Golomb-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, partitioning unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 can perform rearrangement on the bitstream to reconstruct and rearrange the data from a 1D sequence into 2D rearranged blocks using a reverse scan method based on the coding scan order used by the encoder.
[0139] The dequantization module 404 can be configured to dequantize the quantization level at each location of the encoded block (e.g., a 2D reconstructed block) to obtain coefficients at each location. In some embodiments, the dequantization module 404 can also perform dependent dequantization based on quantization parameters provided by the encoder, which include information related to the quantizers used in dependent quantization, such as the quantization step size used by each quantizer.
[0140] The inverse transform module 406 can be configured to perform inverse transforms, such as the inverse discrete cosine transform (DCT), inverse DST, and inverse Karhunen-Loève transform (KLT), respectively, for the DCT, DST, and KLT performed by the encoder, to transform data from the transform domain (e.g., coefficients) back to the pixel domain (e.g., luminance and / or chrominance information). In some embodiments, the inverse transform module 406 can selectively perform transform operations (e.g., DCT, DST, KLT) based on multiple pieces of information such as the prediction method, the size of the current block, and the prediction direction.
[0141] Inter-frame prediction module 408 and intra-frame prediction module 410 can be configured to generate prediction blocks based on information related to the generation of prediction blocks provided by decoding module 402 and information about previously decoded blocks or images provided by buffer module 414. As described above, when intra-frame prediction is performed in the same manner as the encoder, if the size of the prediction unit and the size of the transform unit are the same, intra-frame prediction can be performed on the prediction unit based on the pixels to the left, the upper left, and the top of the prediction unit. However, if the size of the prediction unit and the size of the transform unit are different when performing intra-frame prediction, intra-frame prediction can be performed using reference pixels based on the transform unit.
[0142] For example, the inter-frame prediction module 408 can be configured to receive from the encoder a bitstream including a reference frame, the current frame, and an indication of weighting factors associated with a multiple-hypothesis prediction (MHP) process. The inter-frame prediction module 408 can be configured to perform the MHP process on a CU located in the current frame based on a search block (e.g., the reference frame and / or a reference template) in the reference frame. In some embodiments, to perform the MHP process, the inter-frame prediction module 408 can be configured to perform template matching on the CU located in the current frame based on the search block and weighting factors in the reference frame to obtain motion information. In some embodiments, to perform the MHP process, the inter-frame prediction module 408 can be configured to identify a weighting factor index associated with the weighting factors based on template matching. The inter-frame prediction module 408 can be configured to identify the weighting factor symbol of the weighting factors based on an indication included in the bitstream. The inter-frame prediction module performs the inter-frame prediction process to decode the bitstream based on the current frame, the reference frame, the weighting factor index, and the weighting factor symbol of the weighting factors.
[0143] The reconstructed block or reconstructed image, a combination of the outputs from the inverse transform module 406 and the prediction modules 408 or 410, can be provided to the filter module 412. The filter module 412 may include a deblocking filter, an offset correction module, and an ALF. The buffer module 414 can store the reconstructed image or block and use it as a reference image or reference block for the inter-frame prediction module 408, and can also output the reconstructed image.
[0144] Within the scope of this disclosure, the encoding module 320 and the decoding module 402 can be configured to encode video images using a quantization-level binarization scheme with Rice parameters adapted to bit depth and / or bit rate, in order to improve encoding efficiency.
[0145] BVG-CCCM may not currently be permitted for single-tree partitioning. Therefore, BVG-CCCM cannot be used for slices encoded using single-tree partitioning, including intra-frame and inter-frame slices. Inter-frame slices can currently only be encoded using single-tree partitioning. Therefore, for such slices encoded using single-tree partitioning, the coding performance of chroma blocks can be further improved.
[0146] To overcome these challenges of other techniques, this disclosure proposes allowing BVG-CCCM to be used for single-tree partitioning. Unlike the dual-tree partitioning case described above, each chroma block in a single-tree partition uniquely corresponds to a co-positional luma block. Therefore, skipping as referenced... Figure 12 The inspection process described for the DBV pattern.
[0147] Still referencing Figure 4 To overcome these and other challenges, this disclosure proposes an intra-prediction module 410 configured to perform BVG-CCCM prediction for a single-tree partition.
[0148] For luma blocks encoded with intraTMP or IBC, the reference luma block `ref_luma` is indicated by the block vector `bvL`. Similarly, the corresponding reference chroma block `ref_chroma` is indicated by the block vector `bvC`, where `bvC` is a scaled version of `bvL` that depends on the chroma format of the video sequence. For example, when the chroma format is 4:2:0, `bvC` can be obtained by scaling `bvL` by 0.5 for both the horizontal and vertical directions. When the chroma format is 4:4:4, `bvC` can be equal to `bvL`. When the chroma format is 4:2:2, the horizontal component of `bvC` can be obtained by scaling the horizontal component of `bvL` by 0.5, and the vertical component of `bvC` can be set to be equal to the vertical component of `bvL`.
[0149] Intra-prediction module 410 (or inter-prediction module 408) can determine whether the BVG-CCCM mode meets the enable conditions for a single-tree partition based on the enable conditions. For a single-tree partition, if the co-occurring luma block of the current chroma block is encoded using intraTMP or IBC, then intra-prediction module 410 (or inter-prediction module 408) can determine that the BVG-CCCM mode meets the enable conditions.
[0150] More specifically, when a single-tree partition is used to encode the current slice (including both intra-frame and inter-frame slices), and if the co-position luma block is encoded using IBC or intraTMP, the intra-frame prediction module 410 (or the inter-frame prediction module 408) can use the corresponding block vector bvL to derive the chroma block vector bvC according to the chroma format. The luma reference region and the chroma reference region can be as follows: Figure 16 The diagram shows the determination of the luminance BV and chrominance BV, where the arrows represent the luminance BV and chrominance BV. Instead of spatially adjacent regions, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use these reference regions to calculate the BVG-CCCM filter parameters. When the BVG-CCCM mode is selected, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use the calculated filter parameters and co-located luminance samples to form CCCM predictions. Similar to the BVG-CCCM mode used for dual-tree partitioning, the proposed BVG-CCCM mode for single-tree partitioning can use an 11-tap filter for cross-component prediction, as shown in equation (5) below.
[0151] predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P(C) + c6P(N) + c7P(S) +c8P(W) + c9P(E)+ c 10 B (5) The input to the spatial 5-tap component of the filter consists of a center (C) luminance sample (which is co-located with the chrominance sample to be predicted) and its adjacent samples above / north (N), below / south (S), left / west (W), and right / east (E), as shown below. Figure 14 As shown. The nonlinear term P represents the square of the corresponding brightness sample, and B is the bias term.
[0152] When deriving filter parameters, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use the derived filter to filter the co-position luma block in order to generate a BVG-CCCM prediction for the current chroma block.
[0153] When the BVG-CCCM mode meets the enable conditions, the intra-frame prediction module 410 (or inter-frame prediction module 408) can determine whether to select BVG-CCCM in various ways, as described below.
[0154] In some implementations, when the CCCM flag is true and BVG-CCCM has been determined to meet the enable conditions, the BVG-CCCM flag can be signaled during single-tree partitioning. If the BVG-CCCM flag equals 1, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use BVG-CCCM to generate a prediction for the current chroma block. If the BVG-CCCM flag equals 0, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use regular CCCM to generate a prediction for the current chroma block. Here, BVG-CCCM is conditionally signaled after CCCM (e.g., it can be considered a sub-mode of CCCM). If the corresponding luminance Cbf is non-zero, the BVG-CCCM flag can be signaled.
[0155] The above arrangement describes a method for controlling BVG-CCCM mode at the CU level using CU-level signaling. Additionally, BVG-CCCM mode can be enabled jointly or individually at different levels (e.g., Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, and Slice Header (SH) level).
[0156] Figure 17 A flowchart of a decoding method 1700 according to some embodiments of the present disclosure is illustrated. Method 1700 may be performed by a system (e.g., decoding system 200, decoder 201, inter-frame prediction module 408, intra-frame prediction module 410, etc.). Method 1700 may include operations 1702-1708 as described below. It should be understood that some steps may be optional, and some steps may be performed simultaneously or in combination with... Figure 17 The different orders shown are executed sequentially.
[0157] refer to Figure 17 At 1702, the system can determine whether the corresponding chroma block meets the enabled conditions for prediction using the BVG-CCCM mode in response to whether the single-tree partitioning of the same-position luma block is encoded using intraTMP or IBC. For example, refer to Figure 4 The intra-frame prediction module 410 (or inter-frame prediction module 408) can determine whether the BVG-CCCM mode meets the enable conditions for single-tree partitioning based on the enable conditions. For single-tree partitioning, if the co-occurring luma block of the current chroma block is encoded using intraTMP or IBC, the intra-frame prediction module 410 (or inter-frame prediction module 408) can determine that the BVG-CCCM mode meets the enable conditions.
[0158] At 1704, the system can parse the bitstream to obtain the BVG-CCCM flag. In some implementations, the BVG-CCCM flag can be parsed from the bitstream after the CCCM flag. In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero. In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels. For example, see [reference]. Figure 4 The intra-frame prediction module 410 (or inter-frame prediction module 408) can parse the bitstream to obtain the BVG-CCCM flag.
[0159] At 1706, the system can generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block. In some implementations, the BVG-CCCM mode prediction can be associated with intra-image prediction or inter-image prediction. In some implementations, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the system can generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned iso-luminance block in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block. In some implementations, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the system can filter the single-tree partitioned iso-luminance block in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block. For example, refer to... Figure 4 When a single-tree partition is used to encode the current slice (including both intra-frame and inter-frame slices), and if the co-position luma block is encoded using IBC or intra-frame TMP, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use the corresponding block vector bvL to derive the chroma block vector bvC according to the chroma format. The luma reference region and chroma reference region can be as follows: Figure 16As shown in the figure, where the arrows represent the luminance BV and chrominance BV. Instead of spatially adjacent regions, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use these reference regions to calculate the BVG-CCCM filter parameters. When the BVG-CCCM mode is selected, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use the calculated filter parameters and co-located luminance samples to form CCCM predictions. Similar to the BVG-CCCM mode for dual-tree partitioning, the proposed BVG-CCCM mode for single-tree partitioning can use an 11-tap filter for cross-component prediction, as shown in equation (5) above. The input of the spatial 5-tap component of the filter consists of the center (C) luminance sample (which is co-located with the chrominance sample to be predicted) and its above / north (N), below / south (S), left / west (W), and right / east (E) adjacent samples, as shown in the figure. Figure 14 As shown. The nonlinear term P represents the square of the corresponding luminance sample, and B is the bias term. When deriving the filter parameters, the intra-frame prediction module 410 (or the inter-frame prediction module 408) can use the derived filter to filter the co-position luminance block to generate a BVG-CCCM prediction for the current chrominance block.
[0160] At 1708, the system can respond to the BVG-CCCM flag indicating the selection of the BVG-CCCM mode for the corresponding chroma block, generating a CCCM mode prediction for that chroma block. For example, refer to... Figure 4 If the BVG-CCCM flag is equal to 0, the intra-frame prediction module 410 (or inter-frame prediction module 408) can use the regular CCCM mode to generate CCCM predictions for the current chroma block.
[0161] Figure 18 A flowchart of an encoding method 1800 according to some embodiments of the present disclosure is shown. Method 1800 may be performed by a system (e.g., encoding system 100, encoder 101, inter-frame prediction module 304, or intra-frame prediction module 306, etc.). Method 1800 may include operations 1802-1808 as described below. It should be understood that some steps may be optional, and some steps may be performed simultaneously or in combination with... Figure 18 The different orders shown are executed sequentially.
[0162] refer to Figure 18 At 1802, the system can determine, in response to whether the single-tree partitioned co-positional luma blocks are encoded using intraTMP or IBC, that the corresponding chroma block meets the enabled conditions for prediction using the BVG-CCCM mode. For example, refer to... Figure 3The intra-frame prediction module 306 (or inter-frame prediction module 304) can determine whether the BVG-CCCM mode meets the enable conditions for a single-tree partition based on whether the enable conditions are met. For a single-tree partition, if the co-occurring luma block of the current chroma block is encoded using intraTMP or IBC, the intra-frame prediction module 306 (or inter-frame prediction module 304) can determine that the BVG-CCCM mode meets the enable conditions.
[0163] At position 1804, the system can encode the BVG-CCCM flag. In some implementations, the BVG-CCCM flag can be encoded after the CCCM flag. In some implementations, the BVG-CCCM flag can be signaled when the cbf_luma syntax element is non-zero. In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels. For example, see [reference]. Figure 3 The intra-frame prediction module 306 (or inter-frame prediction module 304) can encode the BVG-CCCM flag.
[0164] At 1806, the system can respond to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, and generate a BVG-CCCM mode prediction for that chroma block. In some implementations, the BVG-CCCM mode prediction can be associated with intra-image prediction or inter-image prediction. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the system can generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned isoluminance block to generate the BVG-CCCM mode prediction for that chroma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the system can filter the single-tree partitioned isoluminance block to generate the BVG-CCCM mode prediction for that chroma block. For example, refer to... Figure 3 When a single-tree partition is used to encode the current slice (including both intra-frame and inter-frame slices), and if the co-position luma block is encoded using IBC or intraTMP, the intra-frame prediction module 306 (or the inter-frame prediction module 304) can use the corresponding block vector bvL to derive the chroma block vector bvC according to the chroma format. The luma reference region and chroma reference region can be defined as follows: Figure 16As shown in the figure, where the arrows represent the luminance BV and chrominance BV. Instead of spatially adjacent regions, the intra-frame prediction module 306 (or inter-frame prediction module 304) can use these reference regions to calculate the BVG-CCCM filter parameters. When the BVG-CCCM mode is selected, the intra-frame prediction module 306 (or inter-frame prediction module 304) can use the calculated filter parameters and co-located luminance samples to form CCCM predictions. Similar to the BVG-CCCM mode for dual-tree partitioning, the proposed BVG-CCCM mode for single-tree partitioning can use an 11-tap filter for cross-component prediction, as shown in equation (5) above. The input of the spatial 5-tap component of the filter consists of the center (C) luminance sample (which is co-located with the chrominance sample to be predicted) and its above / north (N), below / south (S), left / west (W), and right / east (E) adjacent samples, as shown in the figure. Figure 14 As shown. The nonlinear term P represents the square of the corresponding luminance sample, and B is the bias term. When deriving the filter parameters, the intra-frame prediction module 306 (or the inter-frame prediction module 304) can use the derived filter to filter the co-position luminance block to generate the BVG-CCCM prediction of the current chrominance block.
[0165] At 1808, the system can respond to the BVG-CCCM flag indicating the selection of the BVG-CCCM mode for the corresponding chroma block, and generate a CCCM mode prediction for the corresponding chroma block. For example, refer to... Figure 3 If the BVG-CCCM flag is equal to 0, the intra-frame prediction module 306 (or inter-frame prediction module 304) can use the normal CCCM mode to generate CCCM predictions for the current chroma block.
[0166] In all respects of this disclosure, the functions described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media can be processors (such as...) Figure 1 and Figure 2The processor 102 in the computer can access any available medium. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, HDDs (such as disk storage devices or other magnetic storage devices), flash drives, SSDs, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (such as a mobile device or computer). As used herein, disks and optical discs include CDs, laser discs, optical discs, digital video discs (DVDs), and floppy disks, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0167] According to one aspect of this disclosure, a method for decoding by a decoder is provided. The method may include: in response to a single-tree partitioned co-occurrence luma block being encoded using intraTMP or IBC pairs, a processor determining that the corresponding chroma block meets the enable conditions for prediction using a BVG-CCCM mode. The method may include the processor parsing the bitstream to obtain a BVG-CCCM flag. The method may also include: in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block.
[0168] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block may include: the processor generating an eleven-tap filter for cross-component prediction based on a single-tree partitioned isotope luma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block may include: the processor filtering the single-tree partitioned isotope luma block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0169] In some implementations, the method may include: in response to a BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, the processor generates a CCCM mode prediction for the corresponding chroma block.
[0170] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0171] In some implementations, the BVG-CCCM flag can be parsed from the bitstream after the CCCM flag.
[0172] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0173] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0174] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0175] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for a corresponding chroma block, a memory stores instructions to generate a BVG-CCCM mode prediction for the corresponding chroma block. These instructions, when executed by a processor, cause the processor to generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned co-occurrence luma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for a corresponding chroma block, a memory stores instructions to generate a BVG-CCCM mode prediction for the corresponding chroma block. These instructions, when executed by a processor, cause the processor to filter the single-tree partitioned co-occurrence luma block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0176] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate a CCCM mode prediction for the corresponding chroma block in response to a BVG-CCCM mode flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block.
[0177] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0178] In some implementations, the BVG-CCCM flag can be parsed from the bitstream after the CCCM flag.
[0179] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0180] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0181] According to another aspect of this disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to the single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0182] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided for storing instructions for a decoder. When executed by a decoder's processor, the instructions cause the decoder's processor to determine, in response to the single-tree partitioning of the co-occurrence luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using the BVG-CCCM mode. When executed by the decoder's processor, the instructions can cause the decoder's processor to parse the bitstream to obtain a BVG-CCCM flag. When executed by the decoder's processor, the instructions can cause the decoder's processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0183] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the decoder's processor, may cause the decoder's processor to generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned co-occurrence luma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the decoder's processor, may cause the decoder's processor to filter the single-tree partitioned co-occurrence luma block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0184] In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to generate a CCCM mode prediction for the corresponding chroma block in response to a BVG-CCCM mode flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block.
[0185] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0186] In some implementations, the BVG-CCCM flag can be parsed from the bitstream after the CCCM flag.
[0187] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0188] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0189] According to one aspect of this disclosure, a method for encoding by an encoder is provided. The method may include, in response to a single-tree partitioning of a co-occurrence luma block being encoded using intraTMP or IBC, determining by a processor that the corresponding chroma block meets the enable conditions for prediction using a BVG-CCCM mode. The method may include, by the processor, encoding a BVG-CCCM flag. The method may also include, in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block, generating a BVG-CCCM mode prediction for the corresponding chroma block by the processor.
[0190] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block may include: the processor generating an eleven-tap filter for cross-component prediction based on a single-tree partitioned isotope luma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, the processor generating a BVG-CCCM mode prediction for the corresponding chroma block may include: the processor filtering the single-tree partitioned isotope luma block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0191] In some implementations, the method may include: in response to a BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, the processor generates a CCCM mode prediction for the corresponding chroma block.
[0192] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0193] In some implementations, the BVG-CCCM flag can be encoded after the CCCM flag.
[0194] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0195] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0196] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable condition for prediction using the BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to encode a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0197] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for a corresponding chroma block, a memory stores instructions to generate a BVG-CCCM mode prediction for the corresponding chroma block. These instructions, when executed by a processor, cause the processor to generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned isochroma block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for a corresponding chroma block, a memory stores instructions to generate a BVG-CCCM mode prediction for the corresponding chroma block. These instructions, when executed by a processor, cause the processor to filter the single-tree partitioned isochroma block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0198] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate a CCCM mode prediction for the corresponding chroma block in response to a BVG-CCCM mode flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block.
[0199] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0200] In some implementations, the BVG-CCCM flag can be encoded after the CCCM flag.
[0201] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0202] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0203] According to another aspect of this disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to determine, in response to a single-tree partitioning of a co-occurrence luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable conditions for prediction using a BVG-CCCM mode. The memory stores instructions that, when executed by the processor, cause the processor to encode a BVG-CCCM flag. The memory stores instructions that, when executed by the processor, cause the processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0204] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided for storing instructions for an encoder. When executed by a processor of the encoder, the instructions can cause the encoder processor to determine, in response to the single-tree partitioning of the same-position luma block being encoded using intraTMP or IBC, that the corresponding chroma block meets the enable condition for prediction using the BVG-CCCM mode. When executed by the encoder processor, the instructions can also cause the encoder processor to encode a BVG-CCCM flag. When executed by the encoder processor, the instructions can also cause the encoder processor to generate a BVG-CCCM mode prediction for the corresponding chroma block in response to the BVG-CCCM flag indicating that a BVG-CCCM mode is selected for the corresponding chroma block.
[0205] In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate an eleven-tap filter for cross-component prediction based on a single-tree partitioned isochromaticity block. In some implementations, in response to the BVG-CCCM flag indicating the selection of a BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to filter the single-tree partitioned isochromaticity block to generate a BVG-CCCM mode prediction for the corresponding chroma block.
[0206] In some implementations, when the instruction is executed by the encoder's processor, the encoder's processor may generate a CCCM mode prediction for the corresponding chroma block in response to a BVG-CCCM mode flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block.
[0207] In some implementations, BVG-CCCM pattern prediction can be associated with intra-image prediction or inter-image prediction.
[0208] In some implementations, the BVG-CCCM flag can be encoded after the CCCM flag.
[0209] In some implementations, the BVG-CCCM flag can be signaled when a cbf_luma syntax element is non-zero.
[0210] In some implementations, the BVG-CCCM mode can be enabled at one or more of the SPS, PH, PPS, or SH levels.
[0211] According to another aspect of this disclosure, a non-transitory computer-readable medium for storing a bitstream generated according to one or more of the operations described herein.
[0212] The foregoing description of the embodiments will thus reveal the general nature of this disclosure, enabling others to readily modify and / or adapt such embodiments for various applications by applying knowledge within the skill of the art without departing from the general concept of this disclosure, without excessive experimentation. Therefore, based on the teachings and guidance presented herein, such modifications and alterations are intended to fall within the meaning and scope of equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for descriptive and not limiting purposes, and that the terminology or terminology of this specification will be interpreted by those skilled in the art based on the teachings and guidance.
[0213] Embodiments of this disclosure have been described above using functional building blocks that illustrate the implementation of specified functions and their relationships. For ease of description, the boundaries of these functional building blocks have been arbitrarily defined herein. Alternative boundaries may be defined, provided that the specified functions and their relationships are properly performed.
[0214] The summary and abstract may set forth one or more, but not all, exemplary embodiments of this disclosure as contemplated by the inventors, and are therefore not intended to limit this disclosure and the appended claims in any way.
[0215] Various functional blocks, modules, and steps have been disclosed above. The arrangements provided are illustrative and not limiting. Therefore, functional blocks, modules, and steps can be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments include only a subset of functional blocks, modules, and steps, and any such subset is permitted.
[0216] The breadth and scope of this disclosure should not be limited by any of the foregoing exemplary embodiments, but should be defined solely by the appended claims and their equivalents.
Claims
1. A method for decoding by a decoder, comprising: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the processor determines that the corresponding chroma block meets the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCMBVG-CCCM mode. The processor parses the bitstream to obtain the BVG-CCCM flag; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, the processor generates a BVG-CCCM mode prediction for the corresponding chroma block.
2. The method according to claim 1, wherein, In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, the processor generates a BVG-CCCM mode prediction for the corresponding chroma block, including: The processor generates an eleven-tap filter for cross-component prediction based on the single-tree division of co-position brightness blocks; and The processor filters the single-tree segmented luminance blocks to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
3. The method according to claim 1, further comprising: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, the processor generates a CCCM mode prediction for the corresponding chroma block.
4. The method according to claim 1, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
5. The method according to claim 1, wherein, The BVG-CCCM flag is parsed from the bitstream after the CCCM flag.
6. The method according to claim 1, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is signaled.
7. The method according to claim 1, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
8. A decoder, comprising: processor; and The memory stores instructions that, when executed by the processor, cause the processor to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Parse the bitstream to obtain the BVG-CCCM flag; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
9. The decoder according to claim 8, wherein, In response to the BVG-CCCM flag indicating selection of the BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the memory stores an instruction that, when executed by the processor, causes the processor to: An eleven-tap filter for cross-component prediction is generated based on the single-tree division of co-position brightness blocks. as well as The single tree is divided into corresponding luminance blocks and filtered to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
10. The decoder according to claim 8, wherein, The memory stores instructions, which, when executed by the processor, cause the processor to: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, a CCCM mode prediction for the corresponding chroma block is generated.
11. The decoder according to claim 8, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
12. The decoder according to claim 8, wherein, The BVG-CCCM flag is parsed from the bitstream after the CCCM flag.
13. The decoder according to claim 8, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is signaled.
14. The decoder according to claim 8, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
15. An apparatus for decoding, comprising: processor; and The memory stores instructions that, when executed by the processor, cause the processor to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Parse the bitstream to obtain the BVG-CCCM flag; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
16. A non-transitory computer-readable medium storing instructions, said instructions, when executed by a processor of a decoder, causing the processor of the decoder to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Parse the bitstream to obtain the BVG-CCCM flag; and In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
17. The non-transitory computer-readable medium according to claim 16, wherein, In response to the BVG-CCCM flag indicating the selection of the BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the processor of the decoder, causes the processor of the decoder to: An eleven-tap filter for cross-component prediction is generated based on the single-tree division of co-position brightness blocks. as well as The single tree is divided into corresponding luminance blocks and filtered to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
18. The non-transitory computer-readable medium according to claim 16, wherein, When the instruction is executed by the processor of the decoder, the processor of the decoder causes the decoder to: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, a CCCM mode prediction for the corresponding chroma block is generated.
19. The non-transitory computer-readable medium according to claim 16, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
20. The non-transitory computer-readable medium of claim 16, wherein, The BVG-CCCM flag is parsed from the bitstream after the CCCM flag.
21. The non-transitory computer-readable medium according to claim 16, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is signaled.
22. The non-transitory computer-readable medium according to claim 16, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
23. A method for encoding by an encoder, comprising: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the processor determines that the corresponding chroma block meets the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCMBVG-CCCM mode. The processor encodes the BVG-CCCM flag; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, the processor generates a BVG-CCCM mode prediction for the corresponding chroma block.
24. The method according to claim 23, wherein, In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, the processor generates a BVG-CCCM mode prediction for the corresponding chroma block, including: The processor generates an eleven-tap filter for cross-component prediction based on the single-tree division of co-position brightness blocks; and The processor filters the single-tree segmented luminance blocks to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
25. The method of claim 23, further comprising: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, the processor generates a CCCM mode prediction for the corresponding chroma block.
26. The method according to claim 23, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
27. The method according to claim 23, wherein, The BVG-CCCM flag is encoded after the CCCM flag.
28. The method according to claim 23, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is encoded.
29. The method according to claim 23, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
30. An encoder, comprising: processor; and The memory stores instructions that, when executed by the processor, cause the processor to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Encode the BVG-CCCM mark; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
31. The encoder according to claim 30, wherein, In response to the BVG-CCCM flag indicating the selection of the BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the memory stores an instruction that, when executed by the processor, causes the processor to: An eleven-tap filter for cross-component prediction is generated based on the single-tree division of co-position brightness blocks. as well as The single tree is divided into corresponding luminance blocks and filtered to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
32. The encoder according to claim 30, wherein, The memory stores instructions, which, when executed by the processor, cause the processor to: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, a CCCM mode prediction for the corresponding chroma block is generated.
33. The encoder according to claim 30, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
34. The encoder according to claim 30, wherein, The BVG-CCCM flag is encoded after the CCCM flag.
35. The encoder according to claim 30, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is signaled.
36. The encoder according to claim 30, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
37. An apparatus for encoding, comprising: processor; and The memory stores instructions that, when executed by the processor, cause the processor to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Encode the BVG-CCCM mark; as well as In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
38. A non-transitory computer-readable medium storing instructions that, when executed by a processor of an encoder, cause the processor of the encoder to: In response to the fact that the single-tree partitioning of the same-position luma block is encoded using intraTMP or intraBlock Copying (IBC), the corresponding chroma block is determined to meet the conditions for enabling prediction using the block vector-guided BVG-convolutional cross-component intra-prediction model CCCM BVG-CCCM mode. Encode the BVG-CCCM mark; and In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is selected for the corresponding chroma block, a BVG-CCCM mode prediction for the corresponding chroma block is generated.
39. The non-transitory computer-readable medium according to claim 38, wherein, In response to the BVG-CCCM flag indicating the selection of the BVG-CCCM mode for the corresponding chroma block, in order to generate a BVG-CCCM mode prediction for the corresponding chroma block, the instruction, when executed by the processor of the encoder, causes the processor of the encoder to: An eleven-tap filter for cross-component prediction is generated based on the single-tree division of co-position brightness blocks. as well as The single tree is divided into corresponding luminance blocks and filtered to generate the BVG-CCCM mode prediction for the corresponding chrominance blocks.
40. The non-transitory computer-readable medium according to claim 38, wherein, When the instruction is executed by the processor of the encoder, the processor of the encoder causes the processor to: In response to the BVG-CCCM flag indicating that the BVG-CCCM mode is not selected for the corresponding chroma block, a CCCM mode prediction for the corresponding chroma block is generated.
41. The non-transitory computer-readable medium according to claim 38, wherein, The BVG-CCCM pattern prediction is associated with intra-image prediction or inter-image prediction.
42. The non-transitory computer-readable medium according to claim 38, wherein, The BVG-CCCM flag is encoded after the CCCM flag.
43. The non-transitory computer-readable medium according to claim 38, wherein, When the cbf_luma syntax element is non-zero, the BVG-CCCM flag is signaled.
44. The non-transitory computer-readable medium according to claim 38, wherein, The BVG-CCCM mode is enabled at one or more of the following levels: Sequence Parameter Set (SPS) level, Image Header (PH) level, Image Parameter Set (PPS) level, or Slice Header (SH) level.
45. A non-transitory computer-readable medium for storing a bitstream generated by one or more operations according to claims 23-29.