Method and device for decoder-side intra-mode derivation
By unifying PDPC calculations and applying a fusion scheme in decoder-side intra-mode derivation, the method optimizes video decoding for efficient compression and quality maintenance in video coding technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2026-03-30
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing video data while maintaining video quality due to limitations in bandwidth and memory resources, particularly in decoder-side intra-mode derivation (DIMD) processes.
The method involves unifying position-dependent intra-predictive combination (PDPC) calculations by modifying intra-predictions based on boundary reference samples and disabling certain modes like DC, planar, or angular modes, and applying a fusion scheme as a weighted average of predictors in the DIMD mode to optimize video decoding.
This approach enhances video decoding efficiency by reducing bitrate requirements and maintaining video quality, adapting to different intra-predictive coding modes for improved compression performance.
Smart Images

Figure 0007837402000015 
Figure 0007837402000016 
Figure 0007837402000017
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to Provisional Application No. 63 / 250,186, filed on 29 September 2021, which is incorporated herein by reference in its entirety for all purposes.
[0002] This disclosure relates to the encoding and compression of video. More specifically, this disclosure relates to encoders and decoders with decoder-side intra-mode derivation (DIMD). [Background technology]
[0003] Digital video is supported by a variety of electronic devices, including digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video conferencing devices, and video streaming devices. These electronic devices transmit and receive digital video data over communication networks, or otherwise communicate and / or store digital video data in storage devices. Due to limitations in the bandwidth capacity of communication networks and the memory resources of storage devices, video coding may be used to compress video data according to one or more video coding standards before the video data is transmitted or stored. Examples of video coding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding generally utilizes prediction methods (e.g., interpretation, intrapretation) that leverage the inherent redundancy of video data. The goal of video coding is to compress video data into a format that uses a lower bitrate while avoiding or minimizing a decrease in video quality. [Overview of the project] [Problems that the invention aims to solve]
[0004] An example of this disclosure provides a method and apparatus for video decoding using an intra prediction coding mode. [Means for solving the problem]
[0005] A first aspect of this disclosure provides a method for decoding video using an intra-predictive coding mode. The method may include the steps of: determining prediction sample values for one or more video blocks based on PDPC calculations for one or more intra-predictions of one or more video blocks, in order to unify position-dependent intra-predictive combination (PDPC) calculations in an intra-predictive coding mode, wherein the PDPC calculations modify the results of one or more intra-predictions based on a combination of boundary reference samples; and disabling PDPC calculations for DC mode or planar mode in response to a determination that a DC mode or planar mode is applied to one or more intra-predictions of one or more video blocks, in order to unify PDPC calculations in an intra-predictive coding mode.
[0006] A second aspect of this disclosure provides a method for decoding video using an intra-predictive coding mode. The method may include the steps of: determining predicted sample values for one or more video blocks based on PDPC calculations for one or more intra-predictions of one or more video blocks, in order to unify position-dependent intra-predictive combination (PDPC) calculations in an intra-predictive coding mode, wherein the PDPC calculations modify the results of one or more intra-predictions based on a combination of boundary reference samples; and disabling PDPC calculations for an angular mode in response to a determination that an angular mode is applied to one or more intra-predictions of one or more video blocks, in order to unify PDPC calculations in an intra-predictive coding mode.
[0007] A third aspect of the present disclosure provides a method for video decoding using an intra predictive coding mode. The method may include the steps of: the decoder determining whether a fusion scheme is applied in a DIMD mode, wherein the fusion scheme is applied as a weighted average of predictors in the DIMD mode; the decoder applying an offset to the available directional modes in the DIMD mode to obtain an offset directional mode; and the decoder determining whether the offset directional mode should be added to the list of most probable modes (MPMs) based on whether a fusion scheme is applied in the DIMD mode.
[0008] A fourth aspect of the present disclosure provides a method for video coding using an intra-predictive coding mode. The method may include the steps of: determining predictive sample values for one or more video blocks based on PDPC calculations for one or more intra-predictions of one or more video blocks, in order to unify position-dependent intra-predictive combination (PDPC) calculations in an intra-predictive coding mode, wherein the PDPC calculations modify the results of one or more intra-predictions based on a combination of boundary reference samples; and disabling PDPC calculations for DC mode or planar mode in response to a determination that a DC mode or planar mode is applied to one or more intra-predictions of one or more video blocks, in order to unify PDPC calculations in an intra-predictive coding mode.
[0009] A fifth aspect of this disclosure provides a method for video coding using an intra-predictive coding mode. The method may include the steps of: determining predictive sample values for one or more video blocks based on PDPC calculations for one or more intra-predictions of one or more video blocks, in order to unify position-dependent intra-predictive combination (PDPC) calculations in an intra-predictive coding mode, wherein the PDPC calculations modify the results of one or more intra-predictions based on a combination of boundary reference samples; and disabling PDPC calculations for an angular mode in response to a determination that an angular mode is applied to one or more intra-predictions of one or more video blocks, in order to unify PDPC calculations in an intra-predictive coding mode.
[0010] A sixth aspect of the present disclosure provides a method for video coding using an intra predictive coding mode. The method may include the steps of: the encoder determining whether a fusion scheme is applicable in a DIMD mode, wherein the fusion scheme is applied as a weighted average of predictors in a DIMD mode; the encoder applying an offset to the available directional modes in a DIMD mode to obtain an offset directional mode; and the encoder determining whether the offset directional mode should be added to the list of most likely modes (MPMs) based on whether a fusion scheme is applicable in a DIMD mode.
[0011] Please understand that the above general description and the following detailed description are for illustrative and explanatory purposes only and are not intended to limit this disclosure.
[0012] The accompanying drawings incorporated herein and constituting part of herein serve to illustrate the principles of this disclosure, together with the descriptions, and provide examples consistent with this disclosure. [Brief explanation of the drawing]
[0013] [Figure 1]A block diagram showing an exemplary system for encoding and decoding video blocks according to some implementations of the present disclosure. [Figure 2] A block diagram showing an exemplary video encoder according to some implementations of the present disclosure. [Figure 3] A block diagram showing an exemplary video decoder according to some implementations of the present disclosure. [Figure 4A-4E] A block diagram showing how a frame is recursively partitioned into a plurality of video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 5A] A diagram showing the definition of samples used by PDPC applied to the prediction mode according to some implementations of the present disclosure. [Figure 5B] A diagram showing the definition of samples used by PDPC applied to the prediction mode according to some implementations of the present disclosure. [Figure 5C] A diagram showing the definition of samples used by PDPC applied to the prediction mode according to some implementations of the present disclosure. [Figure 5D] A diagram showing the definition of samples used by PDPC applied to the prediction mode according to some implementations of the present disclosure. [[ID=2A]] [Figure 6] A diagram showing examples of allowed GPM partitions according to some implementations of the present disclosure. [Figure 7] A diagram showing examples of selected pixels for which gradient analysis is performed according to some implementations of the present disclosure. [Figure 8] A diagram showing a convolution process according to some implementations of the present disclosure. [Figure 9] A diagram showing prediction fusion by weighted averaging of two HoG modes and one planar mode according to some implementations of the present disclosure. [Figure 10] A diagram showing the template and its reference samples used in TIMD according to some implementations of the present disclosure. [Figure 11A]A block diagram showing a video decoding process using TIMD according to some implementations of the present disclosure. [Figure 11B] A block diagram showing a video decoding process using TIMD according to some implementations of the present disclosure. [Figure 11C] A block diagram showing a video decoding process using TIMD according to some implementations of the present disclosure. [Figure 11D] A block diagram showing a video decoding process using TIMD according to some implementations of the present disclosure. [Figure 11E] A block diagram showing a video encoding process using TIMD according to some implementations of the present disclosure. [Figure 11F] A block diagram showing a video encoding process using TIMD according to some implementations of the present disclosure. [Figure 12A] A block diagram showing a video decoding process using DIMD according to some implementations of the present disclosure. [Figure 12B] A block diagram showing a video decoding process using DIMD according to some implementations of the present disclosure. (continued on next page) [Figure 12C] A block diagram showing a video decoding process using DIMD according to some implementations of the present disclosure. [Figure 12D] A block diagram showing a video encoding process using DIMD according to some implementations of the present disclosure. [Figure 13] A diagram showing fractional bits used in the proposed integerization method according to some implementations of the present disclosure. [Figure 14] A block diagram showing a computing environment coupled with a user interface according to some implementations of the present disclosure. [Figure 15] A block diagram showing video decoding according to some implementations of the present disclosure. [Figure 16] A block diagram showing video decoding according to some implementations of the present disclosure. [Figure 17]This is a block diagram showing video decoding in several implementations of the present disclosure. [Modes for carrying out the invention]
[0014] Next, exemplary embodiments will be referenced in detail, examples of which are shown in the accompanying drawings. The following description refers to the accompanying drawings, and unless otherwise noted, the same numbers in different drawings represent the same or similar elements. The implementations described below in the description of exemplary embodiments do not represent all implementations in accordance with this disclosure. Rather, they are merely examples of apparatus and methods in accordance with the aspects relating to this disclosure described in the accompanying claims.
[0015] The terms used in this disclosure are for the sole purpose of describing specific embodiments and are not intended to limit this disclosure. The singular forms “a,” “an,” and “the” are intended to include the plural forms when used in this disclosure and the appended claims, unless the context clearly indicates otherwise. It should also be understood that the terms “and / or” as used herein are intended to mean, and include, any or all possible combinations of one or more of the related enumerated items.
[0016] In this specification, terms such as “first,” “second,” and “third” may be used to describe various types of information, but it should be understood that these terms should not limit the information. These terms are used solely to distinguish one category of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information, and similarly, second information may be referred to as first information. When used herein, the term “case” may be understood, depending on the context, to mean “when,” “on the occasion of,” or “depending on judgment.”
[0017] Various video coding techniques are sometimes used to compress video data. Video coding is performed according to one or more video coding standards. For example, well-known video coding standards today include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its earlier standard VP9. Audio-Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standards Working Group of China. Most existing video coding standards are built upon well-known hybrid video coding frameworks, that is, using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in video images or sequences, and using transform coding to compact the energy of prediction errors. A key goal of video coding techniques is to compress video data into a format that uses lower bitrates while avoiding or minimizing a loss of video quality.
[0018] The first generation AVS standards include the Chinese national standards "Information Technology, Advanced Audio-Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio-Video Coding, Part 16: Radio, Television, and Video" (known as AVS+). Compared to the MPEG-2 standard, it can achieve approximately 50% bitrate savings at the same perceived quality. The video portion of the AVS1 standard was published as a Chinese national standard in February 2006. The second generation AVS standards include a series of Chinese national standards "Information Technology, Efficient Multimedia Coding" (known as AVS2), primarily aimed at transmitting additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was published as a Chinese national standard. Meanwhile, the video portion of the AVS2 standard was proposed by the Institute of Electrical and Electronics Engineers (IEEE) as one of the international standards for applications. The AVS3 standard is one of the next-generation video coding standards for UHD video applications, aiming to surpass the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS conference, the AVS3-P2 baseline was completed, achieving approximately 30% bitrate savings compared to the HEVC standard. Currently, there is one reference software called the High Performance Model (HPM), which is maintained by the AVS Group to demonstrate a reference implementation of the AVS3 standard.
[0019] Figure 1 is a block diagram illustrating an exemplary system 10 for parallel encoding and decoding of video blocks, according to several implementations of the present disclosure. As shown in Figure 1, system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. The source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, etc. In some implementations, the source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0020] In some implementations, the destination device 14 may receive the encoded video data to be decoded via link 16. Link 16 may comprise any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, link 16 may comprise a communication medium that allows the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may comprise routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device 12 to the destination device 14.
[0021] In some other implementations, the encoded video data may be sent from the output interface 22 to a storage device 32. The encoded video data in the storage device 32 may then be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, digital multipurpose disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or download. The file server may be any type of computer capable of storing the encoded video data and sending the encoded video data to the destination device 14. An exemplary file server includes a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both that is suitable for accessing the encoded video data stored on the file server. Transmission of the encoded video data from the storage device 32 may be streaming transmission, download transmission, or a combination of both.
[0022] As shown in Figure 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as a video acquisition device, e.g., a video camera, a video archive containing previously acquired video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. For example, if the video source 18 is a video camera in a security surveillance system, the source device 12 and the destination device 14 may form a camera phone or video phone. However, the implementations described herein may be generally applicable to video coding and may be applicable to wireless and / or wired applications.
[0023] Captured video, pre-captured video, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored in a storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or transmitter.
[0024] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 includes a receiver and / or a modem and may receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on the storage device 32 may include various syntax elements generated by the video decoder 20 for use by the video decoder 30 when decoding the video data. Such syntax elements may be contained within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0025] In some implementations, the destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to the user and may comprise any of various display devices such as a liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.
[0026] The video encoder 20 and video decoder 30 may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and is applicable to other video encoding / decoding standards as well. It is generally intended that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is generally intended that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.
[0027] The video encoder 20 and the video decoder 30 may each be implemented as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof, among other suitable encoder and / or decoder circuits. If the electronic device is partially implemented in software, it may store software instructions in a suitable non-transient computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed herein. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and each may be integrated as part of a combined encoder / decoder (CODEC) within its respective device.
[0028] Figure 2 is a block diagram showing an exemplary video encoder 20 in several implementation forms described in this application. The video encoder 20 may perform intra-predictive coding and inter-predictive coding of video blocks within a video frame. Intra-predictive coding relies on spatial prediction to reduce or eliminate spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or eliminate temporal redundancy in video data within adjacent video frames or pictures in a video sequence. Note that the term “frame” may be used synonymously with the terms “image” or “picture” in the field of video coding.
[0029] As shown in Figure 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transformation processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-prediction processing unit 46, and an intra-block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transformation processing unit 60, and an adder 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, may be placed between the adder 62 and the DPB 64 to filter block boundaries and remove blocky artifacts from the reconstructed video. In addition to the deblocking filter, other in-loop filters such as a sample adaptive offset (SAO) filter and / or an adaptive in-loop filter (ALF) may also be used to filter the output of adder 62. In some examples, the in-loop filters may be omitted, and the decoded video blocks may be provided directly to DPB64 by adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or it may be divided into one or more of the fixed or programmable hardware units shown.
[0030] The video data memory 40 may store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18 shown in Figure 1. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or reference pictures) used by the video encoder 20 when encoding the video data (e.g., in intra-predictive coding mode or inter-predictive coding mode). The video data memory 40 and the DPB 64 may be formed by any of various memory devices. In various examples, the video data memory 40 may be on-chip with the other components of the video encoder 20 or off-chip with respect to those components.
[0031] As shown in Figure 2, the partitioning unit 45 within the prediction processing unit 41, after receiving the video data, partitions the video data into video blocks. This partitioning may also include slicing, tiling (e.g., sets of video blocks), or dividing the video frame into other larger coding units (CUs) according to a predefined partitioning structure, such as a quad-tree (QT) structure associated with the video data. A video frame may be, or be considered to be, a two-dimensional array or matrix of samples having sample values. Samples in the array may also be called pixels or pels. The number of samples in the horizontal and vertical (or axis) directions of the array or picture defines the size and / or resolution of the video frame. A video frame may be divided into multiple video blocks, for example, by using QT partitioning. A video block may also be, or be considered to be, a two-dimensional array or matrix of samples having sample values, although it is smaller in dimensions than a video frame. The number of samples in the horizontal and vertical (or axis) directions of the video block defines the size of the video block. A video block may be further divided into one or more block divisions or subblocks (which may again form blocks) by iteratively using, for example, QT divisions, binary-tree (BT) divisions, or triple-tree (TT) divisions, or any combination thereof. Note that as used herein, the terms “block” or “video block” can refer to a portion of a frame or picture, particularly a rectangular (square or non-square) portion. For example, with reference to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transformation unit (TU), and / or a corresponding block, for example, a coding tree block (CTB), coding block (CB), prediction block (PB), or transformation block (TB), and / or a subblock.
[0032] The prediction processing unit 41 may select one of several possible predictive coding modes for the current video block, such as one of several intra predictive coding modes or one of several inter predictive coding modes, based on the error result (e.g., coding rate and distortion level). The prediction processing unit 41 may provide the resulting intra predictive coding block or inter predictive coding block to the adder 50 to generate a residual block, and provide it to the adder 62 to reconstruct the coding block for later use as part of a reference frame. The prediction processing unit 41 also provides syntactic elements such as motion vectors, intra-mode indicators, piecewise information, and other such syntactic information to the entropy coding unit 56.
[0033] To select an appropriate intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 may provide spatial predictions by performing intra-predictive coding of the current video block for one or more adjacent blocks in the same frame as the current block to be coded. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 may provide temporal predictions by performing inter-predictive coding of the current video block for one or more prediction blocks in one or more reference frames. The video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0034] In some implementations, the motion estimation unit 42 determines the interprediction mode of the current video frame by generating motion vectors that indicate the displacement of the video block in the current video frame relative to the predicted block in the reference video frame, according to a predetermined pattern in the sequence of video frames. Motion estimation performed by the motion estimation unit 42 is the process of generating motion vectors that estimate the motion of the video block. The motion vectors may indicate, for example, the displacement of the video block in the current video frame or picture relative to the predicted block in the reference frame relative to the current block encoded in the current frame. The predetermined pattern may specify the video frames in the sequence as P frames or B frames. The intraBC unit 48 may determine vectors for intraBC encoding, such as block vectors, in a similar manner to how the motion estimation unit 42 determines the motion vectors for interprediction, or it may determine the block vectors using the motion estimation unit 42.
[0035] The predicted blocks of the video block may be, or correspond to, blocks of a reference frame that are considered to be exactly the same as the video block to be encoded, in terms of pixel differences that can be determined by the Sum of Absolute Difference (SAD), Sum of Square Difference (SSD), or other difference metrics. In some implementations, the video encoder 20 may calculate the values of sub-integer pixel positions of the reference frame stored in the DPB64. For example, the video encoder 20 may interpolate the values of 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference frame. Thus, the motion estimation unit 42 may perform a motion search on the total pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision.
[0036] The motion estimation unit 42 calculates the motion vector of a video block in an interpredictive coded frame by comparing the position of a video block with the position of a predicted block in a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of which identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.
[0037] Motion compensation performed by the motion compensation unit 44 may include fetching or generating a prediction block based on a motion vector determined by the motion estimation unit 42. Upon receiving the motion vector of the current video block, the motion compensation unit 44 may locate the position of the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from the DPB 64, and transfer the prediction block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block may include a luminance (luma) difference component, a chroma difference component, or both. The motion compensation unit 44 may also generate syntactic elements associated with the video block of the video frame, which are used by the video decoder 30 when decoding the video block of the video frame. Syntactic elements may include, for example, syntactic elements that define the motion vector used to identify the prediction block, optional flags indicating the prediction mode, or any other syntactic information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.
[0038] In some implementations, the intraBC unit 48 may generate vectors and fetch prediction blocks in a manner similar to that described above in relation to the motion estimation unit 42 and the motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are called block vectors rather than motion vectors. Specifically, the intraBC unit 48 may determine which intraprediction mode to use to encode the current block. In some examples, the intraBC unit 48 may encode the current block using various intraprediction modes, for example, during separate encoding passes, and test their performance through rate distortion analysis. The intraBC unit 48 may then select an appropriate intraprediction mode to use from among the various tested intraprediction modes and generate an intramode indicator accordingly. For example, the intraBC unit 48 may calculate rate distortion values using rate distortion analysis for the various tested intraprediction modes and select the intraprediction mode with the best rate distortion characteristics from among the tested modes as the appropriate intraprediction mode to use. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to create the encoded block, and the bit rate (i.e., number of bits) used to create the encoded block. The intraBC unit 48 may calculate ratios from the distortion and rates of various encoded blocks to determine which intraprediction mode shows the best rate distortion value for that block.
[0039] In other examples, the intra-BC unit 48 may use the motion estimation unit 42 and the motion compensation unit 44 in whole or in part to perform such functions for intra-BC prediction in the implementation forms described herein. In any case, for intra-block copying, the predicted block may be a block that is considered to be a precise match to the block to be encoded in terms of pixel differences which may be determined by SAD, SSD, or other difference metrics, and the identification of the predicted block may include calculating values for sub-integer pixel positions.
[0040] Regardless of whether the predicted block is a block from the same frame due to intra-prediction or a block from a different frame due to inter-prediction, the video encoder 20 may form a residual video block and form a pixel difference value by subtracting the pixel value of the predicted block from the pixel value of the currently encoded video block. The pixel difference value forming the residual video block may include both luminance component differences and chrominance component differences.
[0041] The intra-prediction processing unit 46 may intra-predict the current video block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra-block copy prediction performed by the intra-BC unit 48, as described above. Specifically, the intra-prediction processing unit 46 may determine the intra-prediction mode to use to encode the current block. To do so, the intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, for example, in a separate encoding pass, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. The intra-prediction processing unit 46 may provide the entropy encoding unit 56 with information indicating the selected intra-prediction mode for the block. The entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode into a bitstream.
[0042] After the prediction processing unit 41 determines the predicted block of the current video block by interpretation or intraprediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be contained in one or more TUs and is provided to the transformation processing unit 52. The transformation processing unit 52 transforms the residual video data into residual transformation coefficients using a transformation such as a discrete cosine transform (DCT) or a conceptually similar transformation.
[0043] The conversion processing unit 52 may send the resulting conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0044] Following quantization, the entropy coding unit 56 entropy-codes the quantized transformation coefficients into a video bitstream using, for example, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), Syntax-based context-adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy coding methodology or technique. The coded bitstream is then transmitted to the video decoder 30 as shown in Figure 1, or archived in the storage device 32 as shown in Figure 1 for later transmission to or retrieval by the video decoder 30. The entropy coding unit 56 may also entropy-code the motion vector and other syntactic elements of the current video frame being coded.
[0045] The inverse quantization unit 58 and the inverse transformation processing unit 60 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual image block in the pixel region to generate a reference block for predicting other image blocks. As described above, the motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0046] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block created by the motion compensation unit 44 to create a reference block for storage in the DPB 64. The reference block may then be used by the intraBC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block for interpreting another video block in a subsequent video frame.
[0047] Figure 3 is a block diagram showing an exemplary video decoder 30 in several implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transformation processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-prediction unit 84, and an intra-BC unit 85. The video decoder 30 may perform a decoding process that is substantially the reverse of the encoding process described above with respect to the video encoder 20 in relation to Figure 2. For example, the motion compensation unit 82 may generate prediction data based on motion vectors received from the entropy decoding unit 80, while the intra-prediction unit 84 may generate prediction data based on an intra-prediction mode indicator received from the entropy decoding unit 80.
[0048] In some examples, units of the video decoder 30 may be tasked with performing the implementation of the present invention. Also, in some examples, the implementation of the present disclosure may be divided into one or more units of the video decoder 30. For example, the intraBC unit 85 may perform the implementation of the present invention alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intraprediction unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intraBC unit 85, and the functions of the intraBC unit 85 may be performed by other components of the predictive processing unit 81, such as the motion compensation unit 82.
[0049] The video data memory 79 may store video data, such as an encoded video bitstream, which is decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from a storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include an encoded picture buffer (CPB) that stores the encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data used by the video decoder 30 when decoding the video data (e.g., in intra-predictive coding mode or inter-predictive coding mode). The video data memory 79 and DPB 92 may be formed by any of various memory devices, such as dynamic random-access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, in Figure 3, the video data memory 79 and DPB 92 are depicted as two separate components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, the video data memory 79 may be on-chip with the other components of the video decoder 30, or it may be off-chip relative to those components.
[0050] During the decoding process, the video decoder 30 receives an encoded video bitstream representing the video blocks and associated syntactic elements of the encoded video frames. The video decoder 30 may also receive syntactic elements at the video frame level and / or video block level. The entropy decoding unit 80 of the video decoder 30 entropy-decodes the bitstream to generate quantized coefficients, motion vectors or intra-predictive mode indicators, and other syntactic elements. The entropy decoding unit 80 then transfers the motion vectors or intra-predictive mode indicators and other syntactic elements to the prediction processing unit 81.
[0051] When a video frame is encoded as an intra-predictive coding (I) frame or for an intra-coded prediction block in another type of frame, the intra-prediction unit 84 of the prediction processing unit 81 may generate prediction data for the video block of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.
[0052] When a video frame is encoded as an interpredictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 creates one or more prediction blocks for the video block of the current video frame based on the motion vector and other syntactic elements received from the entropy decoding unit 80. Each prediction block may be created from a reference frame in one of the reference frame lists. The video decoder 30 may construct the reference frame lists, List 0 and List 1, using a default construction technique based on the reference frames stored in the DPB92.
[0053] In some examples, when a video block is encoded according to the intraBC mode described herein, the intraBC unit 85 of the prediction processing unit 81 creates a prediction block of the current video block based on the block vector and other syntactic elements received from the entropy decoding unit 80. The prediction block may be located within the reconstructed region of the same picture as the current video block defined by the video encoder 20.
[0054] The motion compensation unit 82 and / or intra-BC unit 85 determine the prediction information for the video blocks of the current video frame by analyzing the motion vectors and other syntactic elements, and then use that prediction information to create the prediction blocks for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntactic elements to determine the prediction mode used to encode the video blocks of the video frame (e.g., intra-predict or inter-predict), the inter-predict frame type (e.g., B or P), construction information about one or more of the frame's reference frame lists, the motion vector for each inter-predict coded video block in the frame, the inter-predict status for each inter-predict coded video block in the frame, and other information for decoding the video blocks in the current video frame.
[0055] Similarly, the intraBC unit 85 may use some of the received syntactic elements, such as flags, to determine that the current video block was predicted using intraBC mode, construction information regarding which video blocks in the frame are in the reconstruction area and should be stored in the DPB 92, the block vector of each intraBC predicted video block in the frame, the intraBC predicted status of each intraBC predicted video block in the frame, and other information for decoding the video blocks in the current video frame.
[0056] The motion compensation unit 82 may also perform interpolation using an interpolation filter used by the video encoder 20 during the encoding of the video block to calculate the interpolated values of the sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntactic elements and use that interpolation filter to create the prediction block.
[0057] The inverse quantization unit 86 determines the degree of quantization by inverse quantization of the quantization transformation coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80, using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame. The inverse transformation processing unit 88 applies an inverse transformation, such as an inverse DCT, inverse integer transformation, or a conceptually similar inverse transformation process, to the transformation coefficients in order to reconstruct the residual blocks in the pixel region.
[0058] After the motion compensation unit 82 or intraBC unit 85 generates a predicted block for the current video block based on vectors and other syntactic elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse processing unit 88 with the corresponding predicted block generated by the motion compensation unit 82 and intraBC unit 85. For further processing of the decoded video block, an in-loop filter 91 such as a deblocking filter, SAO filter, and / or ALF may be placed between the adder 90 and the DPB. In some examples, the in-loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores a reference frame used for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may store the decoded video for later presentation on a display device such as the display device 34 in Figure 1.
[0059] In a typical video encoding process, a video sequence typically contains an ordered set of frames or pictures. Each frame may contain three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of chrominance samples (Cb). SCr is a two-dimensional array of chrominance samples (Cr). In other examples, a frame may be monochromatic and therefore contain only one two-dimensional array of luminance samples.
[0060] As shown in Figure 4A, the video encoder 20 (or more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may contain an integer number of CTUs that are sequentially ordered from left to right and top to bottom in the raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTUs are signaled by the video encoder 20 in the sequence parameter set so that all CTUs in the video sequence have the same size, which is one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not necessarily limited to a specific size. As shown in Figure 4B, each CTU may contain one CTB of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntactic elements used to encode the samples in the coding tree blocks. The syntactic elements describe how the video sequence can be reconstructed in the video decoder 30, including the characteristics of different types of units of the encoded blocks of pixels, as well as inter-prediction or intra-prediction, intra-prediction mode, motion vector, and other parameters. For a monochrome picture or a picture with three distinct color planes, the CTU may include a single coding tree block and syntactic elements used to encode samples of the coding tree block. The coding tree block may be an N×N block of samples.
[0061] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary, ternary, quadary, or a combination thereof, on the encoding tree block of the CTU to divide the CTU into smaller CUs. As depicted in Figure 4C, the 64x64 CTU400 is first partitioned into four smaller CUs, each with a block size of 32x32. Of the four smaller CUs, CU410 and CU420 are each partitioned into four 16x16 CUs by their block size. The two 16x16 CUs, CU430 and CU440, are each further partitioned into four 8x8 CUs by their block size. Figure 4D depicts a quadary tree data structure showing the final result of the partitioning process of the CTU400 depicted in Figure 4C, where each leaf node of the quadary tree corresponds to one CU of each size ranging from 32x32 to 8x8. Similar to the CTU depicted in Figure 4B, each CU may include a CB of luminance samples, two corresponding encoded blocks of saturation samples in frames of the same size, and syntactic elements used to encode the samples in the encoded blocks. For a monochrome picture or a picture with three distinct color planes, a CU may include a single encoded block and syntactic structures used to encode the samples in the encoded block. Note that the quadtree partitions depicted in Figures 4C and 4D are for illustrative purposes only, and a single CTU may be split into CUs based on quadtree / ternary / binary partitions to adapt to various localities. In a multi-type tree structure, a single CTU may be partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by binary and ternary structures. As shown in Figure 4E, there are five possible partition types for encoded blocks with width W and height H: 4 partitions, horizontal 2 partitions, vertical 2 partitions, horizontal 3 partitions, and vertical 3 partitions.
[0062] In some implementations, the video encoder 20 may further divide the encoded block of the CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction for inter-prediction or intra-prediction is applied. The PU of the CU may include a PB for luminance samples, two corresponding PBs for chrominance samples, and syntactic elements used to predict the PBs. For a monochrome picture or a picture with three distinct color planes, the PU may include a single PB and syntactic structures used to predict the PBs. The video encoder 20 may generate predicted luminance blocks, predicted Cb blocks, and predicted Cr blocks for the luminance, Cb, and Cr PBs of each PU of the CU.
[0063] The video encoder 20 may generate prediction blocks for the PU using intra-prediction or inter-prediction. If the video encoder 20 generates prediction blocks for the PU using intra-prediction, the video encoder 20 may generate prediction blocks for the PU based on decoded samples of frames associated with the PU. If the video encoder 20 generates prediction blocks for the PU using inter-prediction, the video encoder 20 may generate prediction blocks for the PU based on decoded samples of one or more frames other than the frames associated with the PU.
[0064] After the video encoder 20 has generated predicted luminance blocks, predicted Cb blocks, and predicted Cr blocks for one or more PUs of the CU, the video encoder 20 may generate a luminance residual block of the CU by subtracting the predicted luminance block of the CU from its original luminance coding block, such that each sample in the CU's luminance residual block represents the difference between a luminance sample in one of the CU's predicted luminance blocks and a corresponding sample in the CU's original luminance coding block. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block of the CU, respectively, such that each sample in the CU's Cb residual block represents the difference between a Cb sample in one of the CU's predicted Cb blocks and a corresponding sample in the CU's original Cb coding block, and each sample in the CU's Cr residual block represents the difference between a Cr sample in one of the CU's predicted Cr blocks and a corresponding sample in the CU's original Cr coding block.
[0065] Furthermore, as shown in Figure 4C, the video encoder 20 may use a quadtree partition to decompose the luminance residual blocks, Cb residual blocks, and Cr residual blocks of the CU into one or more luminance transformation blocks, Cb transformation blocks, and Cr transformation blocks, respectively. A transformation block is a rectangular (square or non-square) block of samples to which the same transformation is applied. The TU of the CU may include a transformation block for luminance samples, two corresponding transformation blocks for chrominance samples, and syntactic elements used to transform the transformation block samples. Thus, each TU of the CU may be associated with a luminance transformation block, a Cb transformation block, and a Cr transformation block. In some examples, a luminance transformation block associated with a TU may be a subblock of the luminance residual block of the CU. A Cb transformation block may be a subblock of the Cb residual block of the CU. A Cr transformation block may be a subblock of the Cr residual block of the CU. For a monochrome picture or a picture with three distinct color planes, the TU may include a single transformation block and syntactic structures used to transform the samples of the transformation block.
[0066] The video encoder 20 may generate a luminance coefficient block of the TU by applying one or more transformations to the luminance transformation block of the TU. The coefficient block may be a two-dimensional array of transformation coefficients. The transformation coefficients may be scalar quantities. The video encoder 20 may generate a Cb coefficient block of the TU by applying one or more transformations to the Cb transformation block of the TU. The video encoder 20 may generate a Cr coefficient block of the TU by applying one or more transformations to the Cr transformation block of the TU.
[0067] The video encoder 20 may, after generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), quantize the coefficient block. Quantization generally refers to the process by which the transformation coefficients are quantized to potentially reduce the amount of data used to represent the transformation coefficients, thereby achieving further compression. After the video encoder 20 has quantized the coefficient block, the video encoder 20 may entropy encode the syntactic elements representing the quantized transformation coefficients. For example, the video encoder 20 may perform CABAC on the syntactic elements representing the quantized transformation coefficients. Finally, the video encoder 20 may output a bitstream containing a sequence of bits that form a representation of the encoded frame and associated data, which is stored in the storage device 32 or transmitted to the destination device 14.
[0068] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntactic elements from it. The video decoder 30 may reconstruct frames of video data based at least partially on the syntactic elements obtained from the bitstream. The process of reconstructing video data is almost the reverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 may reconstruct the residual blocks associated with the TU of the current CU by performing an inverse transform on the coefficient blocks associated with the TU of the current CU. The video decoder 30 also reconstructs the encoded blocks of the current CU by adding samples of the prediction blocks of the PU of the current CU to the corresponding samples of the transform blocks of the TU of the current CU. After reconstructing the encoded blocks for each CU of the frame, the video decoder 30 may reconstruct the frame.
[0069] As mentioned above, video coding primarily uses two modes to achieve video compression: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that IBC can be considered an intra-frame prediction or a third mode. Of the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block from a reference video block.
[0070] However, as video data acquisition technology is constantly improving and the video block size for preserving details in video data is becoming finer, the amount of data required to represent the motion vector of the current frame has also increased significantly. One way to overcome this challenge is to benefit from the fact that groups of adjacent CUs in both the spatial and temporal domains not only have similar video data for predictive purposes, but the motion vectors between these adjacent CUs are also similar. Therefore, it is possible to use the motion information of spatially adjacent CUs and / or CUs that are in the same temporal location as an approximation of the motion information (e.g., motion vector) of the current CU by examining their spatial and temporal correlations, which is also called a "motion vector predictor (MVP)" for the current CU.
[0071] In relation to Figure 3, as explained above, instead of encoding the actual motion vector of the current CU determined by the motion estimation unit 42 into the video bitstream, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to create the Motion Vector Difference (MVD) of the current CU. Doing so eliminates the need to encode the motion vector determined for each CU in a frame by the motion estimation unit 42 into the video bitstream, which can significantly reduce the amount of data used to represent motion information in the video bitstream.
[0072] Similar to the process of selecting a prediction block in a reference frame during interframe prediction of a code block, a list of motion vector candidates for the current CU (also called a "merge list") must be constructed using potential candidate motion vectors associated with spatially adjacent CUs and / or CUs at the same temporal position as the current CU. A set of rules must then be employed by both the video encoder 20 and the video decoder 30 to select one element from the motion vector candidate list as the motion vector predictor for the current CU. In this way, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and it is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU, using only the index of the selected motion vector predictor in the motion vector candidate list.
[0073] Location-dependent intra-prediction combinations
[0074] In VVC, the intra-prediction results for DC mode, planar mode, and some angular modes are further modified by a position-dependent intra-prediction combination (PDPC) method. PDPC is an intra-prediction method that calls a combination of HEVC-style intra-prediction with boundary reference samples and filtered boundary reference samples. PDPC is applied without signal transmission to the following intra-modes, namely planar, DC, intra-angles below horizontal, and intra-angles above vertical and below 80 degrees. PDPC is not applied if the current block is in Bdpcm mode or if the MRL index is greater than 0.
[0075] Using a linear combination of intra-prediction modes (DC, plane, angle) and reference samples, the following equation is used: pred(x',y')=Clip(0,(1<<BitDepth)-1,(wL×R-1,y’+wT×Rx’,-1+(64-wL-wT)×pred(x’,y’)+32)> >6) According to this, the predicted sample pred(x',y') is predicted, where Rx,-1 and R-1,y represent the reference samples located above and to the left of the current sample (x,y), respectively.
[0076] When PDPC is applied to DC, planar, horizontal, and vertical intra-modes, the additional boundary filters required for HEVC DC mode boundary filters or horizontal / vertical mode edge filters are not necessary. The PDPC process for DC mode and planar mode is identical. For angular mode, if the current angular mode is HOR_IDX or VER_IDX, the left or top reference sample is not used, respectively. The weights and scale factors of PDPC depend on the prediction mode and block size. PDPC is applied to blocks where both width and height are 4 or greater.
[0077] Figures 5A to 5D show the definitions of PDPC reference samples (Rx,-1 and R-1,y) applied to various prediction modes. Figure 5A shows an example of the diagonal upper right mode. Figure 5B shows an example of the diagonal lower left mode. Figure 5C shows an example of the adjacent diagonal upper right mode. Figure 5D shows an example of the adjacent diagonal lower left mode. The prediction sample pred(x',y') is located at (x',y') within the prediction block. For example, in the diagonal mode, the coordinate x of the reference sample Rx,-1 is given by x=x'+y'+1, and similarly the coordinate y of the reference sample R-1,y is given by y=x'+y'+1. In other angular modes, the reference samples Rx,-1 and R-1,y may be placed at fractional sample positions. In this case, the sample value of the nearest integer sample position is used.
[0078] As previously mentioned, intra-prediction samples are generated from a set of adjacent reference samples that are either unfiltered or filtered, which can introduce discontinuities along the block boundaries between the current encoded block and its adjacent blocks. To address such issues, HEVC applies boundary filtering by utilizing a two-tap filter (for DC mode) or a gradient-based smoothing filter (for horizontal and vertical prediction modes) to combine the first row / column of prediction samples for DC prediction mode, horizontal prediction mode (i.e., mode 18), and vertical prediction mode (i.e., mode 50) with unfiltered reference samples.
[0079] Gradient PDPC
[0080] In VVC, PDPC may not be applied in some scenarios because quadratic reference samples are unavailable. A gradient-based PDPC, extended from the horizontal / vertical modes, is applied. The PDPC weights (wT / wL) and the nScale parameter, which determines the attenuation in the PDPC weights with respect to distance from the left / top boundary, are set to be equal to the corresponding parameters in the horizontal / vertical modes, respectively. Bilinear interpolation is applied when quadratic reference samples are at fractional sample positions.
[0081] Geometric partition mode (GPM)
[0082] VVC supports geometric segmentation modes for interpretation. Geometric segmentation modes are signaled as a special merge mode by a single CU level flag. In the current GPM design, a total of 64 segmentation modes are supported by the GPM modes for each possible CU size where both the width and height are between 8 and 64, excluding 8x64 and 64x8.
[0083] When this mode is used, the CU is split into two parts by geometrically arranged straight lines, as shown in Figure 6. The position of the split line is mathematically derived from the angle and offset parameters of a particular division. Each part of the geometric division within the CU is interpredicted using its own motion, and only uni-prediction is allowed for each division, i.e., each part has one motion vector and one reference index. As with conventional bi-prediction, a uni-prediction motion constraint is applied to ensure that only two motion-compensated predictions are required for each CU. When the geometric division mode is used for the current CU, the geometric division index (angle and offset) indicating the division mode of the geometric division and two merge indices (one for each division) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled at the sequence level.
[0084] Mixing along the edges of geometric divisions
[0085] After each geometric segment is acquired using its own motion, a mixture is applied to the two single-prediction signals to derive samples around the edges of the geometric segment. The mixture weight for each position in the CU is derived based on the distance from the individual sample position to the corresponding segment edge.
[0086] GPM signal transmission design
[0087] According to the current GPM design, the use of GPM is indicated by signaling a single flag at the CU level. The flag is signaled only if the current CU is encoded by merge mode or skip mode. Specifically, if the flag is equal to 1, the flag indicates that the current CU is predicted by GPM. Otherwise (if the flag is equal to zero), the CU is encoded by another merge mode, such as normal merge mode, merge mode with motion vector difference, combined interprediction, and intraprediction. If GPM is enabled for the current CU, one syntactic element, namely merge_gpm_partition_idx, is further signaled to indicate the applied geometric partition mode (specifying the direction and offset of the line from the CU center that splits the CU into two partitions, as shown in Figure 6). Then, two syntactic elements, merge_gpm_idx0 and merge_gpm_idx1, are signaled to indicate the indices of the single-prediction merge candidates used for the first and second GPM partitions. More specifically, these two syntactic elements are used to determine the unidirectional MV of two GPM segments from a single predictive merge list, as described in the section "Constructing a Single Predictive Merge List". According to the current GPM design, the two indices cannot be the same in order to make the two unidirectional MVs more distinct. Based on such prior knowledge, first, the single predictive merge index of the first GPM segment is signaled and used as a predictor to reduce the signaling overhead of the single predictive merge index of the second GPM segment. In detail, if the second single predictive merge index is smaller than the first single predictive merge index, its original value is signaled directly. Otherwise (if the second single predictive merge index is larger than the first single predictive merge index), its value is subtracted by 1 before being signaled to the bitstream. On the decoder side, first, the first single predictive merge index is the decoder.Next, in decoding the second single predictive merge index, if the parsed value is less than the first single predictive merge index, the second single predictive merge index is set to be equal to the parsed value; otherwise (if the parsed value is greater than or equal to the first single predictive merge index), the second single predictive merge index is set to be equal to the parsed value plus 1. Table 1 shows the existing syntactic elements used in GPM mode in the current VVC specification.
[0088] [Table 1]
[0089] On the other hand, the current GPM design uses a truncated unary code for binarizing two single predictive merge indices, namely merge_gpm_idx0 and merge_gpm_idx1. Furthermore, since the two single predictive merge indices cannot be the same, different maximum values are used for truncation of the codewords of the two single predictive merge indices, and these maximum values are set to be equal to MaxGPMMergeCand-1 and MaxGPMMergeCand-2 for merge_gpm_idx0 and merge_gpm_idx1, respectively. MaxGPMMergeCand is the number of candidates in the single predictive merge list.
[0090] When GPM / AWP mode is applied, two different binarization methods are used to convert the syntax merge_gpm_partition_idx into a binary bit string. Specifically, the syntax elements are binarized by fixed-length codes and truncated binary codes, respectively, in the VVC and AVS3 standards. On the other hand, in AVS3's AWP mode, a different maximum value is used for binarization.
[0091] Spatial angular weighted prediction (SAWP)
[0092] AVS offers a Spatial Angular Weighted Prediction (SAWP) mode that extends the GPM mode to intra-blocks. In SAWP mode, two intra-prediction blocks are weighted instead of two intra-prediction blocks. The two intra-prediction blocks are predicted using two different intra-prediction modes selected from the intra-prediction modes. The intra-prediction modes are selected from angle modes 5 to 30. The maximum size is 32x32. Two most probable modes (MPMs) from the normal intra-modes are used for the MPM derivation of SAWP mode.
[0093] Multi-direction intra-predictive design (MDIP) follows the same design philosophy as SAWP, but with some slight differences in specific design details.
[0094] Decoder-side intra-mode derivation (DIMD)
[0095] DIMD is an intra-encoding tool in which the luminance intra-prediction mode (IPM) is not transmitted over the bitstream. Instead, the IPM is derived using previously encoded / decoded pixels in the same manner in the encoder and decoder. The DIMD method performs texture gradient processing to derive two optimal modes. These two modes and the planar mode are then applied to a block, and their predictors are weighted averaged. The DIMD selection result is signaled in the bitstream of the intra-encoded block using a flag. In the decoder, if the DIMD flag is true, the intra-prediction mode is derived in the reconstruction process using the same previously encoded neighboring pixels. If it is not true, the intra-prediction mode is parsed from the bitstream, similar to a classical intra-encoding mode.
[0096] To derive the intra-prediction mode of a block, we must first select a set of adjacent pixels on which gradient analysis will be performed. For the purpose of normativity, these pixels should be in the pool of decoded / reconstructed pixels. As shown in Figure 7, we select a template that encloses the current block by only T pixels to the left and T pixels above it. Next, we perform gradient analysis on the pixels of the template. This allows us to determine the major angular direction of the template, assuming that it is likely to be identical to one of the current blocks (a core assumption of this method). Thus, the following matrix is convolved with the template,
number
[0097] For each pixel in the template, each of these two matrices is multiplied by a 3x3 window centered on the current pixel and consisting of its eight immediate neighbors, and the results are summed. Thus, two values Gx (from multiplication with Mx) and Gy (from multiplication with My) are obtained, corresponding to the horizontal and vertical gradients at the current pixel, respectively.
[0098] Figure 8 shows the convolution process. Blue pixels are the current pixels. Red pixels (including blue) are pixels for which gradient analysis is possible. Gray pixels are pixels for which gradient analysis is not possible because there are no neighbors. Purple pixels are available (reconstructed) pixels outside the considered template and are used for gradient analysis of red pixels. If a purple pixel is unavailable (for example, because a block is too close to the picture boundary), gradient analysis is not performed for all red pixels that use this purple pixel. For each red pixel, the gradient intensity (G) and orientation (O) are calculated using Gx and Gy.
number
[0099] Next, the gradient orientation is converted to an intra-angle prediction mode and used to index the histogram (initialized to zero). The histogram value in that intra-angle mode increases by G. Once all red pixels in the template have been processed, the histogram will contain the cumulative gradient intensity for each intra-angle mode. For the current block, the IPM corresponding to the two highest histogram bars is selected. If the maximum value in the histogram is 0 (meaning that gradient analysis could not be performed or the area making up the template is flat), the DC mode is selected as the intra-prediction mode for the current block.
[0100] The two IPMs corresponding to the two highest HoG bars are combined with the planar mode. In one or more examples, predictive fusion is applied as a weighted average of the three predictors described above. For this purpose, the weights for the planar mode are fixed at 21 / 64 (approximately 1 / 3). In this case, the remaining weights 43 / 64 (approximately 2 / 3) are shared between the two HoG IPMs in proportion to the amplitude of the HoG bars. Figure 9 visualizes this process.
[0101] The derived intra-mode is included in the primary list of intra-most likely modes (MPMs), and therefore the DIMD process is executed before the MPM list is constructed. The primary derived intra-mode of a DIMD block is stored with the block and used for constructing the MPM lists of adjacent blocks.
[0102] Template-based intra-mode derivation (TIMD)
[0103] For each intra-mode in MPM, the sum of absolute transformed differences (SATD) between the predicted and reconstructed samples in the template region shown in Figure 10 is calculated, and the intra-modes with the first two modes having the smallest SATD costs are selected. These are then fused with weights, and such weighted intra-predictions are used to encode the current CU.
[0104] The costs of the two selected modes are compared to a threshold, and in the test, a cost factor of 2 is applied as follows: costMode2 < 2 * costMode1
[0105] If this condition is true, fusion is applied; otherwise, only mode1 is used.
[0106] The weights of the modes are calculated from their SATD costs as follows: weight1=costMode2 / (costMode1+costMode2) weight2 = 1 - weight1
[0107] While DIMD mode can improve intra-predictive efficiency, there is room for further performance enhancement. Meanwhile, some parts of existing DIMD mode need to be simplified for more efficient codec hardware implementation or improved for better encoding efficiency. Furthermore, the trade-off between implementation complexity and the benefits of its encoding efficiency needs further refinement.
[0108] Following the final decision on VVC, the JVET group continued to explore compression efficiencies exceeding VVC. JVET maintained a single reference software called the Extended Compression Model (ECM) by integrating several additional coding tools on top of the VVC Test Model (VTM). In the current ECM, PDPC is used depending on the intra-mode. For DIMD modes, PDPC is used depending on each intra-mode. As shown in Figure 11D, two different positions of the PDPC scheme are used and applied to each intra-mode in DIMD mode. For intra-predictions using angular mode in DIMD mode, PDPC is applied before predictive fusion. For intra-predictions using DC mode or planar mode in DIMD mode, PDPC is applied after predictive fusion. Such a non-unified design may not be optimal from a standardization standpoint.
[0109] Similarly, two different fusion designs are available, applicable to DIMD and TIMD, respectively. Each different fusion design is associated with different candidates and weight calculations. For blocks to which DIMD is applied, the two IPMs and planar modes corresponding to the two highest HoG bars are selected for fusion. The weight of the planar mode is fixed at 21 / 64 (approximately 1 / 3). In this case, the remaining weight of 43 / 64 (approximately 2 / 3) is shared between the two HoG IPMs in proportion to the amplitude of the HoG bars. For blocks to which TIMD is applied, the intra-modes having the first two modes with the smallest SATD costs are selected, and the mode weights are calculated from their SATD costs. Such non-unified designs may not be optimal from a standardization standpoint. In addition to the above, there is room to further improve its performance with various fusion methods.
[0110] In current ECM designs, intra-modes derived from DIMD are included in the primary list of intra-most likely modes (MPMs), regardless of whether the derived intra-mode is already used in DIMD. There is room for further improvement in this performance.
[0111] Existing DIMD and TIMD designs involve multiple floating-point operations (including addition, multiplication, and division) to calculate the parameters used to derive the optimal intra-prediction mode and generate the corresponding prediction samples for a single current DIMD / TIMD coded block. Specifically, the following floating-point operations are applied in existing DIMD and TIMD designs in ECM:
[0112] 1) Derivation of Gradient Orientation in DIMD: As previously explained, in DIMD mode, two optimal intra-prediction modes are selected based on an analysis of the histogram of the gradient (HoG) of adjacent reconstructed samples (i.e., templates) above and to the left of the current block. During such analysis, the gradient orientation of each template sample needs to be calculated, and this gradient orientation is further converted into one of the existing angular intra-prediction directions. In ECM, to calculate such orientation based on horizontal and vertical gradients, a pair of floating-point division and multiplication is applied to each template sample, i.e.,
number
[0113] 2) Mixing of Prediction Samples in DIMD: In existing DIMD designs, prediction samples generated using two angle-intra prediction modes with the largest and second largest gradient histogram amplitudes are mixed with the planar mode prediction samples to form the final prediction samples for the current block. Furthermore, the weights of the two angle-intra predictions are determined based on the histogram amplitudes of their gradients.
number
number
[0114] 3) Mixing of Predicted Samples in TIMD: In existing TIMD designs, if the SATDs of two selected intra-modes are sufficiently close, the predicted samples generated from the two intra-modes are mixed together to produce the final predicted samples for the current block's intra-modes. According to the current design, the weights applied to the two intra-modes are calculated according to their respective SATD values.
number
[0115] All of the above floating-point operations are extremely costly to implement in actual codecs, both in hardware and software.
[0116] This disclosure provides methods for simplifying and / or further improving existing designs of DIMD modes in order to address previously identified issues. In general, the main features of the technology proposed in this disclosure can be summarized as follows:
[0117] 1) Unify the PDPC used under angular mode and DC / planar mode in DIMD mode by applying PDPC to all intra-predictions before predictive fusion. An example of such a method is shown in the block diagram in Figure 12A.
[0118] 2) Unify the PDPC used under angular mode and DC / planar mode in DIMD mode by applying PDPC to all intra-predictions after predictive fusion. An example of such a method is shown in the block diagram in Figure 12B.
[0119] 3) Unify the PDPC used under angular mode and DC / planar mode in DIMD mode by disabling PDPC for all intra predictions in DIMD mode. An example of such a method is shown in the block diagram in Figure 12C.
[0120] 4) Unify the PDPC used under angular mode and DC / planar mode in TIMD or DIMD mode by disabling the PDPC for DC / planar intra-prediction in TIMD or DIMD mode. For example, as shown in Figure 11E, in response to a determination that DC mode or planar mode is applied, the encoder or decoder may disable the PDPC for DC mode or planar mode after fusion in TIMD. In another example, as shown in Figure 12C, in response to a determination that DC mode or planar mode is applied, the encoder or decoder may disable the PDPC for DC mode or planar mode before fusion in DIMD. In yet another example, in Figure 12D, in response to a determination that planar mode is applied in the first intra-prediction mode and the second intra-prediction mode, the encoder or decoder may disable the PDPC for planar intra-prediction before fusion in DIMD mode.
[0121] 5) Unify the PDPC used under the angular mode and DC / planar mode in TIMD or DIMD mode by disabling the PDPC for angular intra-prediction in TIMD or DIMD mode. For example, as shown in Figure 12C, in response to a determination that the angular mode is applied to the first intra-prediction mode and the second intra-prediction mode, the encoder or decoder may disable the PDPC for angular intra-prediction before fusion in DIMD mode. In another example, as shown in Figure 11F, in response to a determination that the angular mode is applied, the encoder or decoder disables the PDPC for angular intra-prediction before fusion in TIMD.
[0122] 6) Unify the fusion method used in DIMD mode and the fusion method used under TIMD mode by applying the fusion method used under DIMD mode to TIMD mode.
[0123] 7) Unify the fusion methods used in DIMD mode and TIMD mode by applying the fusion method used in TIMD mode to DIMD mode.
[0124] 8) The fusion method used under DIMD mode and TIMD mode is unified by transmitting the result of the fusion method selection.
[0125] 9) Derive the intra-mode from DIMD into a list of intra-most likely modes (MPMs), taking into account whether the derived intra-mode is already in use in DIMD.
[0126] 10) Derive the intra-mode from the TIMD into the list of intra-most likely modes (MPM).
[0127] It should be noted that the proposed method may also be applicable to other intra predictive coding modes such as TIMD / MDIP by an encoder or decoder. Another set of examples of its application to the TIMD mode is shown in the block diagrams of Figures 11A-11C. Figure 11A shows an example where all PDPC processes are applied before the TIMD fusion process. Figure 11B shows an example where all PDPC processes are applied after the TIMD fusion process. Figure 11C shows an example where all PDPC processes are disabled in TIMD.
[0128] It should be noted that the proposed method may also be applicable to other combined inter and intra prediction coding modes, such as combined inter and intra prediction (CIIP).
[0129] Please note that the disclosed methods may be applied individually or in combination.
[0130] Harmonization of PDPC used in angular mode and DC / planar mode in DIMD
[0131] According to one or more embodiments of this disclosure, the same PDPC position is applied to both the angular mode and the DC / planar mode under DIMD mode. Various methods may be used to achieve this goal.
[0132] In one example of this disclosure, as shown in Figure 12A, it is proposed to apply PDPC calculations before predictive fusion in DIMD mode. In other words, each intra-predictive mode is applied to the PDPC based on its intra-mode before predictive fusion in DIMD mode. The proposed method may also be applied to other intra-predictive coding modes such as TIMD.
[0133] In another example of this disclosure, as shown in Figure 12B, it is proposed to apply a PDPC operation after predictive fusion in DIMD mode. In other words, a weighted combination of three predictors is applied to the PDPC based on a specific mode, e.g., DC mode, planar mode. In one example, the specific mode is planar mode, and then a PDPC with planar mode is applied after predictive fusion in DIMD mode. In another example, the IPM corresponding to the highest histogram bar is selected as the specific mode, and then a PDPC with the specific mode is applied after predictive fusion in DIMD mode. In yet another example, the IPM corresponding to the second highest histogram bar is selected as the specific mode, and then a PDPC of the specific mode is applied after predictive fusion in DIMD mode.
[0134] Another example of this disclosure proposes disabling PDPC calculations in DIMD mode. In other words, PDPC calculations are not used in DIMD mode, as shown in Figure 12C.
[0135] Another example of this disclosure proposes disabling PDPC calculations for DC / planar intra-prediction in DIMD mode. In other words, PDPC calculations are not used for DC / planar intra-prediction in DIMD mode. In one example, as shown in Figure 15, in step 1502, the encoder or decoder may determine the predicted sample values for one or more video blocks based on PDPC calculations for one or more intra-predictions of one or more video blocks in order to unify the position-dependent intra-prediction combination (PDPC) calculations in intra-prediction coding mode, the PDPC calculations modify the results of one or more intra-predictions based on a combination of boundary reference samples. In step 1504, the encoder or decoder may disable PDPC calculations for DC mode or planar mode in response to a determination that DC mode or planar mode is applied to one or more intra-predictions of one or more video blocks in order to unify the PDPC calculations in intra-prediction coding mode.
[0136] In yet another example of this disclosure, it is proposed to disable the PDPC operation for angular intra-prediction in DIMD mode. In other words, the PDPC operation is not used for angular intra-prediction in DIMD mode. In an example as shown in Figure 16, in step 1602, the encoder or decoder may determine the predicted sample values of one or more video blocks based on the PDPC operation for one or more intra-predictions of one or more video blocks in order to unify the position-dependent intra-prediction combination (PDPC) operation in intra-prediction coding mode, the PDPC operation modifies the result of one or more intra-predictions based on a combination of boundary reference samples. In step 1604, the encoder or decoder may disable the PDPC operation for angular mode in order to unify the PDPC operation in intra-prediction coding mode in order to unify the angular mode in one or more intra-predictions of one or more video blocks.
[0137] It should be noted that the proposed method may also be applicable to other intra predictive coding modes such as TIMD / MDIP.
[0138] Harmonization of fusion schemes used in DIMD mode and TIMD mode
[0139] According to one or more embodiments of this disclosure, the same fusion scheme is applied to both DIMD mode and TIMD mode. Various methods may be used to achieve this goal. The fusion scheme is applied as a weighted average of the predictors in DIMD mode and TIMD mode.
[0140] In one example of this disclosure, it is proposed to apply the fusion scheme used under the DIMD mode to the TIMD mode. In other words, in the TIMD mode, the first two modes having the smallest SATD cost and planar mode are selected as predictors for fusion, and a weighted average of the predictors is calculated. The weight of the planar mode is fixed at 21 / 64 (approximately 1 / 3). In this case, the remaining weight 43 / 64 (approximately 2 / 3) is shared between the other two modes in proportion to the amplitude of the SATD cost.
[0141] In another example of this disclosure, it is proposed to apply the fusion scheme used under TIMD mode to DIMD mode. In other words, in DIMD mode, the first two modes with the highest HoG bars are selected as predictors for fusion, and the mode weights are calculated from HoG IPM in proportion to the amplitude of the HoG bars. If the maximum value in the histogram is 0 (meaning that gradient analysis could not be performed or the area constituting the template is flat), one default mode, e.g., DC mode, planar mode, is selected as the intra-prediction mode for the current block.
[0142] In yet another example of this disclosure, it is proposed to signal the selection result of a fusion scheme in TIMD and / or DIMD modes. In one example, for a given CU, a flag is signaled to the decoder to indicate whether or not the block uses DIMD mode. If encoded using DIMD mode, one more flag is signaled to the decoder to indicate, for example, which fusion scheme is used as the first or second fusion method described above.
[0143] Corrects the DIMD mode used within the MPM list.
[0144] In another aspect of the present disclosure, it is proposed to derive an intra-mode from DIMD into a list of intra-most likely modes (MPMs), depending on whether the derived intra-mode is already used in DIMD. According to one or more embodiments of the present disclosure, if the fusion scheme is used in DIMD mode, the intra-mode derived from DIMD may be used as a candidate for the MPM list. In other words, if the fusion scheme is not used in DIMD mode, the intra-mode derived from DIMD cannot be used as a candidate for the MPM list.
[0145] In other aspects of this disclosure, a directional mode to which an offset from the available directional modes of DIMD has been added may be used as a candidate for the MPM list. In a particular example, the offset may be 1, -1, 2, -2, 3, -3, 4, -4.
[0146] As shown in Figure 17, in one example, in step 1702, the encoder or decoder may determine whether a fusion scheme is applied in DIMD mode, which is applied as a weighted average of predictors in DIMD mode in step 1702. In step 1704, the encoder or decoder may obtain an offset directional mode by applying an offset to the available directional modes in DIMD mode. In step 1706, the decoder may determine whether the offset directional mode should be added to the list of most likely modes (MPMs) based on whether a fusion scheme is applied in DIMD mode.
[0147] As an example, a general MPM list with 22 entries is first constructed, then the first six entries from this general MPM list are included in the primary MPM (PMPM) list, and the remaining entries form the secondary MPM (SMPM) list. The first entry in the general MPM list is the planar mode. The remaining entries consist of intra-modes for adjacent blocks left (L), up (A), bottom left (BL), top right (AR), and top left (AL), DIMD mode (blue section), directional mode with an offset added from the first two available directional modes and DIMD mode (red section) for adjacent blocks, and default mode {DC_IDX(1), VER_IDX(50), HOR_IDX(18), VER_IDX-4(46), VER_IDX+4(54), 14, 22, 42, 58, 10, 26, 38, 62, 6, 30, 34, 66, 2, 48, 52, 16}.
[0148] If the CU blocks are oriented vertically, the order of adjacent blocks is A, L, BL, AR, AL; otherwise, the order is L, A, BL, AR, AL.
[0149] In this example, DIMD modes without an offset (blue section) are added to the MPM list first. If the list is not full, DIMD modes with an offset (red section) are added to the MPM list.
[0150] TIMD mode used in MPM list
[0151] In another aspect of this disclosure, it is proposed to derive intra-modes from TIMD to a list of intra-most probable modes (MPMs). Generally, VVCs have 67 intra-predictive modes, including non-directional modes (planar, DC) and 65 angular modes, which efficiently model the various directional structures typically present in video and image content. In one or more embodiments of this disclosure, intra-modes derived from TIMD may be used as candidates for the MPM list. In one example, intra-modes derived from DIMD cannot be used as candidates for the MPM list, but intra-modes derived from TIMD may be used as candidates for the MPM list.
[0152] In another aspect of the present disclosure, it is proposed to derive an intra-mode from TIMD into a list of intra-most likely modes (MPMs), depending on whether the derived intra-mode is already used in TIMD. According to one or more embodiments of the present disclosure, if the fusion scheme is used in TIMD mode, the intra-mode derived from TIMD may be used as a candidate for the MPM list. In other words, if the fusion scheme is not used in TIMD mode, the intra-mode derived from TIMD cannot be used as a candidate for the MPM list.
[0153] The above methods may be implemented using a device that includes one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The device may use the circuits in combination with other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit disclosed above may be at least partially implemented using one or more circuits.
[0154] Simplified DIMD and TIMD using integer arithmetic
[0155] As previously noted, existing DIMD and TIMD designs involve several floating-point operations (i.e., addition, multiplication, and division) to derive DIMD / TIMD parameters, which are unacceptable for actual codec implementations in both software and hardware. This section proposes a look-up table (LUT) based scheme that simplifies DIMD and TIMD implementations by replacing all floating-point operations with integer addition and multiplication. In one example, the decoder identifies the floating-point division operations to be performed to derive parameters in DIMD or TIMD mode, and the decoder obtains the parameters in DIMD or TIMD by replacing the floating-point division operations with integer addition and multiplication based on the look-up table (LUT).
[0156] Specifically, as shown in Figure 13, an integer L is given by one exponent and K bits of significant part (including the K most significant bits (MSB) after the exponent).
number
number
[0157] If the fractional part 1 / L is quantized with M-bit precision, the above equation becomes:
number
number
number
[0158] In response to this, the proposed integerization scheme can be achieved as follows, with respect to division between any two integers, for example
number
[0159] In practice, various combinations of LUT size (i.e., K) and parameter precision (i.e., M) may be applied to achieve various trade-offs between the accuracy of the derived parameters and the complexity of the implementation. For example, using a larger LUT size and higher parameter precision is beneficial in maintaining high parameter accuracy, but comes with the burden of increased storage size to maintain the LUT and increased bit depth to perform corresponding integer operations (e.g., integer multiplication, addition, and bitwise shifts). Based on such considerations, in a particular example, it is proposed to set the values of K and M to 4. Based on such settings, the corresponding LUT is: divLut[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,0} This is derived as follows.
[0160] Using the DIMD derivation as an example, in the case of the above integerization method (when both K and M are set to 4), the derivation of the gradient orientation for each template sample is exponent = Floor(Log2(G x )) norm MSB = ((G x ≪ 4)≫ exponent)&15 v = G y ·(divLut[norm MSB |8) s = exponent+(norm MSB ≠ 0?4:3)-16
Number
[0161] Note that the values of K and M used in the above example are for illustrative purposes only. In practice, a proposed LUT - based method using different values of K and M may be applied to convert floating - point division to integer operations in the process of other future coding techniques.
[0162] Figure 14 shows a computing environment 1610 coupled to a user interface 1650. The computing environment 1610 may be part of a data - processing server. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.
[0163] The processor 1620 typically controls the overall operation of the computing environment 1610, including operations related to display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors for executing instructions that request all or some of the steps in the manner described above. Furthermore, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.
[0164] Memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. Memory 1630 may include certain software 1632. Examples of such data include instructions for any application or method running on the computing environment 1610, image datasets, image data, etc. Memory 1630 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0165] The I / O interface 1640 provides an interface between the processor 1620 and peripheral interface modules such as a keyboard, click wheel, and buttons. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. The I / O interface 1640 may be coupled with an encoder and a decoder.
[0166] In one embodiment, a non-transient computer-readable storage medium is also provided, for example, in memory 1630, which contains multiple programs for performing the methods described above, which can be executed by a processor 1620 in a computing environment 1610. Alternatively, the non-transient computer-readable storage medium may store a bitstream or datastream containing encoded video information (e.g., video information including one or more syntactic elements) generated by an encoder (e.g., video encoder 20 in Figure 2) using the encoding method described above, for use by a decoder (e.g., video decoder 30 in Figure 3) when decoding video data. The non-transient computer-readable storage medium may be, for example, ROM, random-access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0167] In one embodiment, a computing device is also provided comprising one or more processors (e.g., processor 1620) and a non-transient computer-readable storage medium or memory 1630 storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the method described above when executing the plurality of programs.
[0168] In one embodiment, a computer program product is also provided which includes multiple programs in, for example, memory 1630, that can be executed by a processor 1620 in a computing environment 1610 to perform the method described above. For example, the computer program product may include a non-transient computer-readable storage medium.
[0169] In one embodiment, the computing environment 1610 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0170] The descriptions in this disclosure are provided for illustrative purposes only and are not intended to exhaust or limit the scope of this disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and the accompanying drawings.
[0171] Unless otherwise specified, the order of the steps in the methods provided herein is for illustrative purposes only, and the steps in the methods provided herein are not limited to the order specifically described above and may be modified in accordance with the actual conditions. Furthermore, at least one of the steps in the methods provided herein may be adjusted, combined, or deleted in accordance with the actual requirements.
[0172] The examples are selected and illustrated to illustrate the principles of this disclosure, to enable those skilled in the art to understand this disclosure in various implementations, and to maximize the use of the underlying principles and various implementations by making various modifications to suit specific intended uses. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed implementations, and that modifications and other implementations are intended to be included within the scope of this disclosure.
Claims
1. A method for decoding video, The decoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The decoder determines the prediction sample value of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, the PDPC operation modifies the result of the one or more intra-predictions based on a combination of boundary reference samples, The decoder includes the step of disabling the PDPC calculation for the DC mode or the planar mode in response to a determination that the DC mode or planar mode is applied in the plurality of intra predictions. method.
2. The method for video decoding according to claim 1, wherein the intra predictive coding mode includes a template-based intra mode derivation (TIMD) mode, a decoder-side intra mode derivation (DIMD) mode, or a multidirectional intra predictive (MDIP) mode.
3. The method for video decoding according to claim 2, wherein the decoder disables the PDPC calculation for the DC mode or the planar mode after the fusion in the TIMD mode.
4. The method for video decoding according to claim 2, wherein the decoder disables the PDPC calculation for the planar mode before the fusion in the DIMD mode.
5. A method for decoding video, The decoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The decoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, A method comprising the step of disabling the PDPC calculation for the angular mode in response to the decoder determining that the angular mode is applied to the plurality of intra predictions.
6. The method for video decoding according to claim 5, wherein the intra predictive coding mode includes a template-based intra mode derivation (TIMD) mode, a decoder-side intra mode derivation (DIMD) mode, or a multidirectional intra predictive (MDIP) mode.
7. The method for video decoding according to claim 6, wherein the decoder disables the PDPC calculation for the angle mode before the fusion in the DIMD mode.
8. A method for video encoding, The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The encoder includes the step of disabling the PDPC calculation for the DC mode or the planar mode in response to a determination by the encoder that a DC mode or planar mode is applied to the plurality of intra predictions. method.
9. The method for video coding according to claim 8, wherein the intra predictive coding mode includes a template-based intra mode derivation (TIMD) mode, a decoder-side intra mode derivation (DIMD) mode, or a multidirectional intra predictive (MDIP) mode.
10. The method for video coding according to claim 9, wherein the encoder disables the PDPC operation for the DC mode or the planar mode after the fusion in the TIMD mode.
11. The method for video coding according to claim 9, wherein the encoder disables the PDPC calculation for the planar mode before the fusion in the DIMD mode.
12. A method for video encoding, The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The encoder includes the step of disabling the PDPC calculation for the angular mode in response to a determination that the angular mode is applied in the plurality of intra predictions. method.
13. The method for video coding according to claim 12, wherein the intra predictive coding mode includes a template-based intra mode derivation (TIMD) mode, a decoder-side intra mode derivation (DIMD) mode, or a multidirectional intra predictive (MDIP) mode.
14. The method for video coding according to claim 13, wherein the encoder disables the PDPC calculation for the angular mode before the fusion in the DIMD mode.
15. One or more processors, A memory configured to store instructions that can be executed by the one or more processors An apparatus comprising, wherein one or more processors are configured to execute the method described in any one of claims 1 to 14 when executing the instruction.
16. A non-transient computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause one or more computer processors to perform the method described in any one of claims 1 to 14.
17. A computer program that stores computer executable instructions, and when executed, causes one or more processors to execute the method described in any one of claims 1 to 14.
18. A method for transmitting a bitstream, A step of generating a bitstream by performing an encoding method, wherein the encoding method is: The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The generating step includes, in response to the encoder's determination that a DC mode or a planar mode is applied to the plurality of intra predictions, disabling the PDPC calculation for the DC mode or the planar mode; The step of transmitting the bitstream, The bitstream is decrypted by a decoding method, and the decoding method is The decoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The decoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, A transmission step including the step of disabling the PDPC calculation for the DC mode or the planar mode in response to the decoder's determination that the DC mode or planar mode is applied to the plurality of intra predictions, Methods that include...
19. A method for transmitting a bitstream, A step of generating a bitstream by performing an encoding method, wherein the encoding method is: The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The generating step includes the step of disabling the PDPC calculation for the angular mode in response to the encoder's determination that the angular mode is applied in the plurality of intra predictions, The step of transmitting the bitstream, The bitstream is decrypted by a decoding method, and the decoding method is The decoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The decoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The steps of transmitting include, in response to the decoder's determination that an angular mode is applied in the plurality of intra predictions, disabling the PDPC calculation for the angular mode, Methods that include...
20. A method for transmitting a bitstream, The steps include: executing an encoding method to generate a bitstream, The step of transmitting the bitstream includes, The aforementioned encoding The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The encoder includes the step of disabling the PDPC calculation for the DC mode or the planar mode in response to a determination by the encoder that a DC mode or planar mode is applied to the plurality of intra predictions. method.
21. A method for transmitting a bitstream, The steps include: executing an encoding method to generate a bitstream, The step of transmitting the bitstream includes, The aforementioned encoding The encoder determines that an intra-predictive coding mode, which fuses multiple intra-predictions of the video block, is applied to the video block. The encoder determines the predicted sample values of the video block based on a position-dependent intra-prediction combination (PDPC) operation for one or more of the plurality of intra-predictions, wherein the PDPC operation modifies the results of the one or more intra-predictions based on a combination of boundary reference samples, The encoder includes the step of disabling the PDPC calculation for the angular mode in response to a determination that the angular mode is applied in the plurality of intra predictions. method.
Citation Information
Patent Citations
Method, apparatus, and computer program for multi-line intra-frame prediction
JP2021521721A
Method and system for decoder-side intra mode derivation for block-based video coding
US20190166370A1
Constrained position dependent intra prediction combination (PDPC)
US20200267397A1
Encoding device, decoding device, encoding method, and decoding method
WO2020116402A1