Method and apparatus for intra prediction using interpolation filter
By selecting an interpolation filter based on sub-pixel offset and intra-prediction mode to determine the dominant reference side, the method addresses inefficiencies in existing video coding standards, improving memory and coding efficiency in video encoding and decoding processes.
Patent Information
- Application Number
- JP2025081366
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-11-07
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2039-10-07
AI Technical Summary
Existing video coding standards, such as HEVC and VVC, face challenges in intra-prediction modes that are complex and non-adaptive, leading to inefficient memory usage and coding complexity, particularly due to fixed and non-adaptive index lists for intra-modes.
The method involves selecting a sub-pixel interpolation filter based on the sub-pixel offset and intra-prediction mode to determine the size of the dominant reference side, optimizing memory usage by determining the size of the dominant reference side used in intra-prediction processes, thereby reducing unnecessary memory requirements.
This approach enhances memory efficiency and coding efficiency by optimizing memory usage and simplifying the intra-prediction process, facilitating more efficient video encoding and decoding.
Smart Images

Figure 2025114819000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 742,300, filed October 6, 2018, U.S. Provisional Patent Application No. 62 / 744,096, filed October 10, 2018, U.S. Provisional Patent Application No. 62 / 753,055, filed October 30, 2018, and U.S. Provisional Patent Application No. 62 / 757,150, filed November 7, 2018. The above-mentioned patent applications are incorporated herein by reference in their entireties.
[0002] The present disclosure relates to the technical field of image and / or video coding and decoding, and in particular to a method and apparatus for directional intra prediction with reference sample processing matched to the length of an interpolation filter. [Background technology]
[0003] Since the introduction of DVD discs, digital video has been widely used. Before transmission, the video is encoded and transmitted using a transmission medium. The viewer receives the video and uses a viewing device to decode and display the video. Over the years, video quality has improved, for example, with greater resolution, color depth, and frame rate. This has resulted in larger data streams that are now typically transported over the Internet and mobile communication networks.
[0004] However, videos with higher resolution usually contain more information and therefore require more bandwidth. To reduce bandwidth requirements, video coding standards have been introduced that involve compressing the video. When the video is encoded, the bandwidth requirements (or corresponding memory requirements, in the case of storage) are reduced. Often, this reduction comes at the expense of quality. Therefore, video coding standards attempt to find a balance between bandwidth requirements and quality.
[0005] High Efficiency Video Coding (HEVC) is an example of a video coding standard commonly known to those skilled in the art. HEVC divides coding units (CUs) into prediction units (PUs) or transform units (TUs). The Versatile Video Coding (VVC) next-generation standard is the most recent collaborative video project of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standardization organizations, collaborating in a partnership called the Joint Video Exploration Team (JVET). VVC is also known as the ITU-T H.266 / Next Generation Video Coding (NGVC) standard. In VVC, the concept of multiple partition types, i.e., the separation of CUs, PUs, and TUs, is eliminated except for cases where CUs are too large for the maximum transform length, supporting greater flexibility for CU partition shapes.
[0006] The processing of these coding units (CUs) (also called blocks) depends on their size, spatial location, and the coding mode specified by the encoder. Coding modes can be classified into two groups according to the type of prediction: intra-prediction mode and inter-prediction mode. Intra-prediction mode uses samples from the same picture (also called frame or image) to generate reference samples for calculating prediction values for samples of the block being reconstructed. Intra-prediction is also called spatial prediction. Inter-prediction mode is designed for temporal prediction and predicts samples of a block of the current picture using reference samples from the previous or next picture.
[0007] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC29 / WG11) are considering the potential need for standardization of future video coding technologies with compression capabilities significantly exceeding those of the current HEVC standard (including its current and upcoming extensions for screen content coding and high dynamic range coding). The groups are collaborating on this exploration in a collaborative research effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by their experts in this area.
[0008] The VTM (Versatile Test Model) standard uses 35 intra-modes, while the BMS (Benchmark Set) uses 67 intra-modes.
[0009] The intra-mode coding scheme currently described in BMS is considered complex, and a drawback of the unselected mode set is that the index list is always constant and not adaptive based on the current block characteristics (e.g., relative to its neighboring block intra-modes). Summary of the Invention [Means for solving the problem]
[0010]
[0009] Embodiments of the present application are disclosed that provide an apparatus and method for intra prediction. The apparatus and method use a mapping process to simplify the calculation procedure for intra prediction so as to improve coding efficiency. The scope of protection is defined by the claims.
[0011] These and other objects are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, the description and the figures.
[0012] Particular embodiments are outlined in the accompanying independent claims, with further embodiments in the dependent claims.
[0013] According to a first aspect, the present invention relates to a method of video coding, the method being executed by an encoding device or a decoding device, the method comprising: performing an intra-prediction process for a block, such as a block comprising samples to be predicted or a block of prediction samples, in particular a luma block comprising luma samples to be predicted, wherein a sub-pixel interpolation filter is applied to reference samples (e.g., luminance reference samples) during the intra-prediction process for the block, or a sub-pixel interpolation filter is applied to reference samples (e.g., chrominance reference samples) during the intra-prediction process for the block; The sub-pixel interpolation filter is selected based on a sub-pixel offset, for example, a sub-pixel offset between a position of a reference sample and a position of a sample to be interpolated, or between a reference sample and a sample to be predicted; The size of the primary reference side used in the intra prediction process is determined according to the length of the sub-pixel interpolation filter and the intra prediction mode (e.g., the intra prediction mode among the set of available intra prediction modes) that results in the maximum value of the sub-pixel offset (e.g., the maximum non-integer value), and the primary reference side comprises reference samples.
[0014] A reference sample is a sample on which prediction (here, particularly intra-prediction) is performed. In other words, a reference sample is a sample outside a (current) block that is used to predict samples of the (current) block. The term "current block" refers to a target block on which processing, including prediction, is performed. For example, a reference sample is a sample that neighbors the block on one or more of its sides. In other words, the reference sample used to predict the current block may be included in a line of samples that is at least partially adjacent to and parallel to one or more block boundaries (sides).
[0015] The reference samples may be samples at integer sample positions or interpolated samples at sub-sample positions, e.g., non-integer positions. Integer sample positions may refer to actual sample positions in the image to be coded (encoded or decoded).
[0016] A reference side is a side of a block from which reference samples are used to predict samples of the block. A dominant reference side is a side of a block from which reference samples are taken (in some embodiments, there is only one side from which reference samples are taken). However, generally, a dominant reference side may refer to a side from which reference samples are primarily taken (e.g., from which most of the reference samples are taken, or from which reference samples for predicting most of the block samples are taken, etc.). A dominant reference side includes reference samples used to predict samples of a block. It may be advantageous for memory saving purposes if a dominant reference side consists of reference samples used to predict samples of a block, and if all of the reference samples used to predict samples of the block are included in the dominant reference side. However, the present disclosure is also generally applicable with dominant reference sides including reference samples used to predict a block. These may comprise reference samples used directly for prediction as well as reference samples used for filtering to obtain sub-samples that are then used for prediction of block samples.
[0017] In general, the reference samples of a current block comprise neighboring reconstructed samples of the current block. Thus, if the current block is a current chroma block, the chroma reference samples of the current chroma block comprise neighboring reconstructed samples of the current chroma block. Thus, if the current block is a current luma block, the luma reference samples of the current luma block comprise neighboring reconstructed samples of the current luma block.
[0018] It is understood that the maximum value of the sub-pixel offset determines the memory requirement. Therefore, by determining the size of the dominant reference side according to the present disclosure, the present disclosure facilitates memory efficiency in video coding using intra prediction. In other words, by determining the size of the dominant reference side used in the intra prediction process according to the first aspect described above, memory requirements can be reduced while providing (storing) reference samples for predicting blocks. This can then lead to more efficient implementation of intra prediction for image / video encoding and decoding.
[0019] In one possible implementation of such a method according to the first aspect, the interpolation filter is selected based on a sub-pixel offset between the position of the reference sample and the position of the prediction sample.
[0020] It is understood that predicted samples are interpolated samples in that they are based on the output of an interpolation process.
[0021] In one possible implementation of such a method according to the first aspect, The sub-pixel offset is determined based on a reference line (such as refIdx), or The sub-pixel offset is determined based on the intraPredAngle, which depends on the selected intra prediction mode, or The sub-pixel offset is determined based on the distance between a reference sample (such as a reference line) and the side of the block of prediction samples, i.e., from the reference sample (such as a reference line) to the side of the block of prediction samples.
[0022] In one possible implementation of such a method according to the first aspect, the maximum value of the sub-pixel offset is a maximum non-integer sub-pixel offset (such as a maximum fractional sub-pixel offset or a maximum non-integer value of the sub-pixel offset), and the size of the major reference side is: the integer part of the maximum non-integer sub-pixel offset; the size of a side of the block of prediction samples; It is chosen to be equal to the sum of a portion or the entire length of the interpolation filter (such as half the length of the interpolation filter).
[0023] One of the advantages of such a choice of the size of the primary reference side is the preparation (storing / buffering) of all samples required for intra prediction of the block and a reduction in the number of samples (stored / buffered) that are not used to predict (samples of) the block.
[0024] In one possible implementation of such a method according to the first aspect, If the intra prediction mode is greater than the vertical intra prediction mode (VER_IDX), then the side of the block of prediction samples is the width of the block of prediction samples; or If the intra prediction mode is less than the horizontal intra prediction mode (HOR_IDX), the side of the block of prediction samples is the height of the block of prediction samples.
[0025] For example, in FIG. 10, VER_IDX corresponds to vertical intra-prediction mode #50, and HOR_IDX corresponds to horizontal intra-prediction mode #18.
[0026] In one possible implementation of such a method according to the first aspect, the reference sample of the primary reference side having a position greater than twice the size of said block side is set to be equal to the sample located at twice the size of said size.
[0027] In other words, it is padding to the right by repeating pixels that fall beyond double the side length. Memory buffer sizes are preferably powers of 2, and it is better to use the last sample of a power-of-2 sized buffer (i.e., located at double that size) than to keep a buffer of a size that is not a power of 2.
[0028] In one possible implementation of such a method according to the first aspect, the size of the main reference side is: With the chief of the main block, A fractional or full interpolation filter length (such as the interpolation filter length or half the interpolation filter length) minus 1; The following two values M: Chief of the Block, It is determined as the maximum of the sum of the integer part of the maximum (or greatest) non-integer sub-pixel offset plus some portion or all of the length of the interpolation filter (e.g., half the length of the interpolation filter), or the integer part of the maximum (or greatest) non-integer sub-pixel offset plus some portion or all of the length of the interpolation filter (e.g., half the length of the interpolation filter) + 1.
[0029] One of the advantages of such a choice of the size of the primary reference side is that it reduces or even avoids the preparation (storing / buffering) of all samples required for intra prediction of the block, and the preparation (storing / buffering) of samples that are not used to predict (samples of) the block.
[0030] It should be noted that "block major side", "block side length", "block major side length", and "side size of a block of prediction samples" are the same concepts throughout this disclosure.
[0031] In one possible implementation of such a method according to the first aspect, when the maximum of the two values M is equal to the length of the block major side, no right padding is performed, or Right padding is performed when the maximum of the two values M is equal to the integer part of the largest non-integer sub-pixel offset plus half the length of the interpolation filter, or the integer part of the largest non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1.
[0032] In one possible implementation, padding is performed by repeating the first and / or last reference sample of the primary reference side on the left and / or right side, respectively, specifically as follows: denoting the primary reference side as ref and the size of the primary reference side as refS, the padding is expressed as ref[-1]=p[0] and / or ref[refS+1]=p[refS], where ref[-1] represents the left value of the primary reference side; p[0] represents the value of the first reference sample of the primary reference side; ref[refS+1] represents the right value of the primary reference side, p[refS] represents the value of the last reference sample of the primary reference side.
[0033] In other words, right padding may be performed by ref[refS+1]=p[refS]. Additionally or alternatively, left padding may be performed by ref[-1]=p[0].
[0034] In this way, the padding may facilitate the preparation of all samples required for prediction, taking into account the interpolation filtering as well.
[0035] In one possible implementation of such a method according to the first aspect, the filters used in the intra prediction process are finite impulse response filters, whose coefficients are fetched from a look-up table.
[0036] In one possible implementation of such a method according to the first aspect, the interpolation filter used in the intra prediction process is a 4-tap filter.
[0037] In one possible implementation of such a method according to the first aspect, the coefficients of the interpolation filter are as follows:
[0038] [Table 1]
[0039] The "Subpixel Offset" column is specified in 1 / 32 subpixel resolution, depending on the subpixel offset, such as the non-integer part of the subpixel offset, as in In other words, an interpolation filter (such as a subpixel interpolation filter) is represented by the coefficients in the table above.
[0040] In one possible implementation of such a method according to the first aspect, the coefficients of the interpolation filter are as follows:
[0041] [Table 2]
[0042] The "Subpixel Offset" column is specified in 1 / 32 subpixel resolution, depending on the subpixel offset, such as the non-integer part of the subpixel offset, as in In other words, an interpolation filter (such as a subpixel interpolation filter) is represented by the coefficients in the table above.
[0043] In one possible implementation of such a method according to the first aspect, the coefficients of the interpolation filter are:
[0044] [Table 3]
[0045] The non-integer part of the subpixel offset, such as offset, depends on the subpixel, and the "Subpixel Offset" column is specified in 1 / 32 subpixel resolution. In other words, an interpolation filter (such as a subpixel interpolation filter) is represented by the coefficients in the table above.
[0046] In one possible implementation of such a method according to the first aspect, the coefficients of the interpolation filter are as follows:
[0047] [Table 4]
[0048] The "Subpixel Offset" column is specified in 1 / 32 subpixel resolution, depending on the subpixel offset, such as the non-integer part of the subpixel offset, as in In other words, an interpolation filter (such as a subpixel interpolation filter) is represented by the coefficients in the table above.
[0049] In one possible implementation of such a method according to the first aspect, the subpixel interpolation filter is selected from a set of filters used for the intra-prediction process for a given subpixel offset. In other words, the filter for the intra-prediction process for a given subpixel offset is selected from a set of filters (e.g., only one filter, or one of a set of filters, may be used for the intra-prediction process).
[0050] In one possible implementation of such a method according to the first aspect, the set of filters comprises a Gaussian filter and a cubic filter.
[0051] In one possible implementation of such a method according to the first aspect, the number of sub-pixel interpolation filters is N, and N sub-pixel interpolation filters are used for intra reference sample interpolation, where N>=1 and is a positive integer.
[0052] In one possible implementation of such a method according to the first aspect, the reference samples being used to obtain the values of the predicted samples of a block are not adjacent to the block of predicted samples. The encoder may signal an offset value in the bitstream, so that this offset value indicates the distance between the adjacent line of reference samples and the line of reference samples from which the values of the predicted samples are derived. Figure 24 shows possible positions of the line of reference samples and the corresponding value of the ref_offset variable. The variable "ref_offset" indicates which reference line is used, for example, when ref_offset=0, it indicates that "reference line 0" (as shown in Figure 24) is used.
[0053] Example values of the offset in use in a particular implementation of a video codec (eg, a video encoder or a video decoder) are as follows: Using the adjacent line of the reference sample (ref_offset=0, shown by "Reference Line 0" in Figure 24), Use the first line (closest to the adjacent line) (ref_offset=1, shown by "Reference Line 1" in Figure 24), Use the third line (ref_offset=3, indicated by "Reference Line 3" in FIG. 24).
[0054] The directional intra prediction mode specifies the value of the sub-pixel offset (deltaPos) between two adjacent lines of predicted samples. This value is represented by a fixed-point integer value with 5-bit precision. For example, deltaPos=32 means that the offset between two adjacent lines of predicted samples is exactly one sample.
[0055] If the intra-prediction mode is larger than DIA_IDX (mode #34), for the example described above, the value of the primary reference side size is calculated as follows: From the set of intra-prediction modes available (i.e., that the encoder may indicate for the block of prediction samples), the mode that is larger than DIA_IDX and provides the largest deltaPos value is considered. The value of the desired sub-pixel offset is derived as follows: The block height is summed with ref_offset and multiplied by the deltaPos value. If the result is divided by 32 with a remainder of 0, another maximum value for deltaPos as described above, except that previously considered prediction modes are skipped when deriving a mode from the set of available intra-prediction modes. Otherwise, the result of this multiplication is considered to be the maximum fractional sub-pixel offset. The integer part of this offset is taken by shifting it right by 5 bits. The integer part of the maximum fractional sub-pixel offset is summed with the width of the block of prediction samples and half the length of the interpolation filter.
[0056] Otherwise, if the intra-prediction mode is smaller than DIA_IDX (mode #34), for the example described above, the value of the primary reference side size is calculated as follows: From the set of intra-prediction modes available (i.e., that the encoder may indicate for the block of prediction samples), the mode that is smaller than DIA_IDX and provides the largest deltaPos value is considered. The value of the desired sub-pixel offset is derived as follows: The block width is summed with ref_offset and multiplied by the deltaPos value. If the result is divided by 32 with a remainder of 0, another maximum value for deltaPos as described above, except that previously considered prediction modes are skipped when deriving a mode from the set of available intra-prediction modes. Otherwise, the result of this multiplication is considered to be the maximum fractional sub-pixel offset. The integer part of this offset is taken by shifting it right by 5 bits. The integer part of the maximum fractional sub-pixel offset is summed with the height of the block of prediction samples and half the length of the interpolation filter.
[0057] According to a second aspect, the present invention relates to an intra prediction method for predicting a current block included in a picture. The method comprises: determining a size of a dominant reference side used in the intra prediction based on an intra prediction mode among a plurality of available intra prediction modes (wherein the reference sample is a reference sample among a plurality of reference samples included in the dominant reference side) that results in a maximum non-integer value of a sub-pixel offset between a target sample among a plurality of target samples (such as a current sample among a plurality of current samples) in the current block and a reference sample used to predict the target sample in the current block, and a size of an interpolation filter to be applied to the reference sample included in the dominant reference side. The method further comprises applying an interpolation filter to the reference sample included in the dominant reference side to obtain a filtered reference sample, and predicting a plurality of samples (such as a plurality of current samples or a plurality of target samples) included in the current block based on the filtered reference sample.
[0058] Thus, this disclosure facilitates achieving memory efficiency in video coding using intra prediction.
[0059] For example, the size of the primary reference side is determined as the sum of the integer part of the maximum non-integer value of the sub-pixel offset, the size of the side of the current block, and half the size of the interpolation filter. In other words, the advantages of the second aspect may correspond to the above-mentioned advantages of the first aspect.
[0060] In some embodiments, if the intra prediction mode is greater than the vertical intra prediction mode VER_IDX, the side of the current block is the width of the current block, or if the intra prediction mode is less than the horizontal intra prediction mode HOR_IDX, the side of the current block is the height of the current block.
[0061] For example, the value of a reference sample having a position in the primary reference side that is greater than twice the size of the side of the current block is set equal to the value of a sample having a sample position that is twice the size of the current block.
[0062] For example, the size of the main reference side is The size of the current block side and Half the length of the interpolation filter -1, The following, i.e., - the size of the block sides, and - the integer part of the largest non-integer value of the sub-pixel offset plus half the length of the interpolation filter (thus the additional sample ref[refW+refIdx+x] for x=1..(Max(1,nTbW / nTbH)*refIdx+1) is derived as follows: ref[refW+refIdx+x]=p[-1+refW][-1-refIdx]), or the integer part of the largest non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1 (thus the additional sample ref[refW+refIdx+x] for x=1..(Max(1,nTbW / nTbH)*refIdx+2) is derived as follows: ref[refW+refIdx+x]=p[-1+refW][-1-refIdx]) The largest of is determined as the sum of
[0063] According to a third aspect, the present invention relates to an encoder comprising processing circuitry for carrying out a method according to the first or second aspect of the invention or any possible embodiment of the first or second aspect.
[0064] According to a fourth aspect, the present invention relates to a decoder comprising processing circuitry for carrying out a method according to the first or second aspect of the invention or any possible embodiment of the first or second aspect.
[0065] According to a fifth aspect, the present invention relates to an apparatus for intra prediction of a current block included in a picture, the apparatus comprising: an intra prediction unit configured to predict a target sample included in the current block based on a filtered reference sample, the intra prediction unit comprising: a determination unit configured to determine a size of a dominant reference side used in the intra prediction based on an intra prediction mode among a plurality of available intra prediction modes (where the reference sample is a reference sample among a plurality of reference samples included in a dominant reference side) that results in a maximum non-integer value of a sub-pixel offset between a target sample among a plurality of target samples in the current block and a reference sample used to predict the target sample in the current block, and a size of an interpolation filter to be applied to the reference sample included in the dominant reference side; and a filtering unit configured to apply the interpolation filter to the reference sample included in the dominant reference side to obtain a filtered reference sample.
[0066] Thus, this disclosure facilitates achieving memory efficiency in video coding using intra prediction.
[0067] In some embodiments, the determination unit determines the size of the primary reference side as the sum of the integer part of the maximum non-integer value of the sub-pixel offset, the size of the side of the current block, and half the size of the interpolation filter.
[0068] For example, if the intra prediction mode is greater than the vertical intra prediction mode VER_IDX, the side of the current block is the width of the current block, or if the intra prediction mode is less than the horizontal intra prediction mode HOR_IDX, the side of the current block is the height of the current block.
[0069] For example, the value of a reference sample having a position in the primary reference side that is greater than twice the size of the side of the current block is set equal to the value of a sample having a sample position that is twice the size of the current block.
[0070] In some embodiments, the decision unit comprises: The size of the block sides, Determine the size of the primary reference side as the sum of the integer part of the largest non-integer value of the sub-pixel offset + half the length of the interpolation filter, or the integer part of the largest non-integer value of the sub-pixel offset + half the length of the interpolation filter + 1.
[0071] The determination unit may be configured to not perform right padding when the maximum of the two values M is equal to the size of a side of the block, or to perform right padding when the maximum of the two values M is equal to the integer part of the maximum sub-pixel offset plus half the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1.
[0072] Additionally or alternatively, in some embodiments, the determining unit is configured to perform the padding by repeating the first sample and / or the last sample of the primary reference side on the left side and / or the right side, respectively, specifically as follows: denoting the primary reference side as ref and the size of the primary reference side as refS, the padding is expressed as ref[-1]=p[0] and / or ref[refS+1]=p[refS], where ref[-1] represents the left value of the primary reference side and p[0] represents the value of the first reference sample of the primary reference side; ref[refS+1] represents the right value of the primary reference side, and p[refS] represents the value of the last reference sample of the primary reference side.
[0073] The method according to the second aspect of the invention may be performed by an apparatus according to the fifth aspect of the invention. Further features and implementations of the apparatus according to the fifth aspect of the invention correspond to the features and implementations of the method according to the second aspect of the invention or to any possible embodiment of the second aspect.
[0074] According to a sixth aspect, there is provided an apparatus comprising a module / unit / component / circuit for performing at least some of the steps of any preceding implementation of any preceding aspect or of the above method according to any such preceding aspect.
[0075] The apparatus according to this aspect may be extended to an implementation form corresponding to an implementation form of the method according to any preceding aspect, and therefore, an implementation form of the apparatus comprises the features of the corresponding implementation form of the method according to any preceding aspect.
[0076] The advantages of the apparatus according to any preceding aspect are the same as the advantages for the corresponding implementation of the method according to any preceding aspect.
[0077] According to a seventh aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory storing instructions for causing the processor to perform the method according to the first aspect or any possible embodiment of the first aspect.
[0078] According to an eighth aspect, the present invention relates to a video encoder for encoding a plurality of pictures into a bitstream, comprising an apparatus for intra prediction of a current block according to any of the embodiments described above.
[0079] According to a ninth aspect, the present invention relates to a video decoder for decoding a plurality of pictures from a bitstream, comprising an apparatus for intra prediction of a current block according to any of the embodiments described above.
[0080] According to a tenth aspect, there is proposed a computer-readable storage medium having stored thereon instructions that, when executed, cause one or more configured processors to code video data, the instructions causing the one or more processors to perform a method according to the first aspect or any possible embodiment of the first aspect.
[0081] According to an eleventh aspect, the present invention relates to a computer program comprising a program code for performing the method according to the first aspect or any possible embodiment of the first aspect, when the computer program is executed on a computer.
[0082] In another aspect of the present application, a decoder is disclosed that includes processing circuitry configured to perform the above method.
[0083] In another aspect of the present application, a computer program product is disclosed comprising program code for performing the above method.
[0084] In another aspect of the present application, a decoder for decoding video data is disclosed, the decoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the decoder to perform the method set forth above.
[0085] The processing circuitry may be implemented in hardware or in a combination of hardware and software, such as by a software programmable processor.
[0086] The aspects, embodiments, and implementations described herein may provide the advantageous effects described above with reference to the first and second aspects.
[0087] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
[0088] The following embodiments of the invention will be described in more detail with reference to the accompanying figures and drawings. [Brief explanation of the drawings]
[0089] [Figure 1A] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] FIG. 2 is a block diagram illustrating another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention; [Figure 3] 1 is a block diagram illustrating an exemplary structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating an example of an encoding device or a decoding device. [Figure 5] FIG. 10 is a block diagram showing another example of an encoding device or a decoding device. [Figure 6] FIG. 10 illustrates the directions and modes of angular intra prediction and the associated values of pang for the vertical prediction direction. [Figure 7] 10 shows the conversion of pref to p1,ref for a 4x4 block. [Figure 8] FIG. 10 is a diagram illustrating the configuration of p1,ref for horizontal angle prediction. [Figure 9] FIG. 10 is a diagram illustrating the configuration of p1,ref for vertical angle prediction. [Figure 10A] FIG. 10 illustrates the direction and mode of angular intra prediction and the associated values of pang f A set of intra prediction modes in JEM and BMS-1. [Figure 10B] FIG. 10 illustrates the directions and modes of angular intra prediction and the associated values of the pang f A set of intra prediction modes in VVC Draft 2. [Figure 11] FIG. 1 is a diagram illustrating intra prediction modes in HEVC [1]. [Figure 12] FIG. 10 is a diagram illustrating an example of interpolation filter selection. [Figure 13] FIG. 1 is a diagram illustrating the QTBT to be described. [Figure 14] FIG. 10 is a diagram showing the orientation of rectangular blocks. [Figure 15A] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 15B] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 15C] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 16] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 17] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 18] FIG. 10 is a diagram illustrating an example of intra prediction of a block from reference samples on the dominant reference side. [Figure 19] FIG. 1 illustrates an interpolation filter used in intra prediction. [Figure 20] FIG. 1 illustrates an interpolation filter used in intra prediction. [Figure 21] FIG. 1 illustrates an interpolation filter used in intra prediction. [Figure 22] FIG. 2 illustrates an interpolation filter used in intra prediction, configured to implement an embodiment of the present invention. [Figure 23] FIG. 2 illustrates an interpolation filter used in intra prediction, configured to implement an embodiment of the present invention. [Figure 24] FIG. 10 shows another example of possible positions of a line of reference samples and corresponding values of the ref_offset variable. [Figure 25]10 is a flowchart illustrating an intra prediction method. [Figure 26] FIG. 1 is a block diagram illustrating an intra-prediction device. [Figure 27] 1 is a block diagram illustrating an exemplary structure of a content supply system that provides content distribution services. [Figure 28] FIG. 2 is a block diagram illustrating an example of a terminal device. DETAILED DESCRIPTION OF THE INVENTION
[0090] In the following, identical reference signs refer to identical or at least functionally equivalent features, unless expressly specified otherwise.
[0091] In the following description, reference is made to the accompanying figures, which form a part of this disclosure and which show, by way of illustration, certain aspects of embodiments of the invention or in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other ways and may include structural or logical changes not shown in the figures. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0092] For example, it is understood that disclosure regarding a described method may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, when one or more particular method steps are described, the corresponding device may include one or more units (e.g., one unit that performs one or more steps, or multiple units that each perform one or more of the multiple steps), e.g., functional units, for performing the described one or more method steps, even if such one or more units are not explicitly described or shown in a figure. On the other hand, when a particular apparatus is described, for example, based on one or more units, e.g., functional units, the corresponding method may include one step (e.g., one step that performs the functionality of one or more units, or multiple steps that each perform the functionality of one or more of the multiple units) for performing the functionality of the one or more units, even if such one or more steps are not explicitly described or shown in a figure. Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically stated otherwise.
[0093] Video coding typically refers to the processing of a sequence of pictures that form a video or a video sequence. Instead of the term "picture," the terms "frame" or "image" may be used as synonyms in the field of video coding. Video coding (or, in general, coding) comprises two parts: video encoding and video decoding. Video encoding is performed at the source side and typically comprises processing an original video picture (e.g., by compression) to reduce the amount of data needed to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed at the destination side and typically comprises the reverse processing compared to the encoder to reconstruct the video picture. Embodiments referring to "coding" a video picture (or, in general, a picture) shall be understood to relate to "encoding" or "decoding" the video picture or respective video sequence. The combination of the encoding and decoding parts is also called a codec (coding and decoding).
[0094] In the case of lossless video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming there is no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, further compression, for example by quantization, is performed to reduce the amount of data representing the video picture, and the video picture cannot always be perfectly reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0095] Some video coding standards belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is usually partitioned into a set of non-overlapping blocks, and coding is usually performed at the block level. In other words, at the encoder, video is usually processed or coded at the block (video block) level, for example, by using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to generate a predictive block, subtract the predictive block from a current block (the block currently being / to be processed) to obtain a residual block, transform the residual block, and quantize the residual block in the transform domain to reduce the amount of data to be transmitted (compression); at the decoder, an inverse process compared to the encoder is applied to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder processing loop, both of which generate the same prediction (e.g., intra-prediction and inter-prediction) and / or reconstruction for processing or coding subsequent blocks.
[0096] In the following embodiment of a video coding system 10, a video encoder 20 and a video decoder 30 are described based on FIGS.
[0097] 1A is a schematic block diagram illustrating an example coding system 10, e.g., a video coding system 10 (or short coding system 10) that may utilize the techniques of this application. A video encoder 20 (or short encoder 20) and a video decoder 30 (or short decoder 30) of video coding system 10 represent examples of devices that may be configured to perform techniques according to various examples described in this application.
[0098] As shown in FIG. 1A, coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14 for decoding, for example, encoded picture data 13.
[0099] Source device 12 comprises an encoder 20 and may additionally, i.e., optionally, comprise a picture source 16 , a pre-processor (or pre-processing unit) 18 , for example, a picture pre-processor 18 , and a communication interface or unit 22 .
[0100] Picture source 16 may comprise or be any kind of picture capture device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any kind of other device for obtaining and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture source may be any kind of memory or storage that stores any of the above-mentioned pictures.
[0101] To distinguish between the preprocessor 18 and the processing performed by the preprocessing unit 18, the picture or picture data 17 may also be referred to as a raw picture or raw picture data 17.
[0102] The pre-processor 18 is configured to receive (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. The pre-processing performed by the pre-processor 18 may comprise, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the pre-processing unit 18 may be an optional component.
[0103] Video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (further details are described below, eg, with reference to FIG. 2).
[0104] The communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or a further processed version thereof) via the communication channel 13 to another device, for example, the destination device 14 or any other device, for storage or direct reconstruction.
[0105] Destination device 14 comprises a decoder 30 (e.g., video decoder 30), and may additionally, i.e., optionally, comprise a communications interface or communications unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0106] The communications interface 28 of the destination device 14 is configured to receive the coded picture data 21 (or further processed versions thereof), for example, directly from the source device 12 or from any other source, for example, a storage device, for example, a coded picture data storage device, and to provide the coded picture data 21 to the decoder 30.
[0107] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., a direct wired or wireless connection, or via any type of network, e.g., a wired or wireless network or any combination thereof, or any type of private and public network or any combination thereof.
[0108] The communications interface 22 may be configured, for example, to package the encoded picture data 21 into a suitable format, e.g., packets, and / or process the encoded picture data using any type of transmission encoding or transmission process for transmission over a communications link or network.
[0109] The communications interface 28, which forms the counterpart of the communications interface 22, may be configured, for example, to receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or transmission processing and / or depackaging to obtain the encoded picture data 21.
[0110] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrows for communication channel 13 in FIG. 1A pointing from source device 12 to destination device 14, or as bidirectional communication interfaces, e.g., configured to send and receive messages to, for example, set up a connection and acknowledge, respond, and exchange any other information related to the communication link and / or data transmission, e.g., encoded picture data transmission.
[0111] The decoder 30 is configured to receive the coded picture data 21 and to provide decoded picture data 31 or decoded pictures 31 (further details are described below, for example, based on Figure 3 or Figure 5).
[0112] Post-processor 32 of destination device 14 is configured to post-process decoded picture data 31 (also called reconstructed picture data), e.g., decoded picture 31, to obtain post-processed picture data 33, e.g., post-processed picture 33. The post-processing performed by post-processing unit 32 may comprise, e.g., color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare decoded picture data 31, e.g., for display, by, e.g., display device 34.
[0113] Display device 34 of destination device 14 is configured to receive the post-processed picture data 33, e.g., for displaying the picture to a user or viewer. Display device 34 may be or comprise any type of display for presenting the reconstructed picture, e.g., an integrated or external display or monitor. The display may comprise, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0114] 1A depicts source device 12 and destination device 14 as separate devices, device embodiments may also include both source device 12 or corresponding functionality and destination device 14 or corresponding functionality, or both. In such embodiments, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.
[0115] As will be clear to those skilled in the art based on the description, the functionality of different units, i.e., the presence and (exact) division of functionality within source device 12 and / or destination device 14 as shown in FIG. 1A, may vary depending on the actual device and application.
[0116] Encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) may each be implemented as any of a variety of suitable circuit configurations, such as shown in FIG. 1B , such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. Where techniques are implemented partially in software, a device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the above (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated within the respective device as part of a combined encoder / decoder (codec).
[0117] Source device 12 and destination device 14 may comprise any of a wide range of devices, including any type of handheld or fixed device, e.g., a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system or any type of operating system. In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.
[0118] 1A is merely an example, and the techniques of the present application may be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between encoding and decoding devices. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode and store data in memory, and / or a video decoding device may retrieve and decode data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data to memory and / or retrieve and decode data from memory.
[0119] 1B is an illustrative diagram of another example video coding system 40 including the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3, according to an example embodiment. System 40 may implement techniques according to various examples described herein. In the illustrated implementation, video coding system 40 may include an imaging device 41, a video encoder 100, a video decoder 30 (and / or a video coder implemented via logic circuitry 47 of a processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.
[0120] As shown, imaging device 41, antenna 42, processing unit 46, logic circuitry 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 may be capable of communicating with one another. As described, although illustrated with both video encoder 20 and video decoder 30, video coding system 40 may include only video encoder 20 or only video decoder 30 in various examples.
[0121] As shown, in some examples, video coding system 40 may include antenna 42. Antenna 42 may be configured to transmit or receive, for example, an encoded bitstream of video data. Further, in some examples, video coding system 40 may include display device 45. Display device 45 may be configured to present the video data. As shown, in some examples, logic circuitry 47 may be implemented via processing unit 46. Processing unit 46 may include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Video coding system 40 may also include optional processor 43, which may also include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. In some examples, logic circuitry 47 may be implemented via hardware, video coding-specific hardware, etc., and processor 43 may implement general-purpose software, an operating system, etc. Additionally, memory store 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory store 44 may be implemented by a cache memory. In some examples, logic circuitry 47 may access memory store 44 (e.g., for implementing an image buffer). In other examples, logic circuitry 47 and / or processing unit 46 may include a memory store (e.g., a cache, etc.) for implementing an image buffer, etc.
[0122] In some examples, video encoder 20 implemented via logic circuitry may include an image buffer (e.g., via either processing unit 46 or memory store 44) and a graphics processing unit (e.g., via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 20 implemented via logic circuitry 47 to embody various modules such as those described with respect to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuitry may be configured to perform various operations as described herein.
[0123] Video decoder 30 may be implemented in a manner similar to that implemented via logic circuitry 47 to embody various modules as described with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, video decoder 30 may be implemented via logic circuitry and may include an image buffer (e.g., by either processing unit 420 or memory store 44) and a graphics processing unit (e.g., by processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video decoder 30 as implemented via logic circuitry 47 to embody various modules as described with respect to FIG. 3 and / or any other decoder system or subsystem described herein.
[0124] In some examples, antenna 42 of video coding system 40 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data related to encoding video frames as described herein, such as data related to coding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as described), and / or data defining the coding partitions), indicators, index values, mode selection data, etc. Video coding system 40 may also include video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. A display device 45 configured to present the video frames.
[0125] For ease of explanation, embodiments of the present invention are described herein with reference to, for example, High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) reference software, i.e., the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.
[0126] Encoder and encoding method 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input unit 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210 and an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output unit 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder using a hybrid video codec.
[0127] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming a forward signal path of the encoder 20, and the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming a backward signal path of the video encoder 20, which corresponds to the signal path of a decoder (see video decoder 30 in FIG. 3 ). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may also be referred to as forming a “built-in decoder” of the video encoder 20.
[0128] Pictures and picture divisions (pictures and blocks) The encoder 20 may be configured to receive, e.g., via an input 201, a picture 17 (or picture data 17), e.g., a picture of a sequence of pictures forming a video or a video sequence. The received picture or picture data may also be a preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be called a current picture (particularly in video coding, to distinguish the current picture from other pictures, e.g., previously coded and / or decoded pictures, of the same video sequence, i.e., the video sequence that also comprises the current picture) or a picture to be coded.
[0129] A (digital) picture is or can be considered as a two-dimensional array or matrix of samples with intensity values. The samples in the array are sometimes called pixels (a short form of picture element) or pels. The number of samples in the horizontal and vertical directions (i.e., axes) of the array or picture defines the size and / or resolution of the picture. For color depiction, three color components are usually employed, i.e., a picture may represent or contain three sample arrays. In an RBG format or color space, a picture comprises corresponding red, green, and blue sample arrays. However, in video coding, each pixel is usually represented in a luminance and chrominance format or color space, e.g., YCbCr, comprising a luminance component denoted by Y (sometimes L is used instead) and two chrominance components denoted by Cb and Cr. The luminance (or short luma) component Y represents brightness or gray-level intensity (e.g., as in a grayscale picture), and the two chrominance (or short chroma) components Cb and Cr represent chromaticity or color information components. Thus, a picture in YCbCr format comprises a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in RGB format may be converted or transformed to YCbCr format, or vice versa, a process also called color transformation or color conversion. If a picture is monochrome, the picture may comprise only a luminance sample array. Thus, a picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0130] Embodiments of video encoder 20 may include a picture partition unit (not shown in FIG. 2) configured to partition picture 17 into multiple (usually non-overlapping) picture blocks 203. These blocks are sometimes called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partition unit may be configured to use the same block size for all pictures of a video sequence and a corresponding grid that defines the block size, or to vary the block size among pictures or subsets or groups of pictures and partition each picture into corresponding blocks.
[0131] In further embodiments, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, some, or all of the blocks that form picture 17. Picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.
[0132] Like picture 17, picture block 203 again may be or be considered to be a two-dimensional array or matrix of samples having intensity values (sample values), but with smaller dimensions than picture 17. In other words, block 203 may comprise, for example, one sample array (e.g., a luma array in the case of a monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color picture 17), or any other number and / or type of array depending on the applied color format. The number of samples in the horizontal and vertical directions (i.e., axes) of block 203 define the size of block 203. Thus, a block may be, for example, an M×N (M columns by N rows) array of samples or an M×N array of transform coefficients.
[0133] An embodiment of video encoder 20, such as that shown in FIG. 2, may be configured to code picture 17 block-wise, eg, encoding and prediction is performed for each block 203.
[0134] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as a residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be provided later), for example, by subtracting sample values of the prediction block 265 on a sample-by-sample basis (pixel-by-pixel basis) from sample values of the picture block 203, to obtain the residual block 205 in the sample domain.
[0135] conversion The transform processing unit 206 may be configured to apply a transform, for example, a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207, sometimes referred to as transform residual coefficients, represent the residual block 205 in the transform domain.
[0136] The transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as the transform specified for H.265 / HEVC. Compared to an orthogonal DCT transform, such an integer approximation is typically scaled by several factors. To maintain the norm of the residual block processed by the forward and inverse transforms, additional scaling factors are applied as part of the transform process. The scaling factors are typically chosen based on several constraints, such as the scaling factors being powers of two for shift operations, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. A particular scaling factor may be specified for, e.g., an inverse transform, e.g., by the inverse transform processing unit 212 (and a corresponding inverse transform, e.g., by the inverse transform processing unit 312 in the video decoder 30), and a corresponding scaling factor for the forward transform, e.g., by the transform processing unit 206 in the encoder 20, may be specified accordingly.
[0137] An embodiment of video encoder 20 (respectively, transform processing unit 206) may be configured to output transform parameters, e.g., one or more types of transform, e.g., directly or encoded or compressed via entropy coding unit 270, such that video decoder 30 may receive and use the transform parameters for decoding.
[0138] quantization The quantization unit 208 may be configured to quantize the transform coefficients 207, for example, by applying scalar quantization or vector quantization, to obtain quantized coefficients 209. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.
[0139] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be truncated to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, in the case of scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, and a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index into a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size) and a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may involve division by a quantization step size, and corresponding and / or inverse dequantization, e.g., by the inverse quantization unit 210, may involve multiplication by the quantization step size. Some standards, e.g., HEVC, embodiments may be configured to use a quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified by the scaling used in the fixed-point approximation of the equation for the quantization step size and the quantization parameter. In an exemplary implementation, the scaling of the inverse transform and the inverse quantization may be combined. Alternatively, customized quantization tables may be used and may be signaled, e.g., in the bitstream, from the encoder to the decoder. Quantization is a lossy operation, and loss increases with increasing quantization step size.
[0140] Embodiments of video encoder 20 (respectively, quantization unit 208) may be configured to output a quantization parameter (QP), e.g., encoded directly or via entropy coding unit 270, so that video decoder 30 may receive and apply the quantization parameter for decoding.
[0141] inverse quantization The inverse quantization unit 210 is configured to apply the inverse quantization of the quantization unit 208 to the quantized coefficients, e.g., by applying the inverse of the quantization scheme applied by the quantization unit 208, based on or using the same quantization step size as the quantization unit 208, to obtain the inverse quantized coefficients 211. The inverse quantized coefficients 211 are sometimes referred to as the inverse quantized residual coefficients 211, and correspond to the transform coefficients 207—although they are typically not identical to the transform coefficients due to loss due to quantization.
[0142] Inverse transformation The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0143] Reconstruction The reconstruction unit 214 (e.g., an adder or summer 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, e.g., by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 - sample by sample - to obtain the reconstructed block 215 in the sample domain.
[0144] Filtering The loop filter unit 220 (or short “loop filter” 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or in general, to filter reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may comprise one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, for example, a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is illustrated in FIG. 2 as being an in-loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstruction block 221. Decoded picture buffer 230 may store reconstructed coding blocks after loop filter unit 220 performs filtering operations on the reconstructed coding blocks.
[0145] Embodiments of video encoder 20 (respectively, loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information), e.g., directly or encoded via entropy coding unit 270, such that decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0146] Decoded Picture Buffer The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures, or generally, reference picture data, for encoding video data by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks, e.g., previously reconstructed and filtered blocks 221, of the same current picture or of a different picture, e.g., a previously reconstructed picture, to provide, e.g., for inter-prediction, a complete, previously reconstructed, i.e., decoded picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstruction blocks 215, or generally unfiltered reconstructed samples, for example, if the reconstruction blocks 215 have not been filtered by the loop filter unit 220, or any other further processed version of the reconstruction blocks or samples.
[0147] Mode Selection (Segmentation and Prediction) The mode select unit 260 comprises a partition unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, e.g., original block 203 (current block 203 of current picture 17), and reconstructed picture data, e.g., filtered and / or unfiltered reconstructed samples or blocks of the same (current) picture and / or from one or more previously decoded pictures, e.g., from the decoded picture buffer 230 or other buffer (e.g., line buffer, not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter prediction or intra prediction, to obtain a prediction block 265 or predictor 265.
[0148] The mode selection unit 260 may be configured to determine or select a classification (including no classification) for the current block prediction mode and a prediction mode (e.g., an intra-prediction mode or an inter-prediction mode) and generate a corresponding prediction block 265, which is used for calculating the residual block 205 and for reconstructing the reconstruction block 215.
[0149] Embodiments of the mode selection unit 260 may be configured to select a partition and prediction mode (e.g., from those supported by or available to the mode selection unit 260) that provides the best match, or in other words, the smallest residual (smallest residual means better compression for transmission or storage), or the smallest signaling overhead (smallest signaling overhead means better compression for transmission or storage), or that considers or balances both. The mode selection unit 260 may be configured to determine the partition and prediction mode based on rate distortion optimization (RDO), i.e., select a prediction mode that results in the smallest rate distortion. Terms such as “best,” “minimum,” “optimum,” etc. in this context do not necessarily refer to an overall “best,” “minimum,” “optimum,” etc., but may also refer to the achievement of a termination or selection criterion, such as a value above or below a threshold or other constraint that potentially leads to a “suboptimal selection” but reduces complexity and processing time.
[0150] In other words, the partitioning unit 262 may be configured to partition the block 203 into smaller block partitions or sub-blocks (which again form blocks), for example using quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and to perform prediction on each of the block partitions or sub-blocks, for example, wherein the mode selection comprises selecting a tree structure of the partitioned block 203, and a prediction mode is applied to each of the block partitions or sub-blocks.
[0151] Below, the partitioning (eg, by partitioning unit 260) and prediction processes (by inter-prediction unit 244 and intra-prediction unit 254) performed by exemplary video encoder 20 are described in more detail.
[0152] classification The partitioning unit 262 may partition (i.e., divide) the current block 203 into smaller partitions, e.g., square or rectangular sized smaller blocks. These smaller blocks (sometimes called sub-blocks) may be further partitioned into even smaller partitions. This is also called tree partitioning or hierarchical tree partitioning; for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned, e.g., into two or more blocks at the next lower tree level, e.g., into nodes at tree level 1 (hierarchical level 1, depth 1), which may again be partitioned into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), and so on, until, e.g., a termination criterion is achieved and partitioning is terminated, e.g., to reach a maximum tree depth or a minimum block size. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree that uses a partition into two partitions is called a binary-tree (BT), a tree that uses a partition into three partitions is called a ternary-tree (TT), and a tree that uses a partition into four partitions is called a quad-tree (QT).
[0153] As previously mentioned, the term "block" as used herein may refer to a portion of a picture, particularly a square or rectangular portion. For example, with reference to HEVC and VVC, a block may be or correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0154] For example, a coding tree unit (CTU) may be or comprise a CTB of luma samples, two corresponding CTBs of chroma samples, or a CTB of samples for a monochrome picture or a picture coded using three separate color planes, and a syntax structure used to code the samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some values of N, such that the division of the components into CTBs is partitioned. A coding unit (CU) may be or comprise a coding block of luma samples, two corresponding coding blocks of chroma samples, or a coding block of samples for a monochrome picture or a picture coded using three separate color planes, and a syntax structure used to code the samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N, such that the division of the CTB into coding blocks is partitioned.
[0155] For example, in an HEVC embodiment, coding tree units (CTUs) may be divided into CUs by using a quadtree structure, denoted as a coding tree. The decision of whether a picture area should be coded using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four PUs according to a PU partition type. Inside one PU, the same prediction process is applied, and related information is transmitted to the decoder for each PU. After obtaining residual blocks by applying a prediction process based on the PU partition type, the CUs may be partitioned into transform units (TUs) according to another quadtree structure, similar to the coding tree for CUs.
[0156] For example, in an embodiment according to the latest video coding standard currently under development, called Versatile Video Coding (VVC), quad-tree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, CUs can have either square or rectangular shapes. For example, coding tree units (CTUs) are first partitioned by a quad-tree structure. The quad-tree leaf nodes are further partitioned by a binary tree or a ternary (or ternary) tree structure. The partition tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transform processes without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, multiple partitions, for example, ternary tree partitioning, have also been proposed to be used with the QTBT block structure.
[0157] In one example, mode select unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0158] As described above, video encoder 20 is configured to determine or select a best or optimal prediction mode from a (predetermined) set of prediction modes, which may comprise, for example, intra-prediction modes and / or inter-prediction modes.
[0159] Intra prediction The set of intra prediction modes may comprise 35 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, e.g., as defined in HEVC, or may comprise 67 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, e.g., as defined for VVC.
[0160] The intra prediction unit 254 is configured to generate an intra prediction block 265 according to an intra prediction mode of the set of intra prediction modes using reconstructed samples of neighboring blocks of the same current picture.
[0161] The intra prediction unit 254 (or generally, the mode selection unit 260) is further configured to output the intra prediction parameters (or generally, information indicating the selected intra prediction mode for the block) to the entropy coding unit 270, for example, in the form of a syntax element 266 for inclusion in the coded picture data 21 so that the video decoder 30 may receive and use the prediction parameters for decoding.
[0162] Inter Prediction The set of inter prediction modes (or possible inter prediction modes) depends on the available reference pictures (i.e., previous, at least partially decoded pictures, e.g., stored in DBP230) and other inter prediction parameters, e.g., whether the entire reference picture is used to search for the best matching reference block or whether only a portion of the reference picture, e.g., a search window area around the area of the current block, is used, and / or, e.g., whether pixel interpolation, e.g., half-pel / semi-pel interpolation and / or quarter-pel interpolation, is applied.
[0163] In addition to the above prediction modes, skip mode and / or direct mode may be applied.
[0164] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain a picture block 203 (current picture block 203 of current picture 17) and a decoded picture 231, or at least one or more previously reconstructed blocks, e.g., reconstructed blocks of one or more other / different previously decoded pictures 231, for motion estimation. For example, a video sequence may comprise the current picture and the previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of or form a sequence of pictures that form a video sequence.
[0165] The encoder 20 may be configured to, for example, select a reference block from multiple reference blocks of the same or different pictures among multiple other pictures, and provide the reference picture (or reference picture index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as inter-prediction parameters. This offset is also called a motion vector (MV).
[0166] The motion compensation unit is configured to obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain inter prediction block 265. The motion compensation performed by the motion compensation unit may involve fetching or generating a predictive block based on a motion / block vector determined by motion estimation, possibly performing interpolation to sub-pixel accuracy. Interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate predictive blocks that can be used to code the picture block. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may locate the predictive block to which the motion vector points in one of the reference picture lists.
[0167] The motion compensation unit may also generate syntax elements associated with the blocks and video slices for use by video decoder 30 in decoding picture blocks of the video slices.
[0168] Entropy Coding The entropy encoding unit 270 is configured to apply, for example, an entropy encoding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context-adaptive VLC scheme (CAVLC), an arithmetic coding scheme, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or technique) to the quantized coefficients 209, the inter-prediction parameters, the intra-prediction parameters, the loop filter parameters, and / or other syntax elements, or to bypass (without compression) to obtain coded picture data 21, which may be output via an output 272, for example, in the form of a coded bitstream 21, so that, for example, the video decoder 30 may receive and use the parameters for decoding. The coded bitstream 21 may be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.
[0169] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may, for some blocks or frames, directly quantize the residual signal without using the transform processing unit 206. In another implementation, the encoder 20 may combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.
[0170] Decoder and decoding method 3 shows an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive coded picture data 21 (e.g., coded bitstream 21), e.g., coded by encoder 20, to obtain a decoded picture 331. The coded picture data or bitstream comprises information for decoding the coded picture data, e.g., data representing picture blocks of a coded video slice, and associated syntax elements.
[0171] 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, an inter prediction unit 344, and an intra prediction unit 354. Inter prediction unit 344 may be or may include a motion compensation unit. Video decoder 30 may, in some examples, perform a decoding path that is reciprocal to the encoding path generally described with respect to video encoder 100 from FIG. 2.
[0172] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter prediction unit 344, and intra prediction unit 354 are also referred to as forming a “built-in decoder” of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Accordingly, the descriptions provided for the respective units and functions of video encoder 20 apply correspondingly to the respective units and functions of video decoder 30.
[0173] Entropy Decoding The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally, the coded picture data 21) and, e.g., perform entropy decoding on the coded picture data 21 to obtain, e.g., quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), e.g., any or all of inter-prediction parameters (e.g., reference picture indices and motion vectors), intra-prediction parameters (e.g., intra-prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to an encoding scheme such as described with respect to the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode selection unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0174] inverse quantization Inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally, information related to inverse quantization) and quantized coefficients from coded picture data 21 (e.g., by parsing and / or decoding by entropy decoding unit 304), and apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain inverse quantized coefficients 311, which are sometimes referred to as transform coefficients 311. The inverse quantization process may involve using the quantization parameter determined by video encoder 20 for each video block in a video slice to determine the degree of quantization, and similarly, the degree of inverse quantization to be applied.
[0175] Inverse transformation The inverse transform processing unit 312 may be configured to receive the inverse quantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the inverse quantized coefficients 311 to obtain the reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 313. The transform may be an inverse transform, e.g., an inverse DCT transform, an inverse DST transform, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may further be configured to receive transform parameters or corresponding information from the coded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304) to determine the transform to be applied to the inverse quantized coefficients 311.
[0176] Reconstruction The reconstruction unit 314 (e.g., an adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365, to obtain the reconstructed block 315 in the sample domain.
[0177] Filtering Loop filter unit 320 (either in the coding loop or after the coding loop) is configured to filter reconstruction block 315 to obtain filtered block 321, e.g., to smooth pixel transitions or otherwise improve video quality. Loop filter unit 320 may comprise one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 320 is shown in FIG. 3 as being an in-loop filter, in other configurations, loop filter unit 320 may be implemented as a post-loop filter.
[0178] Decoded Picture Buffer The decoded video blocks 321 of the picture are then stored in a decoded picture buffer 330, which stores the decoded picture 331 as a reference picture for subsequent motion compensation relative to other pictures and / or for outputting a display, respectively.
[0179] The decoder 30 is arranged to output the decoded pictures 311, for example via an output 312, for presentation or viewing to a user.
[0180] prediction The inter prediction unit 344 may be identical to the inter prediction unit 244 (specifically, the motion compensation unit), and the intra prediction unit 354 may be functionally identical to the inter prediction unit 254, performing partition or partition decision and prediction based on partition and / or prediction parameters or respective information received from the coded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304). The mode selection unit 360 may be configured to perform prediction (intra prediction or inter prediction) for each block based on the reconstructed picture, block, or respective sample (filtered or unfiltered) to obtain a prediction block 365.
[0181] When a video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of mode select unit 360 is configured to generate a predictive block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of mode select unit 360 is configured to produce a predictive block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from entropy decoding unit 304. For inter prediction, the predictive block may be produced from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, i.e., List 0 and List 1, using a default construction technique based on the reference pictures stored in DPB 330.
[0182] Mode select unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and use the prediction information to produce predictive blocks for the current video block being decoded. For example, mode select unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra-prediction or inter-prediction) to use for coding the video blocks of the video slice, an inter-prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists for the slice, a motion vector for each inter-coded video block of the slice, an inter-prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0183] Other variations of video decoder 30 may be used to decode encoded picture data 21. For example, decoder 30 may produce an output video stream without using loop filtering unit 320. For example, non-transform-based decoder 30 may, for some blocks or frames, directly inverse quantize the residual signal without using inverse transform processing unit 312. In another implementation, video decoder 30 may combine inverse quantization unit 310 and inverse transform processing unit 312 into a single unit.
[0184] 4 is a schematic diagram of a video coding device 400 according to one embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 of FIG. 1A, or an encoder, such as the video encoder 20 of FIG. 1A.
[0185] Video coding device 400 comprises an ingress port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing the data, a transmitter unit (Tx) 440 and an egress port 450 (or output port 450) for transmitting the data, and a memory 460 for storing the data. Video coding device 400 may also comprise optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to ingress port 410, receiver unit 420, transmitter unit 440, and egress port 450 for the egress or ingress of optical or electrical signals.
[0186] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. The processor 430 is in communication with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 performs, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 470 significantly improves the functionality of the video coding device 400 and affects the transformation of the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0187] Memory 460 may comprise one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device for storing programs when such programs are selected for execution, and for storing instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0188] 5 is a simplified block diagram of an apparatus 500 that may be used as one or both of source device 12 and destination device 14 from FIG. 1 according to an exemplary embodiment. Apparatus 500 is capable of implementing the techniques of the present application described above. Apparatus 500 may take the form of a computing system including multiple computing devices, or a single computing device, such as a mobile phone, tablet computer, laptop computer, notebook computer, desktop computer, etc.
[0189] The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices, now existing or later developed, capable of manipulating or processing information. While the disclosed implementations may be practiced using a single processor as shown, e.g., processor 502, advantages in speed and efficiency may be achieved using two or more processors.
[0190] The memory 504 in the device 500, in one implementation, may be a read-only memory (ROM) device or a random-access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and application programs 510, which include at least one program that enables the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1-N, which further include a video coding application that performs the methods described herein. The device 500 may also include additional memory in the form of secondary storage 514, which may be, for example, a memory card used with a mobile computing device. Because a video communication session may contain a significant amount of information, it may be stored in whole or in part in the secondary storage 514 and loaded into the memory 504 as needed for processing.
[0191] The device 500 may also include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with touch-sensitive elements operable to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512. In addition to, or as an alternative to, the display 518, other output devices may be provided that allow a user to program or otherwise use the device 500. When the output device is or includes a display, the display may be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.
[0192] The device 500 may also include or be in communication with an image sensing device 520, for example, a camera or any other existing or later developed image sensing device 520 that can sense images, such as an image of a user operating the device 500. The image sensing device 520 may be positioned to point towards the user operating the device 500. In one example, the position and optical axis of the image sensing device 520 may be configured such that it is directly adjacent to the display 518 and has a field of view that includes the area from which the display 518 is viewable.
[0193] The device 500 may also include or be in communication with a sound sensing device 522, e.g., a microphone or any other existing or later developed sound sensing device that can sense sound near the device 500. The sound sensing device 522 may be positioned to point towards a user operating the device 500 and may be configured to receive sound, e.g., voice or other utterances made by a user while the user is operating the device 500.
[0194] While FIG. 5 depicts the processor 502 and memory 504 of device 500 as integrated into a single unit, other configurations may be utilized. The operations of processor 502 may be distributed across multiple machines (each machine having one or more of the processors), which may be coupled directly or across a local area network or other network. Memory 504 may be distributed across multiple machines, such as network-based memory or memory in multiple machines that perform the operations of device 500. While shown here as a single bus, bus 512 of device 500 may consist of multiple buses. Furthermore, secondary storage 514 may be directly coupled to other components of device 500 or may be accessed over a network and may comprise a single integrated unit, such as a memory card, or multiple units, such as multiple memory cards. Device 500 may thus be implemented in a wide variety of configurations.
[0195] Acronym Definitions and Glossary JEM Collaborative Exploration Model (Software Code Base for Future Video Coding Exploration) JVET Joint Video Experts Team LUT Lookup Table QT quadtree QTBT Quadrant tree + Binary tree RDO Rate Distortion Optimization ROM Read Only Memory VTM VVC test model VVC Versatile Video Coding, a standardization project developed by JVET CTU / CTB Coding Tree Unit / Coding Tree Block CU / CB coding unit / coding block PU / PB Prediction Unit / Prediction Block TU / TB Conversion Unit / Conversion Block HEVC High Efficiency Video Coding
[0196] Video coding schemes such as H.264 / AVC and HEVC are designed around the successful principle of block-based hybrid video coding, whereby a picture is first partitioned into blocks, and then each block is predicted by using intra-picture or inter-picture prediction.
[0197] Several video coding standards since H.261 belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is usually partitioned into a set of non-overlapping blocks, and coding is usually performed at the block level. In other words, at the encoder, video is usually processed or coded at the block (picture block) level, for example by using spatial (intra-picture) and temporal (inter-picture) prediction to generate a predictive block, subtract the predictive block from a current block (the block currently being / to be processed) to obtain a residual block, transform the residual block and quantize it in the transform domain to reduce the amount of data to be transmitted (compression), and at the decoder, a reverse process compared to the encoder is partly applied to the coded or compressed block to reconstruct the current block for representation. Additionally, the encoder replicates the decoder processing loop, both of which generate the same prediction (eg, intra-prediction and inter-prediction) and / or reconstruction for processing, ie, coding, subsequent blocks.
[0198] As used herein, the term "block" may refer to a portion of a picture or a frame. For ease of explanation, embodiments of the present invention are described herein with reference to High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) reference software developed by the Joint Video Coding Study Group (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC. It may refer to CUs, PUs, and TUs. In HEVC, CTUs are divided into CUs by using a quadtree structure, denoted as a coding tree. The decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder for each PU. After obtaining the residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for CUs. In the latest developments in video compression technology, quadtree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, CUs can have either square or rectangular shapes. For example, coding tree units (CTUs) are first partitioned using a quadtree structure. The quadtree leaf nodes are further partitioned using a binary tree structure. The binary tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transform processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, hybrid partitioning, such as ternary tree partitioning, has also been proposed to be used with the QTBT block structure.
[0199] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC29 / WG11) are considering the potential need for standardization of future video coding technologies with compression capabilities significantly exceeding those of the current HEVC standard (including its current and upcoming extensions for screen content coding and high dynamic range coding). The groups are collaborating on this exploration in a collaborative research effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by their experts in this area.
[0200] For directional intra prediction, intra prediction modes are available, representing different prediction angles from the diagonal up to the diagonal down. To define the prediction angle, an offset value pang on a 32-sample grid is defined. p to the corresponding intra prediction mode ang The association of p is visualized in Figure 6 for the vertical prediction mode. For the horizontal prediction mode, the scheme is flipped to the vertical direction and p ang As mentioned above, all angle prediction modes are available for all applicable intra prediction block sizes. They all use the same 32-sample grid for the definition of the prediction angle. p across the 32-sample grid in Figure 6 ang The distribution of values reveals increased resolution of the predicted angles around the vertical direction and coarser resolution of the predicted angles towards the diagonal direction. The same applies to the horizontal direction. This design comes from the observation that in many video contents, near-horizontal and vertical structures play a significant role compared to diagonal structures.
[0201] For horizontal and vertical prediction directions, the selection of the sample to be used for prediction is straightforward, but for angular prediction this task is more difficult. For modes 11 to 25, the predicted sample p ref When predicting the current block Bc from the set of (also called the primary reference side), pref Samples from both the vertical and horizontal parts of p can be involved. ref Since determining the location of each sample in any of the branches of p requires some computational effort, a unified 1D prediction reference is designed for HEVC intra prediction. The scheme is visualized in Figure 7. Before performing the actual prediction operation, the reference sample p ref The set of is a one-dimensional vector p 1,ref The projection used for the mapping depends on the direction indicated by the intra prediction angle of the respective intra prediction mode. ref Only the reference samples from the part of p 1,ref For each angle prediction mode, p 1,ref The actual mapping of the reference samples to the reference sample set p is shown in Figures 8 and 9 for the horizontal and vertical angle prediction directions, respectively. 1,ref is constructed once for a block of prediction samples. The prediction is then derived from two adjacent reference samples in the set as described in detail below. As can be seen from Figures 8 and 9, the one-dimensional reference sample set is not completely filled for all intra-prediction modes. Only locations within the projection range for the corresponding intra-prediction direction are included in the set.
[0202] Prediction for both horizontal and vertical prediction modes is performed in the same way, just by swapping the x and y coordinates of the block. 1,ref The prediction from p is performed with 1 / 32 pixel accuracy. Depending on the value of the angle parameter pang, 1,ref Sample offset i in idx , and the weighting coefficient i for the sample at position (x,y) fact is determined. Here, the derivation for the vertical mode is provided. The derivation for the horizontal mode follows accordingly, with x and y swapped.
[0203]
number
[0204] ifact is not equal to 0, i.e., the prediction is exactly p 1,ref If p does not span a complete sample location in 1,ref The linear weighting between two adjacent sample locations in (0≦x, y <Nc)であって、
[0205]
number
[0206] It is executed as i idx and i fact Note that the value of ∇ ...
[0207] VTM-1.0 (Versatile Test Model) uses 35 intra modes, while BMS (Benchmark Set) uses 67 intra modes. Intra prediction is a mechanism used in many video coding frameworks to improve compression efficiency when only a given frame can be involved.
[0208] FIG. 10A shows an example of 67 intra prediction modes, for example as proposed for VVC, where several of the 67 intra prediction modes include a planar mode (index 0), a dc mode (index 1), and angular modes with indices 2 to 66, where the bottom-left angular mode in FIG. 10A refers to index 2 and the index numbering is incremented until index 66 is the top-right most angular mode in FIG. 10A.
[0209] As shown in Figure 10B, the latest version of VVC has several modes corresponding to diagonal intra-prediction directions, including a wide-angle mode (shown as a dashed line). In any of these modes, if the corresponding position within the block side is fractional, predicting a sample within the block interpolation of a set of neighboring reference samples should be performed. HEVC and VVC use linear interpolation between two neighboring reference samples. JEM uses a more sophisticated 4-tap interpolation filter. The filter coefficients are selected to be either Gaussian or cubic depending on the width or height value. The decision on whether to use width or height is coordinated with the decision on dominant reference side selection; that is, when the intra-prediction mode is diagonal or higher, the top side of the reference sample is selected to be the dominant reference side, and the width value is selected to determine the interpolation filter to use. Otherwise, the dominant side reference is selected from the left side of the block, and the height controls the filter selection process. Specifically, if the selected side length is 8 samples or less, cubic 4-tap interpolation is applied. Otherwise, the interpolation filter is a 4-tap Gaussian filter.
[0210] The specific filter coefficients used in JEM are given in Table 1. The predicted sample is calculated by convolving with a coefficient selected from Table 1 according to the subpixel offset and filter type as follows:
[0211]
number
[0212] In this expression, ">>" denotes a bitwise right shift operation.
[0213] The offset between the sample to be predicted (or "prediction sample" for short) in the current block and the interpolated sample position may have an integer part and a non-integer part if the offset has a sub-pixel resolution, such as 1 / 32 pixel. In Table 1 (Table 5), as well as Tables 2 (Table 6) and 3 (Table 7), the column "Sub-pixel Offset" refers to the non-integer part of the offset, e.g., a fractional offset, a fractional part of the offset, or a fractional sample position.
[0214] If a cubic filter is selected, the prediction samples are further clipped to a allowed range of values that is either specified in the SPS or derived from the bit depth of the selected component.
[0215] [Table 5]
[0216] Another set of interpolation filters with 6-bit precision is presented in Table 2.
[0217] [Table 6]
[0218] The intra-predicted samples are calculated by convolving with coefficients selected from Table 1 according to the sub-pixel offset and filter type as follows:
[0219]
number
[0220] In this expression, ">>" denotes a bitwise right shift operation.
[0221] Another set of interpolation filters with 6-bit precision is presented in Table 3.
[0222] [Table 7]
[0223] Figure 11 shows a schematic diagram of multiple intra-prediction modes used in the HEVC UIP scheme. For luma blocks, the intra-prediction modes may comprise up to 36 intra-prediction modes, which may include three non-directional modes and 33 directional modes. The non-directional modes may comprise a planar prediction mode, a mean (DC) prediction mode, and a chroma from luma (LM) prediction mode. The planar prediction mode may perform prediction by assuming a block deflection plane with horizontal and vertical gradients derived from the block boundaries. The DC prediction mode may perform prediction by assuming a flat block plane with values that match the mean values of the block boundaries. The LM prediction mode may perform prediction by assuming that the chroma values for the block match the luma values for the block. The directional modes may perform prediction based on neighboring blocks as shown in Figure 11.
[0224] H.264 / AVC and HEVC specify that a low-pass filter may be applied to reference samples before they are used in the intra prediction process. The decision on whether to use a reference sample filter is determined by the intra prediction mode and block size. This mechanism is sometimes called Mode Dependent Intra Smoothing (MDIS). There are also multiple methods related to MDIS. For example, the Adaptive Reference Sample Smoothing (ARSS) method may explicitly (i.e., a flag is included in the bitstream) or implicitly (i.e., data hiding is used, for example, to avoid placing a flag in the bitstream and reduce signaling overhead) signal whether the prediction samples are filtered. In this case, the encoder may make a decision on smoothing by testing the rate-distortion (RD) cost for all possible intra prediction modes.
[0225] As shown in Figure 10B, the latest version of VVC has several modes corresponding to diagonal intra-prediction directions. For any of these modes, if the corresponding position within the block side is fractional, prediction of a sample within the block interpolation of a set of neighboring reference samples should be performed. HEVC and VVC use linear interpolation between two neighboring reference samples. JEM uses a more sophisticated 4-tap interpolation filter. The filter coefficients are selected to be either Gaussian or cubic depending on the width or height value. The decision on whether to use width or height is coordinated with the decision on dominant reference side selection: when the intra-prediction mode is diagonal or higher, the top side of the reference sample is selected as the dominant reference side, and the width value is selected to determine the interpolation filter to use. Otherwise, the dominant side reference is selected from the left side of the block, and the height controls the filter selection process. Specifically, if the selected side length is 8 samples or less, cubic 4-tap interpolation is applied. Otherwise, the interpolation filter is a 4-tap Gaussian filter.
[0226] An example of interpolation filter selection for modes smaller and larger than the diagonal mode (shown as 45°) in the case of a 32×4 block is shown in FIG.
[0227] VVC uses a partitioning mechanism based on both quadtrees and binary trees, called QTBT. As shown in Figure 13, QTBT partitioning can provide square as well as rectangular blocks. Naturally, some signaling overhead and increased computational complexity on the encoder side are the price to pay for QTBT partitioning compared to the conventional quadtree-based partitioning used in the HEVC / H.265 standard. Nevertheless, QTBT-based partitioning is endowed with better segmentation properties and therefore demonstrates significantly higher coding efficiency than conventional quadtrees.
[0228] However, VVC in its current state applies the same filter to both sides (left and top sides) of the reference sample. Whether the block is oriented vertically or horizontally, the reference sample filter is the same for both reference sample sides.
[0229] In this specification, the terms "vertically oriented block" ("vertically oriented block") and "horizontally oriented block" ("horizontally oriented block") are applied to the rectangular blocks generated by the QTBT framework. These terms have the same meaning as shown in FIG. 14.
[0230] For directional intra-prediction modes with positive sub-sample offsets, it is necessary to determine the memory size used to store the values of reference samples. However, this size depends not only on the dimensions of the block of prediction samples, but also on the processing applied to these samples. Specifically, for positive sub-sample offsets, interpolation filtering will require an increased size of the primary reference side compared to the case when it is not applied. Interpolation filtering is performed by convolving the reference sample with a filter core. Therefore, the increase is caused by the additional samples required for the convolution operation to calculate the convolution results for the leftmost and rightmost parts of the primary reference side.
[0231] By using the steps described below, it is possible to determine the size of the primary reference side and therefore reduce the amount of internal memory required to store samples of the primary reference side.
[0232] 15A, 15B, 15C to 18 show some examples of intra prediction of a block from reference samples on the dominant reference side. For each row of samples of the block of prediction samples, a (possibly fractional) sub-pixel offset is determined. This offset is determined based on the selected directional intra prediction mode M and the orthogonal intra prediction mode M. oIt may have an integer or non-integer value depending on the difference between the vertex and the input (either HOR_IDX or VER_IDX, depending on which of them is closer to the selected intra-prediction mode).
[0233] Table 4 (Table 8) and Table 5 (Table 9) show the possible values of the sub-pixel offset for the first row of predicted samples depending on the mode difference. The sub-pixel offset for the other rows of predicted samples is obtained by multiplying the sub-pixel offset by the difference between the row position of the predicted sample and the first row.
[0234] [Table 8]
[0235] [Table 9]
[0236] When Table 4 (Table 8) or Table 5 (Table 9) is used to determine the sub-pixel offset for the bottom right prediction sample, it may be noted that the primary reference side size is equal to the sum of the integer part of the greatest or maximum sub-pixel offset, the size of the block side of the prediction sample (i.e., the block side length), and half the length of the interpolation filter (i.e., half the interpolation filter length), as shown in FIG. 15A.
[0237] To obtain the size of the dominant reference side for a selected directional intra-prediction mode that provides a positive value of the sub-pixel offset, the following steps may be performed.
[0238] 1. Step 1 may consist of determining which side of the block should be taken as the dominant side based on the index of the selected intra-prediction mode, and which neighboring samples should be used to generate the dominant reference side. The dominant reference side is the line of reference samples used in predicting samples in the current block. The "dominant side" is the side of the block to which the dominant reference side is parallel. If the (intra-prediction) mode is a diagonal mode (e.g., mode 34 as shown in FIG. 10A) or higher, the neighboring samples above (at the top of) the block being predicted (i.e., the current block) are used to generate the dominant reference side, and the top side is selected as the dominant side; otherwise, the neighboring samples to the left of the block being predicted are used to generate the dominant reference side, and the left side is selected as the dominant side. In summary, in step 1, the dominant side is determined for the current block based on the intra-prediction mode of the current block. Based on the dominant side, the dominant reference side is determined, including the reference samples (some or all of which) used for predicting the current block. For example, as shown in Figure 15A, the dominant reference side is parallel to the dominant (block) side, but may be longer than the block side. In other words, for example, given an intra-prediction mode, for each sample of the current block, a corresponding reference sample among multiple reference samples (e.g., dominant reference side samples) is also given.
[0239] 2. Step 2 may consist of determining a maximum sub-pixel offset, calculated by multiplying the length of the non-dominant side by the maximum value from either Table 4 or Table 5, so that the result of this multiplication represents a non-integer sub-pixel offset. Tables 4 and 5 provide exemplary values of sub-pixel offsets for samples in the first line of samples of the current block (either the top row of samples, corresponding to the top side selected as the dominant side, or the left-most column of samples, corresponding to the left side selected as the dominant side). Thus, the values shown in Tables 4 and 5 correspond to the sub-pixel offset per line of samples. The maximum offset that appears in the prediction of the entire block is therefore obtained by multiplying this per-line value by the length of the non-dominant side. In particular, in this example, the result should not be a multiple of 32 because the fixed-point resolution is 1 / 32 samples. If multiplication of any value per line, e.g., from Table 4 or Table 5, by the length of the non-dominant reference side results in a multiple of 32 corresponding to an integer total value (i.e., an integer number of samples) of sub-pixel offsets, then the result of this multiplication is discarded. The non-dominant side is the side of the block (either the top side or the left side) that was not selected in step 1. Thus, if the top side is selected as the dominant side, the length of the non-dominant side is the width of the current block, and if the left side is selected as the dominant side, the length of the non-dominant side is the height of the current block.
[0240] 3. Step 3 may consist of taking the integer part of the sub-pixel offset obtained in step 2, which corresponds to the multiplication result described above (i.e., by shifting right by 5 in binary representation), and summing it with the length of the dominant side (block width or block length, respectively) and half the length of the interpolation filter, resulting in the total value of the dominant reference side. The dominant reference side thus comprises a line of samples parallel to and equal in length to the dominant reference side, extended by adjacent samples within the non-integer part of the sub-pixel offset and further adjacent samples within half the length of the interpolation filter. Because the interpolation is performed over samples within the length of the inter-part of the sub-pixel offset and over the same amount of samples located beyond the length of the sub-pixel offset, only half the length of the interpolation filter is required.
[0241] According to another embodiment of the present disclosure, the reference samples being used to obtain the value of the predicted pixel are not adjacent to the block of prediction samples, and the encoder may signal an offset value in the bitstream, such that the offset value indicates the distance between the adjacent line of reference samples and the line of reference samples from which the value of the prediction sample is derived.
[0242] FIG. 24 shows the possible positions of the line of reference samples and the corresponding values of the ref_offset variable.
[0243] Example values of the offset in use in a particular implementation of a video codec (eg, a video encoder or a video decoder) are as follows: - Use the adjacent line of the reference sample (ref_offset=0, shown by "Reference Line 0" in Figure 24), - Use the first line (closest to the adjacent line) (ref_offset=1, shown by "Reference Line 1" in Figure 24), - Use the third line (ref_offset=3, indicated by "Reference Line 3" in Figure 24).
[0244] The variable "ref_offset" has the same meaning as the variable "refIdx" that will be used further. In other words, the variable "ref_offset" or the variable "refIdx" indicates the reference line, and for example, when ref_offset=0, it indicates that "reference line 0" (as shown in FIG. 24) is used.
[0245] The directional intra prediction mode specifies the value of the sub-pixel offset (deltaPos) between two adjacent lines of predicted samples. This value is represented by a fixed-point integer value with 5-bit precision. For example, deltaPos=32 means that the offset between two adjacent lines of predicted samples is exactly one sample.
[0246] If the intra-prediction mode is greater than DIA_IDX (mode #34), for the example described above, the value of the primary reference side size is calculated as follows: From the set of available intra-prediction modes (i.e., the mode that the encoder may indicate for the block of prediction samples) that is greater than DIA_IDX and provides the largest deltaPos value is considered. The value of the desired sub-pixel offset between the reference sample or interpolated sample position and the sample to be predicted is derived as follows: The block height is summed with ref_offset and multiplied by the deltaPos value. If the result is divided by 32 and the remainder is 0, another maximum value for deltaPos as described above, except that the previously considered prediction mode is skipped when deriving a mode from the set of available intra-prediction modes. Otherwise, the result of this multiplication is considered to be the maximum non-integer sub-pixel offset. The integer part of this offset is taken by shifting it right by 5 bits.
[0247] The size of the dominant reference side is obtained by summing the integer part of the maximum non-integer sub-pixel offset, the width of the block of prediction samples, and half the length of the interpolation filter (as shown in FIG. 15A).
[0248] Otherwise, if the intra-prediction mode is smaller than DIA_IDX (mode #34), for the example described above, the value of the dominant reference side size is calculated as follows: Among the set of intra-prediction modes available (i.e., that the encoder may indicate for the block of prediction samples), the mode that is smaller than DIA_IDX and provides the largest deltaPos value is considered. The value of the desired sub-pixel offset is derived as follows: The block width is summed with ref_offset and multiplied by the deltaPos value. If the result is divided by 32 with a remainder of 0, another maximum value of deltaPos as described above, except that previously considered prediction modes are skipped when deriving a mode from the set of available intra-prediction modes. Otherwise, the result of this multiplication is considered to be the maximum fractional sub-pixel offset. The integer part of this offset is taken by shifting it right by 5 bits. The size of the dominant reference side is obtained by summing the integer part of the maximum fractional sub-pixel offset, the height of the block of prediction samples, and half the length of the interpolation filter.
[0249] 15A, 15B, 15C to 18 show some examples of intra prediction of a block from reference samples on the dominant reference side. For each row of samples of the block of prediction samples 1120, a fractional sub-pixel offset 1150 is determined. This offset is used for the selected directional intra prediction mode M and the orthogonal intra prediction mode M. o It may have an integer or non-integer value depending on the difference between the vertex and the input (either HOR_IDX or VER_IDX, depending on which of them is closer to the selected intra-prediction mode).
[0250] State-of-the-art video coding methods and existing implementations of these methods take advantage of the fact that in the case of intra-angle prediction, the size of the dominant reference side is determined as twice the length of the corresponding block side. For example, in HEVC, when the intra-prediction mode is 34 or greater (see FIG. 10A or 10B), dominant reference side samples are taken from the upper and upper-right neighboring blocks if these blocks are available, i.e., not from an already reconstructed and processed slice. The total number of neighboring samples used is set equal to twice the width of the block. Similarly, when the intra-prediction mode is smaller than 34 (see FIG. 10), dominant reference side samples are taken from the left and lower-left neighboring blocks, and the total number of neighboring samples is set equal to twice the height of the block.
[0251] However, when applying the subpixel interpolation filter, additional samples to the left and right edges of the primary reference side are used. To maintain compatibility with existing solutions, it is proposed that these additional samples be obtained by padding the primary reference side to the left and right. The padding is performed by repeating the first and last samples of the primary reference side to the left and right sides, respectively. Denoting the primary reference side as ref and its size as refS, padding can be expressed as the following allocation operation: ref[-1]=p[0] ref[refS+1]=p[refS]
[0252] In practice, the use of negative indices can be avoided by applying a positive integer offset when referencing elements of the array. Specifically, this offset can be set equal to the number of elements to be left-padded on the major reference side.
[0253] Specific examples of how right and left padding should be performed are given in the following two cases shown by FIG. 15B.
[0254] The right padding case shows, for example, |MM equal to 22 for the wide-angle modes 72 and -6 (Fig. 10B). o | (see Table 4)
[0255]
number
[0256] When the block aspect ratio is 2 (i.e., when the prediction block dimensions are equal to 4x8, 8x16, 16x32, 32x64, 8x4, 16x8, 32x16, or 64x32), the corresponding maximum subpixel offset value is, for the bottom-right prediction sample,
[0257]
number
[0258] where S is the smaller side of the block.
[0259] So for an 8x4 block, the maximum sub-pixel offset is
[0260]
number
[0261] that is, the maximum value of the integer sub-pixel part of this offset is equal to 7. When applying a 4-tap intra-interpolation filter to obtain the value of the bottom-right sample with coordinates x=7, y=3, reference samples with indices x+7-1, x+7, x+7+1, and x+7+2 will be used. Because the primary reference side has 16 adjacent samples with indices 0..15, the rightmost sample position x+7+2=16, which means that one sample at the end of the primary reference side is padded by repeating the reference sample with position x+7+1.
[0262] The same steps are performed when Table 5 is in use for modes 71 and -5. The subpixel offsets for this case are
[0263]
number
[0264] is equal to
[0265]
number
[0266] The maximum value is obtained.
[0267] For example, the left padding case occurs for angle modes 35..65 and 19..33 when the subpixel offset is fractional and smaller than one sample. For the top-left predicted sample, the corresponding subpixel offset value is calculated. According to Table 4 and Table 5, this offset corresponds to an integer subsample offset of 0.
[0268]
number
[0269] Applying a 4-tap interpolation filter to calculate a predicted sample with coordinates x=0, y=0 will require reference samples with indices x-1, x, x+1, and x+2. The leftmost sample position x-1=-1. Because the main reference side has 16 adjacent samples with indices 0..15, the sample at this position is padded by repeating the reference sample with position x.
[0270] From the above example, for a block with an aspect ratio, the primary reference side is padded by half the 4-tap filter length, i.e., two samples, one of which is added to the beginning (left edge) of the primary reference side and the other is added to the end (right edge) of the primary reference side. In the case of 6-tap interpolation filtering, following the steps described above, two samples should be added to the beginning and end of the primary reference side. Generally, when an N-tap intra-interpolation filter is used, the primary reference side is
[0271]
number
[0272] padded with samples, of which
[0273]
number
[0274] Pieces are padded on the left side,
[0275]
number
[0276] are padded on the right side, where N is a non-negative even integer value.
[0277] Repeating the steps described above for other block aspect ratios, the following offsets are obtained (see Table 6):
[0278] [Table 10]
[0279] From the values given in Table 6, it follows that for wide-angle intra prediction modes: When Table 4 is in use, for 4-tap interpolation filtering, left and right padding operations are required for block sizes 4x8, 8x4, 8x16, and 16x8. When Table 5 is in use, in the case of 4-tap interpolation filtering, left and right padding operations are only required for block sizes 4x8 and 8x4.
[0280] Details of the proposed method are given in specification format in Table 7. The padding embodiment described above can be expressed as the following modifications to the VVC draft (section 8.2.4.2.7):
[0281] [Table 11]
[0282] Tables 4 and 5 as explained above represent the possible values of the sub-pixel offset between two adjacent lines of prediction samples depending on the intra prediction mode.
[0283] State-of-the-art video coding solutions use different interpolation filters in intra prediction. Specifically, Figures 19 to 21 show various examples of interpolation filters.
[0284] In the present invention, as shown in Figure 22 or Figure 23, an intra-prediction process for a block is performed, and during the intra-prediction process for a block, a sub-pixel interpolation filter is applied to the luminance and chrominance reference samples, and the sub-pixel interpolation filter (such as a 4-tap filter) is selected based on the sub-pixel offset between the position of the reference sample and the position of the sample to be interpolated, and the size of the main reference side used in the intra-prediction process is determined according to the length of the sub-pixel interpolation filter and the intra-prediction mode that results in the maximum value of the sub-pixel offset. The memory requirement is determined by the maximum value of the sub-pixel offset. The memory requirement is determined by the maximum value of the sub-pixel offset.
[0285] 15B shows the case when the top-left sample is not included in the primary reference side, but is instead padded using the leftmost sample that belongs to the primary reference side. However, if the prediction sample is calculated by applying a two-tap sub-pixel interpolation filter (e.g., a linear interpolation filter), the top-left sample is not referenced, and therefore no padding is required in this case.
[0286] 15C shows the case when a 4-tap subpixel interpolation filter (e.g., a Gaussian filter, a DCT-IF filter, or a cubic filter) is used. In this case, it can be noted that four reference samples are required to calculate at least the top-left predicted sample (marked as “A”): the top-left sample (marked as “B”) and the next three samples (marked as “C,” “D,” and “E,” respectively).
[0287] In this case, two alternative methods are disclosed.
[0288] Use the value of C to pad the value of B.
[0289] Using the reconstructed samples of the neighboring blocks in exactly the same way as the other samples of the primary reference side (including "B", "C", and "D") are obtained. In this case, the size of the primary reference side is: the block major side length (i.e., the block side length or block side size of the predicted sample); Half the interpolation filter length - 1, The following two values M: Chief of the Block, The integer part of the maximum subpixel offset + half the interpolation filter length, or the integer part of the maximum subpixel offset + half the interpolation filter length + 1 (the addition of 1 to this sum may or may not be included due to memory considerations). The largest of is determined as the sum of
[0290] It should be noted that "block major side", "block side length", "block major side length", and "side size of a block of prediction samples" are the same concepts throughout this disclosure.
[0291] It can be seen that half the interpolation filter length minus 1 is used to determine the size of the dominant reference side, thus allowing the dominant reference side from time to time to extend to the left.
[0292] It can be seen that the maximum of the two values M is used to determine the size of the primary reference side, and thus allows the primary reference side at any given time to extend to the right.
[0293] In the above description, the block dominant side length is determined according to the intra-prediction mode (FIG. 10B). If the intra-prediction mode is a diagonal intra-prediction mode (#34) or higher, the block dominant side length is the width of the block of prediction samples (i.e., the block to be predicted). Otherwise, the block dominant side length is the height of the block of prediction samples.
[0294] Sub-pixel offset values can be defined for a wider range of angles (see Table 8).
[0295] [Table 12]
[0296] Depending on the aspect ratio, different maximum and minimum values of the intra-prediction mode index (FIG. 10B) are allowed. Table 9 provides an example of this mapping.
[0297] [Table 13]
[0298] According to Table 9, the maximum modal difference value max(|MM o |), integer sub-pixel offsets are used for interpolation (the maximum sub-pixel offset per row is a multiple of 32), which means that the predicted samples of a predicted block are calculated by copying the values of the corresponding reference samples and no sub-sample interpolation filter is applied.
[0299] Table 9 (Table 13) max(|MM o Considering the constraints on | and the values in Table 8, the maximum sub-pixel offset per row that does not require interpolation is defined as follows (see Table 10):
[0300] [Table 14]
[0301] Using Table 10, for a square 4x4 block, the integer part of the maximum sub-pixel offset plus half the interpolation filter length value can be calculated using the following steps:
[0302] Step 1. The block major side length (equal to 4) is multiplied by 29 and the result is divided by 32, thus giving a value of 3.
[0303] Step 2. Half the length of the 4-tap interpolation filter is 2, which is added to the value obtained in step 1 to get the value 5.
[0304] From the above example, it can be observed that the obtained value is larger than the block major side length. In this example, the major reference side size is set to 10, which means Block major side length (equal to 4), Half the interpolation filter length -1 (equals 1), The following two values M: Block major side length (equal to 4), The integer part of the maximum subpixel offset plus half the interpolation filter length (equals 5), or the integer part of the maximum subpixel offset plus half the interpolation filter length + 1 (equals 6). (Due to memory considerations, the addition of 1 to this sum may or may not be included.) The largest of is determined as the sum of
[0305] The total number of reference samples included in the major reference side is greater than the block major side length doubled.
[0306] When the maximum of the two values M is equal to the block major side length, no right padding is performed. Otherwise, right padding is applied to reference samples whose positions are more than 2*nTbS (nTbS indicates the block major side length) horizontally or vertically away from the position of the top-left predicted sample (shown as "A" in Figure 15C). Right padding is performed by assigning the value of the padded sample to the last reference sample value on the major block side whose position is within 2*nTbS.
[0307] When half the interpolation filter length - 1 is greater than 0, either the value of sample "B" (shown in Figure 15C) is obtained by left padding, or the corresponding reference sample can be obtained in just the same way as reference samples "C", "D", and "E" are obtained.
[0308] The details of the proposed method are given in the specification format in Table 11. Instead of right padding or left padding, the corresponding reconstructed adjacent reference samples can be used. The case when left padding is not used can be represented by the following part of the VVC specification (Section 8.2):
[0309] [Table 15]
[0310] Similarly, using Table 10, the integer part of the maximum sub-pixel offset plus half the interpolation filter length for a non-square block that is 4 samples wide and 2 samples high can be calculated using the following steps (where the block major side length is the width):
[0311] Step 1. The block height (equal to 2) is multiplied by 57 and the result is divided by 32, thus giving a value of 3.
[0312] Step 2. Half the length of the 4-tap interpolation filter is 2, which is added to the value obtained in step 1 to get a value of 5.
[0313] The rest of the steps for calculating the total number of reference samples to be included in the major reference side are the same as in the square block case.
[0314] Using the block dimensions from Table 10 (Table 14) and Table 6 (Table 10), it can be noted that the maximum number of reference samples that undergo left or right padding is two.
[0315] If the block to be predicted is not adjacent to the adjacent reconstructed reference samples used in the intra prediction process (the reference lines can be selected as shown in Figure 24), the embodiments described below are applicable.
[0316] The first step is to define the aspect ratio of the block depending on the dominant side of the predicted block according to the intra prediction mode. If the top side of the block is chosen to be the dominant side, then the aspect ratio R a The aspect ratio R (denoted as "whRatio" in the VVC specification) is set equal to the result of integer division of the block width (denoted as "nTbW" in the VVC specification) by the block height (denoted as "nTbH" in the VVC specification). Otherwise, in the case when the dominant side is the left side of the predicted block, the aspect ratio R a (denoted as "hwRatio" in the VVC specification) is set equal to the result of the integer division of the block height by the block width. In either case, if the value of Ra is less than 1 (i.e., the numerator value of the integer division operator is less than the denominator value), it is set equal to 1.
[0317] The second step is to add a portion of the reference sample (denoted as "p" in the VVC specification) to the dominant reference side. Depending on the value of refIdx, either a neighboring reference sample or a non-neighboring reference sample is used. The reference sample to be added to the dominant reference side is selected using an offset relative to the dominant block side in the direction of the dominant side's orientation. Specifically, if the dominant side is the top side of the prediction block, the offset is horizontal and defined as -refIdx samples. If the dominant side is the left side of the prediction block, the offset is vertical and defined as -refIdx samples. In this step, nTbS+1 samples are added starting from the top-left reference sample (denoted as the "B" sample in Figure 15C) plus the value of the offset described above (nTbS indicates the dominant side length). Please note that the description or definition of RefIdx is presented in this disclosure in combination with Figure 24.
[0318] The next step that is performed depends on whether the sub-pixel offset (denoted as "intraPredAngle" in the VVC specification) is positive or negative. A 0 value for the sub-pixel offset corresponds to a horizontal intra-prediction mode (in the case when the dominant side of the block is the left block side) or a vertical intra-prediction mode (in the case when the dominant side of the block is the top block side).
[0319] If the subpixel offset is negative (e.g., step 3, negative subpixel offset), in the third step, the dominant reference side is extended to the left using the reference sample corresponding to the non-dominant side. The non-dominant side is the side that is not selected as the dominant side; that is, when the intra-prediction mode is 34 or higher (FIG. 10B), the non-dominant side is the left side of the block to be predicted; otherwise, the non-dominant side is the left side of the block. The extension is performed as shown in FIG. 7, and a description of this process can be found in the related description to FIG. 7. The reference sample corresponding to the non-dominant side is selected according to the process disclosed in the second step, with the difference that the non-dominant side is used instead of the dominant side. After this step is completed, the dominant reference side is extended from the beginning to the end using its first sample and last sample, respectively; in other words, in step 3, negative subpixel offset padding is performed.
[0320] If the sub-pixel offset is positive (e.g., step 3, positive sub-pixel offset), in the third step, the primary reference side is extended to the right by an additional nTbS samples in the same way as described in step 2. If the value of refIdx is greater than 0 (the reference sample is not adjacent to the block to be predicted), right padding is performed. The number of right-padded samples is equal to the aspect ratio Ra calculated in the first step multiplied by the refIdx value. If a 4-tap filter is in use, the number of right-padded samples is increased by 1.
[0321] The details of the proposed method are given in specification format in Table 12. The VVC specification modifications for this embodiment may be as follows (refW is set to nTbS-1):
[0322] [Table 16]
[0323] The above-described part of the VVC specification is also applicable to the case when the dominant reference side is left-padded by one sample in the third step for positive values of the sub-pixel offset. Details of the proposed method are given in Table 13 in the specification format.
[0324] [Table 17]
[0325] The present disclosure provides an intra-prediction method for predicting a current block included in a picture, such as a video frame. Method steps of the intra-prediction method are shown in Figure 25. The current block is the block described above that comprises samples to be predicted (or "predicted samples" or "prediction samples"), e.g., luminance samples or chrominance samples.
[0326] The method includes a step of determining (S2510) the size of the primary reference side based on the intra prediction mode among multiple available intra prediction modes (e.g., shown in Figures 10-11) that results in the largest non-integer value of the sub-pixel offset and the size (i.e., length) of the interpolation filter.
[0327] A sub-pixel offset is an offset between a sample in a current block to be predicted (or a "target sample") and a reference sample (or reference sample position) based on which the sample in the current block is predicted. If the reference sample includes a sample that is not directly or linearly above (e.g., a mode having a number greater than or equal to a diagonal mode) or to the left (e.g., a mode having a number less than or equal to a diagonal mode) of the current block, but includes a sample that is offset or shifted relative to the position of the current block, the offset may be related to the angular prediction mode. Because not all modes point to integer reference sample positions, the offset has sub-pixel resolution, and this sub-pixel offset may take a non-integer value and may have an integer part plus a non-integer part. In the case of a non-integer-valued sub-pixel offset, interpolation between reference samples is performed. Thus, the offset is the offset between the position of the sample to be predicted and the interpolated reference sample position. The maximum non-integer value may be the maximum non-integer value (integer part plus non-integer part) for any sample in the current block. For example, as shown in FIGS. 15A to 15C, the target sample associated with the maximum non-integer sub-pixel offset may be the bottom-right sample in the current block. Note that intra prediction modes that result in integer offsets greater than the maximum non-integer value of the sub-pixel offset are ignored.
[0328] Possible sizes (ie, lengths) of the interpolation filter include 4 (eg, the filter is a 4-tap filter) or 6 (eg, the filter is a 6-tap filter).
[0329] The method further includes applying an interpolation filter to reference samples included in the main reference side (S2520), and predicting target samples included in the current block based on the filtered reference samples (S2530).
[0330] In correspondence with the method shown in Figure 26, an apparatus 2600 for intra prediction of a current block included in a picture is also provided. The apparatus 2600 is shown in Figure 26 and may be included in the video encoder shown in Figure 2 or the video decoder shown in Figure 3. In one example, the apparatus 2600 may correspond to the intra prediction unit 254 in Figure 2. In another example, the apparatus 2600 may correspond to the intra prediction unit 354 in Figure 3.
[0331] The apparatus 2600 includes an intra prediction unit 2610 configured to predict a target sample included in a current block based on a filtered reference sample. The intra prediction unit 2610 may be the intra prediction unit 254 shown in FIG. 2 or the intra prediction unit 354 shown in FIG. 3.
[0332] The intra prediction unit 2610 includes a determination unit 2620 (or a "primary reference size determination unit") configured to determine the size of a primary reference side used in intra prediction. Specifically, the size is determined based on an intra prediction mode (among a plurality of available intra prediction modes) that results in the largest non-integer value of the sub-pixel offset between a target sample (among a plurality of target samples) in a current block and a reference sample (hereinafter referred to as a "target reference sample") used to predict the target sample in the current block, and based on the size of an interpolation filter to be applied to the reference sample included in the primary reference side. The target sample is any sample of the block to be predicted. The target reference sample is one of the reference samples of the primary reference side.
[0333] The intra prediction unit 2610 further comprises a filtering unit configured to apply an interpolation filter to the reference samples included in the dominant reference side to obtain filtered reference samples.
[0334] In summary, the maximum value of the sub-pixel offset determines the memory requirement. Therefore, by determining the size of the dominant reference side according to the present disclosure, the present disclosure facilitates memory efficiency in video coding using intra prediction. Specifically, the memory (buffer) used by the encoder and / or decoder to perform intra prediction can be allocated in an efficient manner according to the determined size of the dominant reference side. This is because, first, the size of the dominant reference side determined according to the present disclosure includes all reference samples to be used to predict the current block. Therefore, access to additional samples is not required to perform intra prediction. Second, this is not necessary for all already processed samples of neighboring blocks, but rather, the memory size may be allocated specifically to the dominant reference side, i.e., specifically, those reference samples belonging to the determined size.
[0335] The following is a description of an example application of the encoding method and the decoding method as shown in the above embodiment and the system using them.
[0336] 27 is a block diagram showing a content supply system 3100 for realizing a content distribution service. The content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.
[0337] The capture device 3102 may generate data and encode the data using an encoding method such as that shown in the above embodiment. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown), which then encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 may include, but is not limited to, a camera, a smartphone or pad, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 described above. When the data includes video, a video encoder 20 included in the capture device 3102 may actually perform the video encoding process. When the data includes audio (i.e., voice), an audio encoder included in the capture device 3102 may actually perform the audio encoding process. In some practical scenarios, the capture device 3102 delivers the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 delivers the encoded audio data and the encoded video data separately to the terminal device 3106 .
[0338] In the content delivery system 3100, a terminal device 3106 receives and plays encoded data. The terminal device 3106 may be a device having data reception and recovery capabilities, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, capable of decoding the above-mentioned encoded data. For example, the terminal device 3106 may include the destination device 14 described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0339] In the case of a terminal device having its display, for example, a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can provide the decoded data to its display. In the case of a terminal device not equipped with a display, for example, an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is brought into contact with them to receive and display the decoded data.
[0340] When each device in this system performs encoding or decoding, it may use a picture encoding device or a picture decoding device as shown in the above embodiments.
[0341] 28 is a diagram illustrating the structure of an example of a terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, a protocol progression unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), HyperText Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real Time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any kind of combination thereof.
[0342] After the protocol progression unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As explained above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0343] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. The video decoder 3206, which includes the video decoder 30 as described in the above embodiment, decodes the video ES by the decoding method as shown in the above embodiment to generate video frames, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212.
[0344] The synchronization unit 3212 synchronizes video and audio frames and provides the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video with audio information, which may be coded in the syntax using timestamps related to the presentation of the coded audio and visual data, as well as timestamps related to the delivery of the data stream itself.
[0345] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.
[0346] The present invention is not limited to the above-mentioned system, and any of the picture encoding devices or picture decoding devices in the above-mentioned embodiments may be incorporated into other systems, for example, automobile systems.
[0347] Although embodiments of the present invention are described primarily in the context of video coding, it should be noted that embodiments of coding system 10, encoder 20, and decoder 30 (and correspondingly, system 10), as well as other embodiments described herein, may also be configured for still image processing or coding, i.e., processing or coding of individual pictures independent of any preceding or subsequent pictures, as in video coding. Generally, when picture processing coding is limited to a single picture 17, only inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functionality (also called tools or techniques) of the video encoder 20 and the video decoder 30 may be equally used for still image processing, e.g., residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, partitioning 262 / 362, intra prediction 254 / 354 and / or loop filtering 220, 320, as well as entropy coding 270 and entropy decoding 304.
[0348] The embodiments of, and functions described herein with reference to, for example, the encoder 20 and the decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted over a communication medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media or communication media including, for example, any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include computer-readable media.
[0349] By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0350] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques may be implemented entirely within one or more circuits or logic elements.
[0351] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by various hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, along with suitable software and / or firmware. [Explanation of symbols]
[0352] 10 Video coding system, coding system 12 Source Devices 13 Communication Channels 14 Destination Device 16 Picture Source 17 Picture, Picture Data, Raw Picture, Raw Picture Data 18 Preprocessor, preprocessing unit, picture preprocessor 19 Preprocessed Picture, Preprocessed Picture Data 20 Video Encoder, Encoder 21 Encoded Picture Data, Bitstream, Encoded Bitstream 22, 28 Communication interface, communication unit 30 Video decoder, decoder 31 Decoded Picture, Decoded Picture Data 32 Post-processor, post-processing unit 33 Post-processed picture, post-processed picture data 34 Display Devices 40 Video Coding System 41 Imaging Devices 42 Antenna 43 processors 44 Memory Store 45 Display Devices 46 Processing Unit 47 Logic Circuit Configuration 201 Input unit, input interface 203 Picture Block 204 Residual Calculation Unit 205 Residual Block, Residual 206 Conversion Processing Unit 207 Conversion Factor 208 quantization units 209 Quantized Coefficients, Quantized Transform Coefficients, Quantized Residual Coefficients 210 Inverse Quantization Unit 211 Inverse quantization coefficients, inverse quantization residual coefficients 212 Inverse Transformation Processing Unit 213 Reconstructed residual block, transform block 214 Reconstruction Unit 215 Reconstruction Block 216 buffers 220 Loop filter, loop filter unit 221 Filtered Block, Filtered Reconstructed Block 230 Decoded Picture Buffer (DPB) 231 Decoded Picture 244 Inter Prediction Units 254 intra prediction units 260 Mode Selection Unit 262 division units 265 predicted blocks 266 Syntax Elements 270 Entropy Coding Unit 272 Output section, output interface 304 Entropy Decoding Unit 309 Quantization Coefficients 310 Inverse Quantization Unit 311 Transform coefficients, inverse quantization coefficients 312 Inverse Transformation Processing Unit 313 Transform Block, Reconstructed Residual Block 314 Reconstruction Unit 315 Reconstruction Block 320 Loop filter, loop filter unit 321 Filtered Blocks 330 Decoded Picture Buffer (DBP) 331 Decoded Picture 344 Inter Prediction Unit 354 intra prediction units 360 Mode Selection Unit 365 predicted blocks 400 Video Coding Device 410 inlet port, input port 420 Receiver Unit (Rx) 430 Processor, Logic Unit, Central Processing Unit (CPU) 440 Transmitter Unit (Tx) 450 outlet port, output port 460 memory 470 Coding Module 500 devices 502 processor 504 memory 506 Data 508 Operating Systems 510 Application Program 512 Bus 514 Secondary Storage 518 Display 520 Image sensing device 522 Sound sensing device 2600 equipment 2610 Intra Prediction Unit 2620 Decision Unit 3100 Contents Supply System 3102 Capture Device 3104 Communication Links 3106 Terminal Device 3108 Smartphones, Pads 3110 Computers, Laptops 3112 Network Video Recorder (NVR), Digital Video Recorder (DVR) 3114 TV 3116 Set-top box (STB) 3118 Video Conference System 3120 Video Surveillance System 3122 Personal Digital Assistant (PDA) 3124 In-Vehicle Devices 3126 Display 3202 Protocol Progression Unit 3204 Demultiplexing Unit 3206 Video Decoder 3208 Audio Decoder 3210 Subtitle Decoder 3212 Synchronous Unit 3214 Video / Audio Display 3216 Video / Audio / Subtitle Display
Claims
1. 1. A method of video coding, comprising: performing an intra prediction process for a block comprising samples to be predicted, wherein an interpolation filter is applied to reference samples of said block during said intra prediction process for said block; the interpolation filter is selected based on a sub-pixel offset between the reference sample and the sample to be predicted; a size of a dominant reference side used in the intra prediction process is determined according to a length of the interpolation filter and an intra prediction mode that results in a maximum non-integer value of the sub-pixel offset among a set of available intra prediction modes, the dominant reference side comprising the reference samples; method.
2. The size of the major reference side is: the integer portion of the maximum non-integer value of the sub-pixel offset; the size of the side of the block; a portion or the entire length of the interpolation filter; determined as the sum of The method of claim 1.
3. if the intra prediction mode is greater than a vertical intra prediction mode VER_IDX, the side of the block of prediction samples is the width of the block; or If the intra prediction mode is smaller than the horizontal intra prediction mode HOR_IDX, the side of the block is the height of the block. The method of claim 2.
4. 4. The method of claim 2 or 3, wherein the value of a reference sample in the main reference side that is located at a position greater than twice the size of the side of the block is set to be equal to the value of a sample located at twice the size of the side of the block.
5. The size of the major reference side is: - the size of the sides of the block; half the length of the interpolation filter minus 1; The following two values M: - the size of the side of the block; and - the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter The largest of determined as the sum of The method of claim 1.
6. The size of the major reference side is: - the size of the sides of the block; half the length of the interpolation filter minus 1; The following two values M: - the size of the side of the block; and - the integer part of the maximum non-integer value of the sub-pixel offset + half the length of the interpolation filter + 1 The largest of determined as the sum of The method of claim 1.
7. when the maximum of the two values M is equal to the size of a side of the block, no right padding is performed, or right-padding is performed when the maximum of the two values M is equal to the integer part of the maximum value of the sub-pixel offset plus half the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1; The method of claim 5.
8. The padding is performed by repeating the first reference sample and / or the last reference sample of the primary reference side on the left side and / or the right side, respectively, specifically as follows: denoting the primary reference side as ref and the size of the primary reference side as refS: ref[-1]=p[0] and / or ref[refS+1]=p[refS] The padding is expressed as ref[-1] represents the left value of the primary reference side; p[0] represents the value of the first reference sample of the primary reference side; ref[refS+1] represents the right value of the primary reference side; p[refS] represents the value of the last reference sample of the primary reference side; 8. The method according to any one of claims 1 to 7.
9. 9. The method of claim 1, wherein the interpolation filters used in the intra prediction process are finite impulse response filters, the coefficients of which are fetched from a look-up table.
10. 10. The method of claim 1, wherein the interpolation filter used in the intra prediction process is a 4-tap filter.
11. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 but as follows, i.e., 【Table 1】 where the "non-integer part of sub-pixel offset" column is defined at 1 / 32 sub-pixel resolution, depending on the non-integer part of the sub-pixel offset, such as The method of claim 10.
12. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 but as follows, i.e., 【Table 2】 where the "non-integer part of sub-pixel offset" column is defined at 1 / 32 sub-pixel resolution, depending on the non-integer part of the sub-pixel offset, such as The method of claim 10.
13. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 but as follows, i.e., 【Table 3】 where the "non-integer part of sub-pixel offset" column is defined at 1 / 32 sub-pixel resolution, depending on the non-integer part of the sub-pixel offset, such as The method of claim 10.
14. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 but as follows, i.e., 【Table 4】 where the "non-integer part of sub-pixel offset" column is defined at 1 / 32 sub-pixel resolution, depending on the non-integer part of the sub-pixel offset, such as The method of claim 10.
15. 15. The method of claim 1, wherein the interpolation filter is selected from a set of filters used for the intra prediction process for a given sub-pixel offset.
16. The method of claim 15 , wherein the set of filters comprises a Gaussian filter and a cubic filter.
17. 17. The method of claim 1, wherein the number of interpolation filters is N, and the N interpolation filters are used for intra-reference sample interpolation, where N>=1 and is a positive integer.
18. 18. The method of claim 1, wherein the reference samples include samples that are not adjacent to the block.
19. 1. An intra prediction method for predicting a current block included in a picture, comprising: an intra-prediction mode among a plurality of available intra-prediction modes that results in the largest non-integer value of a sub-pixel offset between a target sample among a plurality of target samples in the current block and a reference sample used to predict the target sample in the current block, the reference sample being among a plurality of reference samples included in a dominant reference side; and based on the size of an interpolation filter to be applied to the reference samples included in the major reference side, determining a size of the dominant reference side used in the intra prediction; applying an interpolation filter to the reference samples included in the dominant reference side to obtain filtered reference samples; predicting the target samples contained in the current block based on the filtered reference samples; A method for providing
20. The size of the major reference side is: the integer portion of the maximum non-integer value of the sub-pixel offset; The size of a side of the current block; half the size of the interpolation filter; determined as the sum of 20. The method of claim 19.
21. If the intra prediction mode is greater than a vertical intra prediction mode VER_IDX, the side of the current block is the width of the current block; or If the intra prediction mode is smaller than the horizontal intra prediction mode HOR_IDX, the side of the current block is the height of the current block.
21. The method of claim 20.
22. a value of a reference sample having a position within the primary reference side that is greater than twice the size of the side of the current block is set equal to a value of a sample having a sample position that is twice the size of the current block; 22. The method of claim 20 or 21.
23. The size of the major reference side is: - the size of the side of the current block; half the length of the interpolation filter minus 1; The following two values M: - the size of the block sides, and - the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter The largest of determined as the sum of 20. The method of claim 19.
24. The size of the major reference side is: - the size of the side of the current block; half the length of the interpolation filter minus 1; The following, namely: - the size of the block sides, and - the integer part of the maximum non-integer value of the sub-pixel offset + half the length of the interpolation filter + 1 The largest of determined as the sum of 20. The method of claim 19.
25. when the maximum of the two values M is equal to the size of a side of the block, no right padding is performed, or right-padding is performed when the maximum of the two values M is equal to the integer part of the maximum value of the sub-pixel offset plus half the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1; 24. The method of claim 23.
26. The padding is performed by repeating the first reference sample and / or the last reference sample of the primary reference side on the left side and / or the right side, respectively, specifically as follows: denoting the primary reference side as ref and the size of the primary reference side as refS: ref[-1]=p[0] and / or ref[refS+1]=p[refS] The padding is expressed as ref[-1] represents the left value of the primary reference side; p[0] represents the value of the first reference sample of the primary reference side; ref[refS+1] represents the right value of the primary reference side; p[refS] represents the value of the last reference sample of the primary reference side; 26. The method of any one of claims 19 to 25.
27. 27. An encoder comprising processing circuitry configured to perform the method of any one of claims 1 to 26.
28. 27. A decoder comprising processing circuitry configured to perform the method of any one of claims 1 to 26.
29. 27. A computer program product comprising program code for carrying out the method of any one of claims 1 to 26.
30. A decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder to perform the method of any one of claims 1 to 26. decoder.
31. 1. An encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the encoder to perform the method of any one of claims 1 to 26. Encoder.
32. 1. An apparatus for intra prediction of a current block included in a picture, comprising: an intra prediction unit configured to predict a plurality of target samples included in the current block based on filtered reference samples, the intra prediction unit comprising: an intra-prediction mode among a plurality of available intra-prediction modes that results in the largest non-integer value of a sub-pixel offset between a target sample among the plurality of target samples in the current block and a reference sample used to predict the target sample in the current block, the reference sample being among a plurality of reference samples included in a dominant reference side; and based on the size of an interpolation filter to be applied to the reference samples included in the major reference side, a determining unit configured to determine a size of the dominant reference side used in the intra prediction; a filtering unit configured to apply an interpolation filter to the reference samples included in the dominant reference side to obtain the filtered reference samples. Device.
33. The decision unit: the integer portion of the maximum non-integer value of the sub-pixel offset; The size of a side of the current block; half the size of the interpolation filter; configured to determine the size of the primary reference side as the sum of 33. The apparatus of claim 32.
34. If the intra prediction mode is greater than a vertical intra prediction mode VER_IDX, the side of the current block is the width of the current block; or If the intra prediction mode is smaller than the horizontal intra prediction mode HOR_IDX, the side of the current block is the height of the current block.
34. The apparatus of claim 33.
35. a value of a reference sample having a position within the primary reference side that is greater than twice the size of the side of the current block is set equal to a value of a sample having a sample position that is twice the size of the current block; 35. Apparatus according to claim 33 or 34.
36. The decision unit: - the size of the side of the current block; half the length of the interpolation filter minus 1; The following, namely: - the size of the block sides, and - the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter The largest of configured to determine the size of the primary reference side as the sum of 33. The apparatus of claim 32.
37. The decision unit: - the size of the side of the current block; half the length of the interpolation filter minus 1; The following, namely: - the size of the block sides, and - the integer part of the maximum non-integer value of the sub-pixel offset + half the length of the interpolation filter + 1 The largest of configured to determine the size of the primary reference side as the sum of 33. The apparatus of claim 32.
38. The decision unit: no right padding is performed when the maximum of the two values M is equal to the size of a side of the block; or configured to perform right padding when the maximum of the two values M is equal to the integer part of the maximum value of the sub-pixel offset plus half the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset plus half the length of the interpolation filter + 1.
37. The apparatus of claim 36.
39. The determining unit is configured to perform padding by repeating the first reference sample and / or the last reference sample of the primary reference side to the left side and / or the right side, respectively, specifically as follows: denoting the primary reference side as ref and the size of the primary reference side as refS: ref[-1]=p[0] and / or ref[refS+1]=p[refS] The padding is expressed as ref[-1] represents the left value of the primary reference side; p[0] represents the value of the first reference sample of the primary reference side; ref[refS+1] represents the right value of the primary reference side; p[refS] represents the value of the last reference sample of the primary reference side; 39. Apparatus according to any one of claims 32 to 38.
40. 40. A video encoder for encoding a plurality of pictures into a bitstream, comprising an apparatus according to any one of claims 32 to 39.
41. A video decoder for decoding a plurality of pictures from a bitstream, comprising an apparatus according to any one of claims 32 to 39.
42. 27. A non-transitory computer-readable medium storing computer instructions for performing intra prediction that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 26.
Citation Information
Patent Citations
Method and apparatus for image interpolation using smoothing interpolation filter
JP2013542666A
Low-complexity interpolation filtering with adaptive tap size
JP2014502822A
Interpolation filters for intra prediction in video coding
US20180091825A1