Method and apparatus for intra prediction using an interpolation filter
By applying a sub-pixel interpolation filter based on the sub-pixel offset and determining the primary reference side size according to the interpolation filter length and intra prediction mode, the method enhances video coding efficiency and reduces memory requirements.
Patent Information
- Application Number
- JP2023045836
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-07
- Filing Date
- 2023-03-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-10-07
AI Technical Summary
Current video coding standards face challenges in balancing bandwidth requirements and video quality, particularly in handling high-resolution videos that require larger data streams.
The method involves performing an intra prediction process in video coding, where a sub-pixel interpolation filter is applied to reference samples based on a sub-pixel offset, and the size of the primary reference side is determined according to the length of the interpolation filter and the intra prediction mode.
This approach simplifies the calculation procedure for intra prediction, improving coding efficiency and reducing memory requirements by optimizing the size of the primary reference side.
Smart Images

Figure 0007684344000034 
Figure 0007684344000035 
Figure 0007684344000036
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 742,300, filed on October 6, 2018, U.S. Provisional Patent Application No. 62 / 744,096, filed on October 10, 2018, U.S. Provisional Patent Application No. 62 / 753,055, filed on October 30, 2018, and U.S. Provisional Patent Application No. 62 / 757,150, filed on November 7, 2018. This is a divisional application of Japanese Patent Application No. 2021-518698 The above - mentioned patent applications are hereby incorporated by reference in their entirety into this specification.
[0002] The present disclosure relates to the technical field of image and / or video coding and decoding, and more particularly, to a method and apparatus for directional intra - prediction involving reference sample processing coordinated with the length of an interpolation filter.
Background Art
[0003] Since the introduction of DVD discs, digital video has been widely used. Before transmission, video is encoded and transmitted using a transmission medium. Viewers receive the video and use a viewing device to decode and display the video. Over the years, the quality of video has been improved, for example, by higher resolution, color depth, and frame rate. This has led to larger data streams that are typically transported today via the Internet and mobile communication networks.
[0004] However, videos with higher resolution usually have more information and thus require a larger bandwidth. To reduce the bandwidth requirements, video coding standards involving video compression have been introduced. When a video is encoded, the bandwidth requirements (or in the case of storage, the corresponding memory requirements) are reduced. Often, this reduction is achieved at the expense of quality. Thus, video coding standards attempt to find a balance between bandwidth requirements and quality.
[0005] High Efficiency Video Coding (HEVC) is an example of a video coding standard generally known to those skilled in the art. In HEVC, in order to divide a coding unit (CU) into a prediction unit (PU) or a transform unit (TU). The Versatile Video Coding (VVC) next-generation standard is the most recent joint video project of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standardization bodies, collaborating in a partnership called the Joint Video Exploration Team (JVET). VVC is also known as the ITU-T H.266 / Next Generation Video Coding (NGVC) standard. In VVC, the concept of multiple partition types, i.e., the separation of CUs, PUs, and TUs, is removed as needed, excluding cases where the CU is too large for the maximum transform length, and greater flexibility is supported for the CU partition shape.
[0006] The processing of these coding units (CUs) (also called blocks) depends on their size, spatial position, and coding mode specified by the encoder. The coding mode can be classified into two groups according to the type of prediction, i.e., intra prediction mode and inter prediction mode. The intra prediction mode uses samples of the same picture (also called frame or image) to generate reference samples for calculating prediction values for samples of the block being reconstructed. Intra prediction is also called spatial prediction. The inter prediction mode is designed for temporal prediction and predicts samples of a block in the current picture using reference samples from the previous or next picture.
[0007] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC29 / WG11) are considering the potential needs for the standardization of future video coding technologies that have a compression capability significantly exceeding that of the current HEVC standard (including its current extensions and near-future extensions for screen content coding and high dynamic range coding). That group is collaborating in this exploration activity within a joint research effort called the Joint Video Exploration Team (JVET) to evaluate the compression technology designs proposed by their experts in this area.
[0008] The VTM (Versatile Test Model) standard uses 35 intra modes, while the BMS (Benchmark Set) uses 67 intra modes.
[0009] The intra mode coding method currently described in BMS is considered complex, and the fact that the index list is not always constant (e.g., for its adjacent block intra mode) and not adaptive based on the current block characteristics is a drawback of the unselected mode set. Summary of the Invention Means for Solving the Problems
[0010] Embodiments of the present application are disclosed that provide an apparatus and method for intra prediction. The apparatus and method simplify the calculation procedure for intra prediction using a mapping process so as to improve coding efficiency. The scope of protection is defined by the claims.
[0011] The above and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the description, and the figures.
[0012] Certain embodiments are outlined in the appended independent claims, together with other embodiments in the dependent claims.
[0013] According to a first aspect, the present invention relates to a method of video coding. The method is performed by an encoding device or a decoding device. The method includes performing an intra prediction process of a block, such as a block comprising samples to be predicted or a block of prediction samples, specifically, such as a luma block comprising luma samples to be predicted, and during the intra prediction process of the block, a sub-pixel interpolation filter is applied to a reference sample (for example, a luminance reference sample), or during the intra prediction process of the block, a sub-pixel interpolation filter is applied to a reference sample (for example, a chrominance reference sample), wherein the sub-pixel interpolation filter is selected based on a sub-pixel offset, for example, between the position of the reference sample and the position of the interpolated sample, or between the reference sample and the sample to be predicted, and the size of the main reference side used in the intra prediction process is determined according to the length of the sub-pixel interpolation filter and the intra prediction mode (for example, the intra prediction mode among the set of available intra prediction modes) that results in the maximum value of the sub-pixel offset (for example, the maximum non-integer value), and the main reference side comprises reference samples.
[0014] A reference sample is a sample on which a prediction (here, in particular, an intra prediction) is performed. In other words, a reference sample is a sample outside the (current) block that is used to predict the sample of the (current) block. The term "current block" indicates the target block on which the processing including the prediction is performed. For example, a reference sample is a sample adjacent to the block at one or more of the block sides. In other words, the reference sample used to predict the current block may be included in a line of samples that at least partially adjoins one or more block boundaries (sides) and is parallel to one or more block boundaries (sides).
[0015] A reference sample may be a sample at an integer sample position or an interpolated sample at a sub-sample position, for example, a non-integer position. The integer sample position may refer to an actual sample position in the image to be coded (encoded or decoded).
[0016] The reference side is the side of the block from which reference samples are used to predict the samples of the block. The primary reference side is the side of the block from which the reference samples are taken (in some embodiments, there is only one side from which the reference samples are taken). However, generally, the primary reference side may refer to the side from which the reference samples are mainly taken (e.g., most of the reference samples are taken from it, or the reference samples for predicting most of the block samples are taken from it). The primary reference side includes the reference samples used to predict the samples of the block. If the primary reference side consists of the reference samples used to predict the samples of the block and all of those reference samples used to predict the samples of the block are included within the primary reference side, that can be advantageous for memory savings purposes. However, the present disclosure also generally applies with the primary reference side that includes the reference samples used to predict the block. These may comprise the reference samples directly used for prediction, as well as the reference samples used for filtering to obtain sub-samples that are then used for prediction of the block samples.
[0017] Generally, the reference samples of the current block comprise the adjacent reconstructed samples of the current block. Thus, if the current block is the current chroma block, the chroma reference samples of the current chroma block comprise the adjacent reconstructed samples of the current chroma block. Thus, if the current block is the current luma block, the luma reference samples of the current luma block comprise the adjacent reconstructed samples of the current luma block.
[0018] It is understood that the memory requirement is determined by the maximum value of the sub-pixel offset. Therefore, by determining the size of the primary reference side according to the present disclosure, the present disclosure facilitates bringing about memory efficiency in video coding using intra prediction. In other words, by determining the size of the primary reference side used in the intra prediction process according to the first aspect described above, the memory requirement can be reduced while providing (storing) reference samples for predicting blocks. This can in turn lead to a more efficient implementation of intra prediction for image / video encoding and decoding.
[0019] In a possible implementation of the method according to such a first aspect, the interpolation filter is selected based on the sub-pixel offset between the position of the reference sample and the position of the predicted sample.
[0020] It is understood that the predicted samples are interpolated samples in that they are based on the output of the interpolation process.
[0021] In a possible implementation of the method according to such a first aspect, the sub-pixel offset is determined based on a reference line (such as refIdx), or the sub-pixel offset is determined based on intraPredAngle that depends on the selected intra prediction mode, or the sub-pixel offset is determined based on the distance between the side of the block of the reference sample and the predicted sample, i.e., from the side of the block of the reference sample (such as the reference line) to the side of the block of the predicted sample.
[0022] In a possible implementation of the method according to such a first aspect, the maximum value of the sub-pixel offset is the maximum non-integer sub-pixel offset (such as the maximum fractional sub-pixel offset or the maximum non-integer value of the sub-pixel offset), and the size of the primary reference side is the integer part of the maximum non-integer sub-pixel offset, and The size of the side of the block of prediction samples is selected to be equal to the sum with a part or the whole of the length of the interpolation filter (such as half of the length of the interpolation filter). (Such as half of the length of the interpolation filter) is selected to be equal to the sum with a part or the whole of the length of the interpolation filter.
[0023] One of the advantages of such a selection of the size of the main reference side is the preparation (storing / buffering) of all the samples required for the intra prediction of the block, and the reduction in the number of samples that are not used (stored / buffered) for predicting the block (of samples).
[0024] In a possible implementation of the method according to such a first aspect, when the intra prediction mode is greater than the vertical intra prediction mode (VER_IDX), the side of the block of prediction samples is the width of the block of prediction samples, or when the intra prediction mode is less than the horizontal intra prediction mode (HOR_IDX), the side of the block of prediction samples is the height of the block of prediction samples.
[0025] For example, in FIG. 10, VER_IDX corresponds to the vertical intra prediction mode #50, and HOR_IDX corresponds to the horizontal intra prediction mode #18.
[0026] In a possible implementation of the method according to such a first aspect, the reference samples of the main reference side having a position greater than twice the size of the block side are set to be equal to the samples located at twice the size of the size.
[0027] In other words, this is padding to the right by repeating the pixels that exceed twice the side length. It is preferable that the memory buffer size is a power of 2, and it is better to use the last sample of a buffer of a size that is a power of 2 (that is, located at twice the size of the size) than to hold a buffer of a size that is not a power of 2.
[0028] In one possible implementation of the method according to such a first aspect, the size of the main reference side part is the block main side length, and (such as the length of the interpolation filter or half of the length of the interpolation filter) a part or the whole of the length of the interpolation filter - 1, and the following two values M, namely the block main side length, the integer part of the maximum (or greatest) non-integer sub-pixel offset + a part or the whole of the length of the interpolation filter (such as half of the length of the interpolation filter), or the integer part of the maximum (or greatest) non-integer sub-pixel offset + a part or the whole of the length of the interpolation filter (such as half of the length of the interpolation filter) + 1, is determined as the sum of the maximum of these.
[0029] One of the advantages of such a selection of the size of the main reference side part is to reduce or even avoid the preparation (storing / buffering) of all the samples required for the intra prediction of the block, and the preparation (storing / buffering) of the samples not used for predicting the block (of the samples).
[0030] Note that "block main side part", "block side length", "block main side length", and "size of the side part of the block of prediction samples" are the same concept throughout this disclosure.
[0031] In one possible implementation of the method according to such a first aspect, when the maximum of the two values M is equal to the block main side length, right padding is not performed, or when the maximum of the two values M is equal to the integer part of the maximum non-integer sub-pixel offset + half of the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter + 1, right padding is performed.
[0032] In one possible implementation, padding is performed by repeating the first reference sample and / or the last reference sample of the main reference side portion on the left side portion and / or the right side portion respectively. Specifically, as follows: That is, when the main reference side portion is denoted as ref and the size of the main reference side portion is denoted as refS, padding is represented as ref[-1]=p[0] and / or ref[refS + 1]=p[refS]. ref[-1] represents the left value of the main reference side portion, p[0] represents the value of the first reference sample of the main reference side portion, ref[refS + 1] represents the right value of the main reference side portion, p[refS] represents the value of the last reference sample of the main reference side portion.
[0033] In other words, right padding can be performed by ref[refS + 1]=p[refS]. Additionally or alternatively, left padding can be performed by ref[-1]=p[0].
[0034] In this way, padding can facilitate the preparation of all samples necessary for prediction, also taking into account interpolation filtering.
[0035] In one possible implementation form of the method according to such a first aspect, the filters used in the intra prediction process are finite impulse response filters, and their coefficients are fetched from a look-up table.
[0036] In one possible implementation form of the method according to such a first aspect, the interpolation filter used in the intra prediction process is a 4-tap filter.
[0037] In one possible implementation form of the method according to such a first aspect, the coefficients of the interpolation filter are as follows: That is,
[0038]
Table 1
[0039] As such, the "sub-pixel offset" column, which depends on the sub-pixel offset such as the non-integer part of the sub-pixel offset, is defined at a 1 / 32 sub-pixel resolution. In other words, an interpolation filter (such as a sub-pixel interpolation filter) is represented by the coefficients in the above table.
[0040] In a possible implementation of the method according to such a first aspect, the coefficients of the interpolation filter are as follows, that is,
[0041] [Table 2]
[0042] As such, the "sub-pixel offset" column, which depends on the sub-pixel offset such as the non-integer part of the sub-pixel offset, is defined at a 1 / 32 sub-pixel resolution. In other words, an interpolation filter (such as a sub-pixel interpolation filter) is represented by the coefficients in the above table.
[0043] In a possible implementation of the method according to such a first aspect, the coefficients of the interpolation filter are as follows, that is,
[0044] [Table 3]
[0045] As such, the "sub-pixel offset" column, which depends on the sub-pixel such as the non-integer part of the sub-pixel offset and the offset, is defined at a 1 / 32 sub-pixel resolution. In other words, an interpolation filter (such as a sub-pixel interpolation filter) is represented by the coefficients in the above table.
[0046] In a possible implementation of the method according to such a first aspect, the coefficients of the interpolation filter are as follows, that is,
[0047] [Table 4]
[0048] As such, the "sub-pixel offset" column is defined at a 1 / 32 sub-pixel resolution and depends on the sub-pixel offset, such as the fractional part of the sub-pixel offset. In other words, an interpolation filter (such as a sub-pixel interpolation filter) is represented by the coefficients in the above table.
[0049] In a possible implementation of the method according to such a first aspect, the sub-pixel interpolation filter is selected from a set of filters used for the intra prediction process for a given sub-pixel offset. In other words, the filter for the intra prediction process for a given sub-pixel offset is selected from a set of filters (for example, one of the filters, or one of the filter sets, can be used for the intra prediction process).
[0050] In a possible implementation of the method according to such a first aspect, the set of filters comprises a Gaussian filter and a cubic filter.
[0051] In a possible implementation of the method according to such a first aspect, the number of sub-pixel interpolation filters is N, and N sub-pixel interpolation filters are used for intra reference sample interpolation, where N >= 1 and is a positive integer.
[0052] In one possible implementation of the method according to such a first aspect, the reference sample in use for obtaining the value of the predicted sample of the block is not adjacent to the block of the predicted sample. The encoder may signal an offset value in the bitstream, and as a result, this offset value indicates the distance between the adjacent line of the reference sample and the line of the reference sample from which the value of the predicted sample is derived. FIG. 24 represents possible positions of the lines of the reference sample and the corresponding values of the ref_offset variable. The variable "ref_offset" indicates which reference line is used. For example, when ref_offset = 0, it represents that "reference line 0" (as shown in FIG. 24) is used.
[0053] Examples of the values of the offset in use in a particular implementation of a video codec (e.g., a video encoder or a video decoder) are as follows. Use the adjacent line of the reference sample (ref_offset = 0 indicated by "reference line 0" in FIG. 24), Use the first line (ref_offset = 1 indicated by "reference line 1" in FIG. 24) that is closest to the adjacent line, Use the third line (ref_offset = 3 indicated by "reference line 3" in FIG. 24).
[0054] The directional intra prediction mode specifies the value of the sub-pixel offset (deltaPos) between two adjacent lines of the predicted sample. This value is represented by a fixed-point integer value with 5-bit precision. For example, deltaPos = 32 means that the offset between two adjacent lines of the predicted sample is exactly 1 sample.
[0055] When the intra prediction mode is larger than DIA_IDX (mode #34), for the example described above, the value of the major reference side size is calculated as follows. Among the set of available (i.e., encoder-allowed for the prediction sample block) intra prediction modes that are larger than DIA_IDX and provide the maximum deltaPos value, the mode is considered. The desired sub-pixel offset value is derived as follows. That is, the block height is summed with ref_offset and multiplied by the deltaPos value. If the result is divisible by 32 with a remainder of 0, another maximum value of deltaPos as described above, provided that when obtaining the mode from the set of available intra prediction modes, the previously considered prediction mode is skipped. Otherwise, the result of this multiplication is considered as the maximum non-integer sub-pixel offset. The integer part of this offset is taken by shifting it 5 bits to the right. Sum the integer part of the maximum non-integer sub-pixel offset, the width of the prediction sample block, and half of the length of the interpolation filter.
[0056] Instead, when the intra prediction mode is smaller than DIA_IDX (mode #34), for the example described above, the value of the major reference side size is calculated as follows. Among the set of available (i.e., encodable by the encoder for the block of prediction samples) intra prediction modes that are smaller than DIA_IDX and provide the largest deltaPos value, the mode is considered. The value of the desired sub-pixel offset is derived as follows. That is, the block width is summed with ref_offset and multiplied by the deltaPos value. If the result of this is divisible by 32 with a remainder of 0, another maximum value of deltaPos as described above, however, when obtaining the mode from the set of available intra prediction modes, the previously considered prediction mode is skipped. Otherwise, the result of this multiplication is considered to be the maximum non-integer sub-pixel offset. The integer part of this offset is taken by shifting it right by only 5 bits. Sum the integer part of the maximum non-integer sub-pixel offset, the height of the block of prediction samples, and half the length of the interpolation filter.
[0057] According to a second aspect, the present invention relates to an intra prediction method for predicting a current block included in a picture. The method determines a size of a primary reference side used in intra prediction based on, among a plurality of available intra prediction modes, an intra prediction mode that results in a maximum non-integer value of a sub-pixel offset between a target sample among a plurality of target samples (such as a current sample among a plurality of current samples) in the current block and a reference sample used to predict the target sample in the current block (the reference sample is a reference sample among a plurality of reference samples included in a primary reference side), and a size of an interpolation filter to be applied to the reference samples included in the primary reference side. The method further includes applying an interpolation filter to the reference samples included in the primary reference side to obtain filtered reference samples, and predicting a plurality of samples (such as a plurality of current samples or a plurality of target samples) included in the current block based on the filtered reference samples.
[0058] Thus, the present disclosure facilitates achieving memory efficiency in video coding using intra prediction.
[0059] For example, the size of the primary reference side is determined as a sum of an integer part of the maximum non-integer value of the sub-pixel offset, a size of a side of the current block, and half of the size of the interpolation filter. In other words, the advantages of the second aspect may correspond to the above-described advantages of the first aspect.
[0060] In some embodiments, when the intra prediction mode is greater than the vertical intra prediction mode VER_IDX, a side of the current block is a width of the current block, or when the intra prediction mode is less than the horizontal intra prediction mode HOR_IDX, a side of the current block is a height of the current block.
[0061] For example, the value of a reference sample having a position within a major reference side that is greater than twice the size of the side of the current block is set to be equal to the value of a sample having a sample position that is twice the size of the current block.
[0062] For example, the size of the major reference side is · the size of the side of the current block, and · half of the length of the interpolation filter - 1, and · hereinafter, that is, - the size of the side of the block, and - the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter (therefore, the additional samples ref[refW+refIdx+x] (x = 1..(Max(1,nTbW / nTbH)*refIdx+1)) are derived as follows. That is, ref[refW+refIdx+x]=p[-1+refW][-1-refIdx]), or the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter + 1 (therefore, the additional samples ref[refW+refIdx+x] (x = 1..(Max(1,nTbW / nTbH)*refIdx+2)) are derived as follows. That is, ref[refW+refIdx+x]=p[-1+refW][-1-refIdx]) the maximum of and is determined as the sum of
[0063] According to a third aspect, the present invention relates to an encoder comprising a processing circuit configuration for performing a method according to the first or second aspect of the present invention or any possible embodiment of the first or second aspect.
[0064] According to a fourth aspect, the present invention relates to a decoder comprising a processing circuit configuration for performing a method according to the first or second aspect of the present invention or any possible embodiment of the first or second aspect.
[0065] According to a fifth aspect, the present invention relates to an apparatus for intra prediction of a current block included in a picture, the apparatus comprising an intra prediction unit configured to predict a target sample included in the current block based on filtered reference samples. The intra prediction unit determines, among a plurality of available intra prediction modes, an intra prediction mode that results in the maximum non-integer value of the sub-pixel offset between a target sample among a plurality of target samples in the current block and the reference sample used to predict the target sample in the current block (where the reference sample is a reference sample among a plurality of reference samples included in a primary reference side), and a determination unit configured to determine the size of the primary reference side used in the intra prediction based on the size of the interpolation filter to be applied to the reference sample included in the primary reference side, and a filtering unit configured to apply an interpolation filter to the reference sample included in the primary reference side to obtain the filtered reference samples.
[0066] Accordingly, the present disclosure facilitates bringing about memory efficiency in video coding using intra prediction.
[0067] In some embodiments, the determination unit determines the size of the primary reference side as the sum of the integer part of the maximum non-integer value of the sub-pixel offset, the size of the side of the current block, and half of the size of the interpolation filter.
[0068] For example, when the intra prediction mode is greater than the vertical intra prediction mode VER_IDX, the side of the current block is the width of the current block, or when the intra prediction mode is less than the horizontal intra prediction mode HOR_IDX, the side of the current block is the height of the current block.
[0069] For example, the value of a reference sample having a position in a major reference side greater than twice the size of the side of the current block is set to be equal to the value of a sample having a sample position that is twice the size of the current block.
[0070] In some embodiments, the determination unit · the size of the side of the block, and · the sum of the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter + 1, determines the size of the major reference side.
[0071] The determination unit may be configured not to perform right padding when the maximum of the two values M is equal to the size of the side of the block, or to perform right padding when the maximum of the two values M is equal to the integer part of the maximum value of the sub-pixel offset + half of the length of the interpolation filter, or the integer part of the maximum non-integer value of the sub-pixel offset + half of the length of the interpolation filter + 1.
[0072] Additionally or alternatively, in some embodiments, the determination unit is configured to perform padding by repeating the first sample and / or the last sample of the major reference side to the left side and / or the right side respectively. Specifically, as follows: that is, when the major reference side is denoted as ref and the size of the major reference side is denoted as refS, the padding is represented as ref[-1]=p[0] and / or ref[refS + 1]=p[refS], where ref[-1] represents the left value of the major reference side, p[0] represents the value of the first reference sample of the major reference side, ref[refS + 1] represents the right value of the major reference side, and p[refS] represents the value of the last reference sample of the major reference side.
[0073] The method according to the second aspect of the present invention can be executed by the apparatus according to the fifth aspect of the present invention. Further features and implementations of the apparatus according to the fifth aspect of the present invention correspond to the features and implementations of the method according to the second aspect of the present invention or any possible implementation of the second aspect.
[0074] According to a sixth aspect, there is provided an apparatus comprising a module / unit / component / circuit for performing at least a part of the steps of the method according to any preceding aspect or any preceding implementation of any preceding aspect.
[0075] The apparatus according to this aspect can be extended to an implementation corresponding to the implementation of the method according to any preceding aspect. Therefore, the implementation of the apparatus has the features of the corresponding implementation of the method according to any preceding aspect.
[0076] The advantages of the apparatus according to any preceding aspect are the same as the advantages for the corresponding implementation of the method according to any preceding aspect.
[0077] According to a seventh aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the first aspect or any possible implementation of the first aspect.
[0078] According to an eighth aspect, the present invention relates to a video encoder for encoding a plurality of pictures into a bitstream, comprising an apparatus for intra prediction of a current block according to any of the embodiments described above.
[0079] According to a ninth aspect, the present invention relates to a video decoder for decoding a plurality of pictures from a bitstream, comprising an apparatus for intra prediction of a current block according to any of the embodiments described above.
[0080] According to a tenth aspect, there is provided a computer-readable storage medium storing instructions which, when executed, cause one or more configured processors to code video data. The instructions cause the one or more processors to execute a method according to the first aspect or any possible embodiment of the first aspect.
[0081] According to an eleventh aspect, the present invention relates to a computer program comprising program code for performing a method according to the first aspect or any possible embodiment of the first aspect when executed on a computer.
[0082] In another aspect of the present application, a decoder comprising a processing circuit configuration configured to perform the above method is disclosed.
[0083] In another aspect of the present application, a computer program product comprising program code for performing the above method is disclosed.
[0084] In another aspect of the present application, a decoder for decoding video data is disclosed, the decoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the decoder to perform the above method when executed by the processors.
[0085] The processing circuit configuration may be implemented in hardware, or in a combination of hardware and software, such as by a software programmable processor.
[0086] The aspects, embodiments, and implementations described herein may provide the advantageous effects described above with reference to the first aspect and the second aspect.
[0087] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims.
[0088] The following embodiments of the present invention will be described in more detail with reference to the accompanying figures and drawings.
Brief Description of the Drawings
[0089]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15A
Figure 15B
Figure 15C
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Mode for Carrying Out the Invention
[0090] Hereinafter, the same reference numerals refer to the same or at least functionally equivalent features unless otherwise explicitly specified.
[0091] In the following description, reference is made to the accompanying drawings, which form a part of the present disclosure and illustrate specific aspects of embodiments of the present invention, or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the figures. Therefore, the following mode for carrying out the invention should not be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0092] It will be understood that, for example, the disclosure regarding a method of description may also apply to a corresponding device or system configured to execute the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., one unit that executes one or more steps, or multiple units each of which executes one or more of a plurality of steps), e.g., functional units, for executing the one or more method steps described, even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific device is described based on one or more units, e.g., functional units, the corresponding method may include one step (e.g., one step that executes the functionality of one or more units, or multiple steps each of which executes the functionality of one or more of a plurality of units) for executing the functionality of the one or more units, even if such one or more steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other unless otherwise specifically stated.
[0093] Video coding generally refers to the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture", the terms "frame" or "image" may be used as synonyms in the field of video coding. Video coding (or generally coding) consists of two parts, namely, video encoding and video decoding. Video encoding is performed on the source side and generally involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and generally involves performing the inverse process compared to the encoder to reconstruct the video picture. Embodiments that refer to the "coding" of a video picture (or generally a picture) are understood to relate to the "encoding" or "decoding" of the video picture or respective video sequence. The combination of the encoding part and the decoding part is also called a codec (coding and decoding).
[0094] In the case of reversible video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss during storage or transmission). In the case of irreversible video coding, for example, further compression by quantization is performed to reduce the amount of data representing the video picture, and the video picture may not be fully reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0095] Some video coding standards belong to the group of "irreversible hybrid video coders" (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video is typically processed or encoded at the block (video block) level, for example, by using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to generate a prediction block, subtracting the prediction block from the current block (the block being currently processed / to be processed) to obtain a residual block, transforming the residual block and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression), and in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed block to reconstruct the current block for rendering. Further, the encoder duplicates the decoder processing loop such that both have the same prediction (e.g., intra and inter prediction), and / or generate a reconstruction for processing or coding subsequent blocks.
[0096] In the following embodiments of the video coding system 10, the video encoder 20 and the video decoder 30 are described with reference to FIGS. 1 - 3.
[0097] FIG. 1A is a schematic block diagram showing an exemplary coding system 10, for example, a video coding system 10 (or short coding system 10) that can utilize the techniques of this application. The video encoder 20 (or short encoder 20) and the video decoder 30 (or short decoder 30) of the video coding system 10 represent examples of devices that can be configured to perform the techniques according to various examples described in this application.
[0098] As shown in FIG. 1A, the coding system 10 comprises a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, for decoding the encoded picture data 13.
[0099] The source device 12 comprises an encoder 20 and may additionally, i.e., optionally, comprise a picture source 16, a pre-processor (or pre-processing unit) 18, for example, a picture pre-processor 18, and a communication interface or communication unit 22.
[0100] The picture source 16 may comprise any kind of picture capture device, for example, a camera for capturing real-world pictures, and / or any kind of picture generation device, for example, a computer graphics processor for generating computer-animated pictures, or any kind of other device for acquiring and / or providing real-world pictures, computer-generated pictures (for example, screen content, virtual reality (VR) pictures), and / or any combination thereof (for example, augmented reality (AR) pictures), or may be those. The picture source may be any kind of memory or storage for storing any of the above-described pictures.
[0101] Distinguished from the pre-processor 18 and the processing performed by the pre-processing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.
[0102] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture 19 or preprocessed picture data 19. The preprocessing executed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.
[0103] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21 (further details will be described below, for example, based on FIG. 2).
[0104] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or a further processed version thereof) via the communication channel 13 to another device, for example, the destination device 14 or any other device, for storage or direct reconstruction.
[0105] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and additionally, i.e., optionally, may include a communication interface or communication unit 28, a postprocessor 32 (or postprocessing unit 32), and a display device 34.
[0106] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or a further processed version thereof), for example, directly from the source device 12 or from any other source, such as a storage device, for example, an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.
[0107] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired network or a wireless network or any combination thereof, or via any type of private network and public network or any combination thereof.
[0108] Communication interface 22 may be configured to, for example, package encoded picture data 21 into an appropriate format, such as a packet, and / or process the encoded picture data using any type of transmission encoding or transmission processing for transmission via a communication link or communication network.
[0109] Communication interface 28, which forms the counterpart of communication interface 22, may be configured to, for example, receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or transmission processing and / or unpacking to obtain the encoded picture data 21.
[0110] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by an arrow with respect to communication channel 13 in FIG. 1A pointing from source device 12 to destination device 14, for example, to set up a connection and recognize responses and exchange any other information related to the communication link and / or data transmission, such as encoded picture data transmission, for example, by sending and receiving messages.
[0111] Decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (further details will be described below, for example, based on FIG. 3 or FIG. 5).
[0112] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, post-processed picture 33. The post-processing executed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing, for example, by the display device 34, to prepare the decoded picture data 31 for display.
[0113] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33, for example, to display a picture to a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor, or may include them. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0114] FIG. 1A shows source device 12 and destination device 14 as separate devices, but embodiments of the devices may also include both source device 12 or corresponding functionality and destination device 14 or corresponding functionality, or both. In such embodiments, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software or any combination thereof.
[0115] As will be apparent to those skilled in the art based on the description, the functionality of the different units, i.e., the presence and (exact) partitioning of functionality within source device 12 and / or destination device 14 as shown in FIG. 1A, can vary depending on the actual device and application.
[0116] Encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) can each be implemented as any of a variety of suitable circuit configurations as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technique is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium for executing the techniques of the present disclosure, and may execute the instructions in hardware using one or more processors. Any of the above (including hardware, software, combinations of hardware and software, etc.) may be considered to be one or more processors. Each of video encoder 20 and video decoder 30 may be included within one or more encoders or decoders, and any of them may be integrated into their respective devices as part of a combined encoder / decoder (codec).
[0117] Source device 12 and destination device 14 may comprise any of a wide variety of devices, such as any kind of handheld or stationary device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiver device, a broadcast transmitter device, etc., and may not use an operating system, or may use any kind of operating system. In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.
[0118] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the techniques of the present application may be applied to video coding settings (such as video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, the data is retrieved from local memory, streamed over a network, etc. The video encoding device may encode and store the data in memory, and / or the video decoding device may retrieve and decode the data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data in memory and / or retrieve and decode data from memory.
[0119] FIG. 1B is an exemplary diagram of another exemplary video coding system 40 that includes encoder 20 of FIG. 2 and / or decoder 30 of FIG. 3 according to an exemplary embodiment. System 40 can implement techniques according to the various examples described in this application. In the illustrated implementation, video coding system 40 may include imaging device 41, video encoder 100, video decoder 30 (and / or a video encoder implemented via logic configuration 47 of processing unit 46), antenna 42, one or more processors 43, one or more memory stores 44, and / or display device 45.
[0120] As illustrated, imaging device 41, antenna 42, processing unit 46, logic configuration 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 may be capable of communicating with each other. As described, both video encoder 20 and video decoder 30 are illustrated, but video coding system 40 may, in various examples, include only video encoder 20 or only video decoder 30.
[0121] As shown, in some examples, the video coding system 40 may include an antenna 42. The antenna 42 may be configured to transmit or receive, for example, an encoded bitstream of video data. Further, in some examples, the video coding system 40 may include a display device 45. The display device 45 may be configured to present video data. As shown, in some examples, the logic circuit configuration 47 may be implemented via the processing unit 46. The processing unit 46 may include, for example, application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. The video coding system 40 may also include an optional processor 43, which may similarly include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. In some examples, the logic circuit configuration 47 may be implemented via hardware, video coding dedicated hardware, and the like, and the processor 43 may implement general-purpose software, an operating system, and the like. Additionally, the memory store 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, the memory store 44 may be implemented by cache memory. In some examples, the logic circuit configuration 47 may access the memory store 44 (e.g., for the implementation of an image buffer). In other examples, the logic circuit configuration 47 and / or the processing unit 46 may include a memory store (e.g., a cache, etc.) for the implementation of an image buffer, etc.
[0122] In some examples, the video encoder 20 implemented via a logic circuit configuration may include an image buffer (e.g., by either the processing unit 46 or the memory store 44) and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include a video encoder 20 implemented via a logic circuit configuration 47 to embody various modules as described with respect to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit configuration may be configured to perform various operations as described herein.
[0123] The video decoder 30 may be implemented in a similar manner as implemented via a logic circuit configuration 47 to embody various modules as described with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 30 may be implemented via a logic circuit configuration and may include an image buffer (e.g., by either the processing unit 420 or the memory store 44) and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include a video decoder 30 implemented via a logic circuit configuration 47 to embody various modules as described with respect to FIG. 3 and / or any other decoder system or subsystem described herein.
[0124] In some examples, the antenna 42 of the video coding system 40 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data related to coding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as described), and / or data defining coding partitions), data related to encoding video frames as described herein, indicators, index values, mode selection data, etc. The video coding system 40 may also include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. A display device 45 configured to present video frames.
[0125] For the sake of convenience of explanation, embodiments of the present invention are described herein, for example, by reference to the reference software of High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), that is, the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0126] Encoder and encoding method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of this application. In the example of FIG. 2, the video encoder 20 includes an input section 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output section 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in FIG. 2 may also be referred to as a hybrid video encoder, or a video encoder by a hybrid video codec.
[0127] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 are sometimes referred to as forming the forward signal path of the encoder 20, and the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are sometimes referred to as forming the reverse signal path of the video encoder 20, and the reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also sometimes referred to as forming the "built-in decoder" of the video encoder 20.
[0128] Picture and picture partitioning (picture and block) Encoder 20 may be configured to receive a picture 17 (or picture data 17), for example, a picture of a sequence of pictures forming a video or a video sequence, via, for example, input section 201. The received picture or picture data may also be preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be referred to as the current picture (especially in video coding, to distinguish the current picture from other pictures of the same video sequence, i.e., also the video sequence comprising the current picture, for example, previously encoded and / or decoded pictures).
[0129] (Digital) pictures are, or can be regarded as, two-dimensional arrays or matrices of samples having intensity values. Samples in the array are sometimes called pixels (short for picture elements) or pels. The number of samples in the horizontal and vertical directions (i.e., axes) of the array or picture defines the size and / or resolution of the picture. For color depiction, usually three color components are employed, i.e., the picture may represent or include three sample arrays. In the RGB format or color space, the picture comprises corresponding sample arrays of red, green, and blue. However, in video coding, each pixel usually comprises a luminance component denoted by Y (sometimes L is also used instead), and two chrominance components denoted by Cb and Cr, in a luminance and chrominance format or color space, e.g., represented by YCbCr. The luminance (or short for luma) component Y represents luminance or gray level intensity (such as in a grayscale picture), and the two chrominance (or short for chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format comprises a luminance sample array of luminance sample values (Y), and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted or transformed to the YCbCr format, and vice versa, and the process is also called color transformation or color conversion. If the picture is monochrome, the picture may only comprise a luminance sample array. Thus, the picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0130] An embodiment of the video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into a plurality of (usually non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size for all pictures of the video sequence and the corresponding grid that defines the block size, or to change the block size between pictures or subsets or groups of pictures and partition each picture into corresponding blocks.
[0131] In a further embodiment, the video encoder may be configured to directly receive the blocks 203 of picture 17, such as one, some, or all of the blocks that form picture 17. Picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.
[0132] Similar to picture 17, picture blocks 203 may again be or be regarded as a two-dimensional array or matrix of samples having intensity values (sample values), but of a smaller dimension than picture 17. In other words, block 203 may comprise, for example, one sample array (e.g., a luma array in the case of a monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color picture 17), or any other number and / or type of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (i.e., axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples or an M×N array of transform coefficients.
[0133] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode picture 17 in block units. For example, encoding and prediction are performed for each block 203.
[0134] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be provided later) by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 in sample units (pixel units), and obtain the residual block 205 in the sample area.
[0135] Transformation The transformation processing unit 206 may be configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207 may also be referred to as transformed residual coefficients and represent the residual block 205 in the transform domain.
[0136] The conversion processing unit 206 may be configured to apply integer approximations of DCT / DST such as the conversion specified for H.265 / HEVC. Compared with the orthogonal DCT transform, such integer approximations are typically scaled by some coefficients. To maintain the norm of the residual blocks processed by the forward and inverse transforms, additional scaling coefficients are applied as part of the conversion process. The scaling coefficients are typically selected based on several constraints such as the scaling coefficient being a power of 2 for shift operations, the bit depth of the conversion coefficients, and the trade-off between accuracy and implementation cost. A specific scaling coefficient may be specified for, for example, the inverse conversion by the inverse conversion processing unit 212 (and, for example, the corresponding inverse conversion by the inverse conversion processing unit 312 in the video decoder 30), and the corresponding scaling coefficient for the forward conversion by, for example, the conversion processing unit 206 in the encoder 20 may be specified accordingly.
[0137] Embodiments of the video encoder 20 (each, the conversion processing unit 206) may be configured to output conversion parameters, for example, one or more types of conversions, which are encoded or compressed directly or via the entropy encoding unit 270, for example, so that the video decoder 30 can receive and use the conversion parameters for decoding.
[0138] Quantization The quantization unit 208 may be configured to quantize the conversion coefficients 207 to obtain quantized coefficients 209, for example, by applying scalar quantization or vector quantization. The quantized coefficients 209 may also be referred to as quantized conversion coefficients 209 or quantized residual coefficients 209.
[0139] The quantization process may reduce the bit depth associated with some or all of the conversion coefficients 207. For example, an n-bit conversion coefficient may be truncated to an m-bit conversion coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, different scalings may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, and a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may be, for example, an index to a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to finer quantization (small quantization step size), a large quantization parameter may correspond to coarser quantization (large quantization step size), or vice versa. Quantization may include division by the quantization step size. For example, inverse quantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, such as embodiments according to HEVC, may be configured to use the quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation that includes division. Additional scaling factors may be introduced for quantization and inverse quantization to restore the norm of the residual block that may be modified by the scaling used in the fixed-point approximation of the equation for the quantization step size and quantization parameter. In an exemplary implementation, the scaling of inverse transform and inverse quantization may be combined. Alternatively, a customized quantization table may be used, for example, signaled from the encoder to the decoder in the bitstream. Quantization is an irreversible operation, and the loss increases with increasing quantization step size.
[0140] Embodiments of the video encoder 20 (each quantization unit 208) may be configured to output quantization parameters (QP), for example, encoded directly or via the entropy encoding unit 270, such that the video decoder 30 may receive and apply the quantization parameters for decoding.
[0141] Inverse quantization The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantization coefficients, for example, based on or using the same quantization step size as the quantization unit 208, or by applying the inverse of the quantization method applied by the quantization unit 208, to obtain inverse quantization coefficients 211. The inverse quantization coefficients 211 may also be referred to as inverse quantization residual coefficients 211 and - although usually not identical to the transform coefficients due to losses caused by quantization - correspond to the transform coefficients 207.
[0142] Inverse transformation The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding inverse quantization coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0143] Reconstruction The reconstruction unit 214 (for example, an adder or summer 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 - on a sample-by-sample basis - to obtain a reconstructed block 215 in the sample domain.
[0144] Filter processing The loop filter unit 220 (or short "loop filter" 220) is configured to filter the reconstruction block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured to, for example, smooth pixel transitions or improve video quality in another way. The loop filter unit 220 may comprise one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bidirectional filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. The loop filter unit 220 is shown in FIG. 2 as being an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstruction block 221. The decoded picture buffer 230 may store the reconstructed coding block after the loop filter unit 220 has performed a filtering operation on the reconstructed coding block.
[0145] Embodiments of the video encoder 20 (each, loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information), encoded, for example, directly or via the entropy coding unit 270, such that the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0146] Decoded picture buffer The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures, or generally reference picture data, for encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks, for example, previously reconstructed and filtered blocks 221 of the same current picture, or different pictures, for example, previously reconstructed pictures, for example, for inter prediction, to provide complete, previously reconstructed, i.e., decoded pictures (and corresponding reference blocks and samples), and / or partially reconstructed current pictures (and corresponding reference blocks and samples). The decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, or generally unfiltered reconstructed samples, or any other further processed version of a reconstructed block or sample, for example, if the reconstructed block 215 has not been filtered by the loop filter unit 220.
[0147] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a classification unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, for example, from the same (current) picture and / or one or more previously decoded pictures, for example, from the decoded picture buffer 230 or another buffer (for example, a line buffer, not shown), filtered and / or unfiltered reconstructed samples or blocks. The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, to obtain a prediction block 265 or a predictor 265.
[0148] The mode selection unit 260 may be configured to determine or select a classification for the current block prediction mode (including no classification) and a prediction mode (for example, an intra prediction mode or an inter prediction mode), and generate a corresponding prediction block 265, which is used for the calculation of the residual block 205 and for the reconstruction of the reconstruction block 215.
[0149] Embodiments of the mode selection unit 260 may be configured to select partitioning and prediction modes (e.g., supported by or available to the mode selection unit 260) that provide the best match or, in other words, the minimum residual (the minimum residual implies better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead implies better compression for transmission or storage), or both, or to consider or balance both. The mode selection unit 260 may be configured to determine the partitioning and prediction modes based on rate distortion optimization (RDO), i.e., to select the prediction mode that results in the minimum rate distortion. Terms such as "best," "minimum," "optimal," etc. in this context do not necessarily refer to the overall "best," "minimum," "optimal," etc., but may refer to the achievement of termination or selection criteria such as thresholds or other constraints that potentially lead to "quasi-optimal selections" but reduce complexity and processing time, either exceeding or falling below such values.
[0150] In other words, the partitioning unit 262 may be configured to repeatedly use, for example, quad-tree-partitioning (QT), binary partitioning (BT), or triple-tree-partitioning (TT), or any combination thereof, to partition block 203 into smaller block partitions or sub-blocks (which again form blocks), and, for example, to perform prediction for each of the block partitions or sub-blocks. The mode selection may comprise the selection of the tree structure of the block 203 to be partitioned, and the prediction mode is applied to each of the block partitions or sub-blocks.
[0151] The following describes in more detail the splitting (e.g., by the splitting unit 260) and prediction processing (by the inter-prediction unit 244 and the intra-prediction unit 254) performed by the exemplary video encoder 20.
[0152] Splitting The splitting unit 262 may split the current block 203 into smaller splits, e.g., smaller blocks of square or rectangular size (also sometimes called sub-blocks). These smaller blocks may be further split into even smaller splits. This is also called a tree split or a hierarchical tree split. For example, at the root tree level 0 (hierarchical level 0, depth 0), the root block may be recursively split, e.g., into two or more blocks at the next lower tree level, e.g., nodes at tree level 1 (hierarchical level 1, depth 1), and these blocks may again be split, e.g., into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), until a termination criterion is reached, e.g., the maximum tree depth or the minimum block size, at which point the splitting ends. Blocks that are no longer split are also called leaf blocks or leaf nodes of the tree. A tree that uses splitting into two splits is called a binary tree (BT), a tree that uses splitting into three splits is called a ternary tree (TT), and a tree that uses splitting into four splits is called a quad tree (QT).
[0153] As described above, the term "block" as used herein may be a portion of a picture, particularly a square or rectangular portion. For example, referring to HEVC and VVC, a block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB:coding block), a transform block (TB:transform block), or a prediction block (PB:prediction block), or may correspond thereto.
[0154] For example, a coding tree unit (CTU) may be a CTB of luma samples, two corresponding CTBs of chroma samples, or a CTB of samples of a picture coded using a monochrome picture or three separate color planes, of a picture having three sample arrays, and a syntax structure used to code the samples, or may comprise them. Correspondingly, a coding tree block (CTB) may be samples of an N×N block for some values of N such that the partitioning of components into CTBs is a division. A coding unit (CU) may be a coding block of luma samples, two corresponding coding blocks of chroma samples, or a coding block of samples of a picture coded using a monochrome picture or three separate color planes, of a picture having three sample arrays, and a syntax structure used to code the samples, or may comprise them. Correspondingly, a coding block (CB) may be samples of an M×N block for some values of M and N such that the partitioning of CTBs into coding blocks is a division.
[0155] For example, in an embodiment according to HEVC, a coding tree unit (CTU) can be divided into coding units (CUs) by using a quadtree structure shown as a coding tree. The decision as to whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU division type. Inside one PU, the same prediction process is applied, and the relevant information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU division type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0156] For example, in an embodiment according to the current latest video coding standard under development, called Versatile Video Coding (VVC), quadtree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node is further partitioned by a binary tree or a ternary (or tritree) structure. The partition tree leaf node is called a coding unit (CU), and its segmentation is used for prediction processing and transform processing without further partitioning. This means that the CU, PU, and TU have the same block size within the QTBT coding block structure. In parallel, it has been proposed that multiple partitions, for example, tritree partitioning, be used together with the QTBT block structure.
[0157] In one example, the mode selection unit 260 of the video encoder 20 can be configured to perform any combination of the partitioning techniques described herein.
[0158] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0159] Intra prediction The set of intra prediction modes may include 35 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes as defined, for example, in HEVC, or may include 67 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes as defined, for example, for VVC.
[0160] The intra prediction unit 254 is configured to generate an intra prediction block 265 according to an intra prediction mode of the set of intra prediction modes, using the reconstructed samples of adjacent blocks of the same current picture.
[0161] The intra prediction unit 254 (or generally, the mode selection unit 260) is further configured to output, in the form of a syntax element 266 for inclusion in the encoded picture data 21, an intra prediction parameter (or generally, information indicating the selected intra prediction mode for a block) to the entropy encoding unit 270, so that, for example, the video decoder 30 can receive and use the prediction parameters for decoding.
[0162] Inter prediction The set of inter prediction modes (or possible inter prediction modes) depends on the available reference pictures (i.e., for example, previously at least partially decoded pictures stored in DBP230) and other inter prediction parameters, for example, whether the entire reference picture is used to search for the best matching reference block, or only a portion of the reference picture, for example, the search window area around the area of the current block, and / or, for example, whether pixel interpolation, for example, half-pel / semi-pel interpolation and / or quarter-pel interpolation, is applied.
[0163] In addition to the above prediction modes, a skip mode and / or a direct mode may be applied.
[0164] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or acquire, for motion estimation, a picture block 203 (the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one or more blocks previously reconstructed, for example, reconstructed blocks of one or more other / different pictures 231 previously decoded. For example, the video sequence may comprise the current picture and a previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of a sequence of pictures forming the video sequence, or may form them.
[0165] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures among a plurality of other pictures, and provide a reference picture (or a reference picture index), and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter-prediction parameter. This offset is also called a motion vector (MV).
[0166] The motion compensation unit is configured to obtain, for example, receive, an inter-prediction parameter, and perform inter-prediction based on or using the inter-prediction parameter to obtain an inter-prediction block 265. The motion compensation performed by the motion compensation unit may involve performing interpolation to sub-pixel accuracy as much as possible to fetch or generate a prediction block based on the motion / block vector determined by motion estimation. The interpolation filtering process may generate additional pixel samples from known pixel samples, and thus potentially increase the number of candidate prediction blocks that can be used to code a picture block. When receiving a motion vector for the PU of the current picture block, the motion compensation unit may identify the position of the prediction block pointed to by the motion vector within one of the reference picture lists.
[0167] The motion compensation unit may also generate syntax elements related to the block and the video slice for use by the video decoder 30 when decoding the picture blocks of the video slice.
[0168] Entropy Coding The entropy encoding unit 270 is configured to apply or bypass (no compression) an entropy encoding algorithm or an entropy encoding method (e.g., variable length coding (VLC) method, context adaptive VLC method (CAVLC), arithmetic coding method, binarization, context adaptive binary arithmetic coding (CABAC), syntax based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or entropy encoding technique) to the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements to obtain the encoded picture data 21, and the encoded picture data 21 can be output via the output unit 272, for example, in the form of an encoded bitstream 21 so that, for example, the video decoder 30 can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30 or stored in memory so that it can be transmitted or retrieved later by the video decoder 30.
[0169] Other structural variants of the video encoder 20 may be used to encode the video stream. For example, the non-transform based encoder 20 can directly quantize the residual signal without using the transform processing unit 206 for some blocks or frames. In another implementation, the encoder 20 can combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.
[0170] Decoder and Decoding Method Figure 3 shows an example of a video decoder 30 configured to implement the techniques of this application. The video decoder 30 is configured to receive, for example, encoded picture data 21 (e.g., an encoded bitstream 21) encoded by an encoder 20 and obtain a decoded picture 331. The encoded picture data or bitstream comprises information for decoding the encoded picture data, e.g., data representing picture blocks of an encoded video slice, and associated syntax elements.
[0171] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may execute a decoding path generally opposite to the encoding path described with respect to the video encoder 100 of FIG. 2.
[0172] As described with respect to encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 are also referred to as forming the "built-in decoder" of video encoder 20. Thus, the inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transform processing unit 312 may have the same function as the inverse transform processing unit 212, the reconstruction unit 314 may have the same function as the reconstruction unit 214, the loop filter 320 may have the same function as the loop filter 220, and the decoded picture buffer 330 may have the same function as the decoded picture buffer 230. Thus, the descriptions provided for each unit and function of video 20 encoder apply correspondingly to each unit and function of video decoder 30.
[0173] Entropy decoding The entropy decoding unit 304 parses the bitstream 21 (or generally, the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or decoding method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode selection unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or video block level.
[0174] Inverse quantization The inverse quantization unit 310 receives quantization parameters (QP) (or generally, information related to inverse quantization) and quantization coefficients from the encoded picture data 21 (e.g., by the entropy decoding unit 304, e.g., by parsing and / or decoding), and based on the quantization parameters, applies inverse quantization to the decoded quantization coefficients 309 to obtain inverse quantization coefficients 311, which may also be called transform coefficients 311. The inverse quantization process may include the use of quantization parameters determined by the video encoder 20 for each video block in a video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0175] Inverse transformation The inverse transformation processing unit 312 may be configured to receive the inverse quantization coefficients 311, also called transformation coefficients 311, in order to obtain the reconstructed residual block 213 in the sample region and apply a transformation to the inverse quantization coefficients 311. The reconstructed residual block 213 may also be called the transformation block 313. The transformation may be an inverse transformation, such as an inverse DCT transformation, an inverse DST transformation, an inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive transformation parameters or corresponding information from the coded picture data 21 (e.g., by the entropy decoding unit 304, e.g., by syntax analysis and / or decoding) to determine the transformation to be applied to the inverse quantization coefficients 311.
[0176] Reconstruction The reconstruction unit 314 (e.g., an adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365, to obtain the reconstructed block 315 in the sample region.
[0177] Filter processing The loop filter unit 320 (either within or after the coding loop) is configured to filter the reconstructed block 315 to obtain a filtered block 321, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 320 may comprise one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, for example, a bidirectional filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. The loop filter unit 320 is shown in FIG. 3 as a in-loop filter, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0178] Decoded picture buffer The decoded video block 321 of the picture is then stored in the decoded picture buffer 330, which stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for outputting a display, respectively.
[0179] The decoder 30 is configured to output the decoded picture 311, for example, via the output unit 312, for presentation or viewing by the user.
[0180] Prediction The inter prediction unit 344 may be the same as the inter prediction unit 244 (specifically, the motion compensation unit), and the intra prediction unit 354 may have the same function as the inter prediction unit 254. Based on the respective information received from the coded picture data 21 (for example, by the entropy decoding unit 304, such as by syntax analysis and / or decoding) of the division and / or prediction parameters, it performs the determination and prediction of division or segmentation. The mode selection unit 360 may be configured to perform prediction (intra prediction or inter prediction) for each block based on the reconstructed picture, block, or each sample (filtered or unfiltered) to obtain the prediction block 365.
[0181] When the video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the mode selection unit 360 is configured to generate a prediction block 365 for the picture block of the current video slice based on the signaled intra prediction mode and the data from the previously decoded blocks of the current picture. When the video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (for example, the motion compensation unit) of the mode selection unit 360 is configured to create a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In the case of inter prediction, the prediction block may be created from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists, i.e., list 0 and list 1, using default configuration techniques based on the reference pictures stored in the DPB 330.
[0182] The mode selection unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors and other syntax elements, and use the prediction information to create a prediction block for the current video block being decoded. For example, the mode selection unit 360 may use some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to code video blocks of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists for the slice, a motion vector for each inter-coded video block of the slice, an inter prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0183] Other variants of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may be able to create an output video stream without using the loop filter processing unit 320. For example, a non-transform-based decoder 30 may be able to directly inverse quantize the residual signal for some blocks or frames without using the inverse transform processing unit 312. In another implementation, the video decoder 30 may combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.
[0184] FIG. 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video coding device 400 may be a decoder such as the video decoder 30 of FIG. 1A or an encoder such as the video encoder 20 of FIG. 1A.
[0185] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. The video coding device 400 may also include optoelectrical (OE: optical-to-electrical) components and electro-optical (EO: electrical-to-optical) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for the outlet or inlet of optical or electrical signals.
[0186] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with the inlet port 410, the receiver unit 420, the transmitter unit 440, the outlet port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 performs, processes, prepares, or provides various coding operations. Thus, including the coding module 470 brings a significant improvement to the functionality of the video coding device 400 and affects the conversion of the video coding device 400 to different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0187] Memory 460 may include one or more disks, tape drives, and solid state drives, and can be used as an overflow data storage device for storing programs when such programs are selected for execution and for storing instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0188] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as one or both of the source device 12 and the destination device 14 from FIG. 1 according to an exemplary embodiment. Apparatus 500 can implement the techniques of this present application described above. Apparatus 500 can be in the form of a computing system that includes a plurality of computing devices, or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
[0189] The processor 502 in apparatus 500 can be a central processing unit. Alternatively, processor 502 can be any other type of device or devices capable of manipulating or processing information, existing or to be developed in the future. The disclosed implementation can be practiced using a single processor as illustrated, e.g., processor 502, but the advantages in terms of speed and efficiency can be achieved using two or more processors.
[0190] In one implementation, the memory 504 in the device 500 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 to N, and the applications 1 to N further include a video coding application that executes the methods described herein. The device 500 can also include additional memory in the form of secondary storage 514, which can be, for example, a memory card used with a mobile computing device. Since a video communication session can contain a significant amount of information, it can be stored in whole or in part in the secondary storage 514 and loaded into the memory 504 as needed for processing.
[0191] The device 500 can also include one or more output devices, such as a display 518. In one example, the display 518 can be a touch-sensitive display that combines the display with a touch-sensitive element operable to sense touch input. The display 518 can be coupled to the processor 502 via the bus 512. In addition to or instead of the display 518, other output devices can be provided that enable a user to program or otherwise use the device 500. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light-emitting diode (LED) display, such as an organic LED (OLED) display.
[0192] Device 500 may also include, or be communicable with, an image sensing device 520, such as a camera, or any other existing or future-developed image sensing device 520 capable of sensing an image, such as an image of a user operating the device 500. The image sensing device 520 may be arranged to be directed towards the user operating the device 500. In one example, the position and optical axis of the image sensing device 520 may be configured to be directly adjacent to the display 518 and to include the area from which the display 518 is visible within the field of view.
[0193] Device 500 may also include, or be communicable with, a sound sensing device 522, such as a microphone, or any other existing or future-developed sound sensing device capable of sensing sound near the device 500. The sound sensing device 522 may be arranged to be directed towards the user operating the device 500 and may be configured to receive sounds, such as voices or other utterances made by the user while the user operates the device 500.
[0194] FIG. 5 shows the processor 502 and the memory 504 of the device 500 as being integrated within a single unit, although other configurations may be utilized. The operations of the processor 502 may be distributed across a plurality of machines (each machine having one or more of the processors) that may be coupled directly or over a local area network or other network. The memory 504 may be distributed across a plurality of machines, such as network-based memory or memory within the plurality of machines that execute the operations of the device 500. Although shown here as a single bus, the bus 512 of the device 500 may consist of a plurality of buses. Further, the secondary storage 514 may be directly coupled to the other components of the device 500 or may be accessed via a network and may comprise a single integrated unit, such as a memory card, or a plurality of units, such as a plurality of memory cards. The device 500 may be implemented in such a wide variety of configurations.
[0195] Initialisms Definition and Glossary JEM Joint Exploration Model (Software Codebase for Future Video Coding Exploration) JVET Joint Video Exploration Team LUT Look-up Table QT Quad Tree QTBT Quad Tree + Binary Tree RDO Rate-Distortion Optimization ROM Read-Only Memory VTM VVC Test Model VVC Versatile Video Coding, i.e., the Standardization Project Developed by JVET CTU / CTB Coding Tree Unit / Coding Tree Block CU / CB Coding Unit / Coding Block PU / PB Prediction Unit / Prediction Block TU / TB Transform Unit / Transform Block HEVC High Efficiency Video Coding
[0196] Video coding methods such as H.264 / AVC and HEVC are designed in accordance with the principles of the favorable results of block-based hybrid video coding. Using this principle, a picture is first divided into blocks, and then each block is predicted by using intra-picture prediction or inter-picture prediction.
[0197] Several video coding standards since H.261 belong to the group of "irreversible hybrid video coders" (i.e., combining spatial and temporal prediction in the sample domain and 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video is typically processed or encoded at the block (picture block) level by using, for example, spatial (intra-picture) prediction and temporal (inter-picture) prediction to generate a predicted block, subtracting the predicted block from the current block (the block being currently processed / to be processed) to obtain a residual block, transforming the residual block and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression), and in the decoder, the reverse process compared to the encoder is partially applied to the encoded or compressed block to reconstruct the current block for rendering. Further, the encoder replicates the decoder processing loop such that both have the same prediction (e.g., intra prediction and inter prediction), and / or generate a reconstruction for processing or coding subsequent blocks.
[0198] As used herein, the term "block" may be part of a picture or frame. For the sake of convenience in description, embodiments of the present invention are described herein with respect to the reference software of High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), which was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC. It may refer to CU, PU, and TU. In HEVC, a Coding Tree Unit (CTU) is divided into Coding Units (CUs) by using a quadtree structure shown as a coding tree. The decision as to whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four Prediction Units (PUs) according to the PU partition type. Inside one PU, the same prediction process is applied, and the relevant information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU may be divided into Transform Units (TUs) according to another quadtree structure similar to the coding tree for the CU. In the latest development of video compression technology, quadtree and quadtree binary (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a Coding Tree Unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a quadtree binary structure. The quadtree binary leaf nodes are called Coding Units (CUs), and their segmentation is used for prediction processing and transform processing without further partitioning. This means that CUs, PUs, and TUs have the same block size within the QTBT coding block structure. In parallel, it has been proposed that composite partitioning, for example, ternary partitioning, be used together with the QTBT block structure.
[0199] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC29 / WG11) are considering the potential needs for the standardization of future video coding technologies that have a compression capability significantly exceeding that of the current HEVC standard (including its current extensions and near-future extensions for screen content coding and high dynamic range coding). The group is collaborating in this exploration activity within a joint research effort called the Joint Video Exploration Team (JVET) to evaluate the compression technology designs proposed by their experts in this area.
[0200] In the case of directional intra prediction, different prediction angles from diagonally up to diagonally down are represented, and the intra prediction mode is available. For the definition of the prediction angle, an offset value pang on a 32-sample grid is defined. The association of p ang to the corresponding intra prediction mode is visualized in Figure 6 for the vertical prediction mode. For the horizontal prediction mode, the scheme is flipped vertically, and accordingly p ang values are assigned. As described above, all angle prediction modes are available for all applicable intra prediction block sizes. They all use the same 32-sample grid for the definition of the prediction angle. The distribution of p ang values across the 32-sample grid in Figure 6 reveals an increased resolution of the prediction angle around the vertical direction and a coarser resolution of the prediction angle towards the diagonal direction. The same applies in the horizontal direction. This design is derived from the observation that in many video contents, nearly horizontal and vertical structures play a more important role compared to diagonal structures.
[0201] For the horizontal prediction direction and the vertical prediction direction, the selection of the samples to be used for prediction is straightforward, but this task requires more effort in the case of angle prediction. For modes 11 to 25, when predicting the current block Bc from a set of prediction samples p ref (also called the primary reference side)ref Samples from both the vertical and horizontal portions can be involved. p ref Since determining the location of each sample at any of the branches of p requires some computational effort, an integrated one-dimensional prediction reference has been designed for HEVC intra prediction. The scheme is visualized in FIG. 7. Before performing the actual prediction operation, the reference sample p ref set is mapped to a one-dimensional vector p 1,ref . The projection used for mapping depends on the direction indicated by the intra prediction angle of each intra prediction mode. Only the reference samples from the portion of p ref that will be used for prediction are mapped to p 1,ref . The actual mapping of the reference samples to p 1,ref for each angular prediction mode is shown in FIGS. 8 and 9 for the horizontal and vertical angular prediction directions, respectively. The reference sample set p 1,ref is configured once for the block of prediction samples. The prediction is then derived from two adjacent reference samples in the set as detailed below. As can be seen from FIGS. 8 and 9, the one-dimensional reference sample set is not always completely filled for all intra prediction modes. Only the locations within the projection range for the corresponding intra prediction direction are included in the set.
[0202] Predictions for both the horizontal and vertical prediction modes are performed in the same way by simply swapping the x and y coordinates of the block. Prediction from p 1,ref is performed at 1 / 32 pel accuracy. Depending on the value of the angle parameter pang, the sample offset i 1,ref in p idx , and the weighting coefficient i fact for the sample at position (x,y) are determined. Here, the derivation for the vertical mode is provided. The derivation for the horizontal mode follows accordingly by swapping x and y.
[0203]
Number
[0204] if the fact is not equal to 0, i.e., the prediction does not cover the exact sample location in p 1,ref above the complete sample location in p, then p 1,ref the linear weighting between two adjacent sample locations in p is (0 ≦ x, y < Nc), where
[0205] [Number]
[0206] is executed as. i idx and i fact Note that the values of and i depend only on y, and thus (for the vertical prediction mode) need to be calculated only once per row.
[0207] VTM - 1.0 (Versatile Test Model) uses 35 intra - modes, while BMS (Benchmark Set) uses 67 intra - modes. Intra - prediction is a mechanism used in many video coding frameworks to improve compression efficiency when only a given frame can be involved.
[0208] Figure 10A shows an example of 67 intra - prediction modes, such as those proposed for VVC. A plurality of the 67 intra - prediction modes include a planar mode (index 0), a dc mode (index 1), and angular modes with indices 2 to 66. The lower - left angular mode in Figure 10A refers to index 2, and the indexing is incremented until index 66 becomes the top - rightmost angular mode in Figure 10A.
[0209] As shown in FIG. 10B, the latest version of VVC has several modes corresponding to diagonal intra prediction directions including a wide-angle mode (illustrated as a dashed line). In any of these modes, when the corresponding position within the block side is fractional, it should be executed to predict samples within the block interpolation of a set of adjacent reference samples. HEVC and VVC use linear interpolation between two adjacent reference samples. JEM uses a more sophisticated 4-tap interpolation filter. The filter coefficients are selected to be either a Gaussian filter or a cubic filter depending on the width or height value. The decision on whether to use the width or the height is coordinated with the decision on the major reference side selection, that is, when the intra prediction mode is above the diagonal mode, the side above the reference samples is selected as the major reference side, and the width value is selected to determine the interpolation filter during use. Otherwise, the major side reference is selected from the left side of the block, and the height controls the filter selection process. Specifically, when the selected side length is 8 samples or less, cubic interpolation 4-taps is applied. Otherwise, the interpolation filter is a 4-tap Gaussian filter.
[0210] The specific filter coefficients used in JEM are given in Table 1 (Table 5). The predicted sample is calculated by convolving as follows with the coefficients selected from Table 1 (Table 5) according to the sub-pixel offset and the filter type.
[0211]
Equation
[0212] In this equation, ">>" indicates a bitwise right shift operation.
[0213] The offset between the sample to be predicted (or simply "predicted sample") in the current block and the sample position to be interpolated may have an integer part and a non-integer part if the offset has a sub-pixel resolution such as 1 / 32 pixel. In Table 1 (Table 5), as well as in Table 2 (Table 6) and Table 3 (Table 7), the column "sub-pixel offset" refers to the non-integer part of the offset, e.g., fractional offset, fractional part of the offset, or fractional sample position.
[0214] If the cubic filter is selected, the predicted sample is further clipped to the allowable range of values, either as defined in the SPS or derived from the bit depth of the selected component.
[0215] [Table 5]
[0216] Another set of interpolation filters with 6-bit precision is presented in Table 2 (Table 6).
[0217] [Table 6]
[0218] The intra prediction sample is calculated by convolving the coefficients selected from Table 1 (Table 5) according to the sub-pixel offset and filter type as follows.
[0219] [Equation]
[0220] In this equation, ">>" indicates a bitwise right shift operation.
[0221] Another set of interpolation filters with 6-bit precision is presented in Table 3 (Table 7).
[0222]
Table 7
[0223] FIG. 11 shows a schematic diagram of a plurality of intra prediction modes used in the HEVC UIP mode. In the case of a luminance block, the intra prediction mode may comprise up to 36 intra prediction modes, which may include three non-directional modes and 33 directional modes. The non-directional modes may comprise a planar prediction mode, an average (DC) prediction mode, and a chroma from luma (LM) prediction mode. The planar prediction mode may perform prediction by assuming a block bevel surface having a horizontal gradient and a vertical gradient derived from the block boundary. The DC prediction mode may perform prediction by assuming a flat block surface having a value that matches the average value of the block boundary. The LM prediction mode may perform prediction by assuming that the chroma value for the block matches the luma value for the block. The directional modes may perform prediction based on adjacent blocks as shown in FIG. 11.
[0224] H.264 / AVC and HEVC specify that a low-pass filter can be applied to reference samples before being used in the intra prediction process. The decision on whether to use the reference sample filter is determined by the intra prediction mode and the block size. This mechanism is sometimes called Mode Dependent Intra Smoothing (MDIS). There are also multiple methods related to MDIS. For example, the Adaptive Reference Sample Smoothing (ARSS) method can signal whether the prediction samples are filtered, either explicitly (i.e., a flag is included in the bitstream) or implicitly (i.e., data hiding is used to avoid placing a flag in the bitstream and reduce signaling overhead). In this case, the encoder may make a decision on smoothing by testing the rate-distortion (RD) cost for all possible intra prediction modes.
[0225] As shown in FIG. 10B, the latest version of VVC has several modes corresponding to diagonal intra prediction directions. In any of these modes, when the corresponding position within the block side is fractional, it should be executed to predict samples within the block interpolation of the set of adjacent reference samples. HEVC and VVC use linear interpolation between two adjacent reference samples. JEM uses a more sophisticated 4-tap interpolation filter. The filter coefficients are selected to be either a Gaussian filter or a cubic filter depending on the value of the width or height. The decision on whether to use the width or the height is coordinated with the decision on the major reference side selection, that is, when the intra prediction mode is above the diagonal mode, the upper side of the reference samples is selected to be the major reference side, and the width value is selected to determine the interpolation filter during use. Otherwise, the major side reference is selected from the left side of the block, and the height controls the filter selection process. Specifically, when the selected side length is 8 samples or less, cubic interpolation 4-taps are applied. Otherwise, the interpolation filter is a 4-tap Gaussian filter.
[0226] An example of interpolation filter selection for modes smaller and larger than the diagonal mode (shown as 45°) in the case of a 32×4 block is shown in FIG. 12.
[0227] In VVC, a partitioning mechanism based on both quad-trees and binary-trees, called QTBT, is used. As shown in FIG. 13, the QTBT partitioning can provide not only square but also rectangular blocks. Of course, some signaling overhead and increased computational complexity on the encoder side are the price of the QTBT partitioning compared to the conventional quad-tree based partitioning used in the HEVC / H.265 standard. Nevertheless, the QTBT-based partitioning gives better segmentation characteristics and thus demonstrates significantly higher coding efficiency than the conventional quad-tree.
[0228] However, in its current state, VVC applies the same filter to both side parts (left and upper side parts) of the reference samples. Regardless of whether the block is oriented vertically or horizontally, the reference sample filter is the same for both reference sample side parts.
[0229] In this specification, the terms "vertically oriented block" ("block with a vertical orientation") and "horizontally oriented block" ("block with a horizontal orientation") are applied to rectangular blocks generated by the QTBT framework. These terms have the same meaning as shown in FIG. 14.
[0230] In the case of a directional intra prediction mode with a positive sub-sample offset, it is necessary to determine the memory size used to store the values of the reference samples. However, this size depends not only on the dimensions of the block of prediction samples but also on the processing further applied to these samples. Specifically, in the case of a positive sub-sample offset, the interpolation filter processing will require a major reference side part of an increased size compared to the case when it is not applied. The interpolation filter processing is performed by convolving the reference samples with a filter core. Therefore, the increase is caused by the additional samples required by the convolution operation to calculate the convolution results for the leftmost and rightmost parts of the major reference side part.
[0231] By using the steps described below, it is possible to determine the size of the major reference side part and thus reduce the amount of internal memory required to store the samples of the major reference side part.
[0232] FIGS. 15A, 15B, 15C to 18 show some examples of intra prediction of a block from the reference samples of the major reference side part. For each row of samples of the block of prediction samples, a (possibly fractional) sub-pixel offset is determined. This offset is orthogonal to the selected directional intra prediction mode M and the intra prediction mode M o(Depending on which of them is closer to the selected intra prediction mode, it may have an integer value or a non-integer value according to the difference between) and HOR_IDX or VER_IDX.
[0233] Table 4 (Table 8) and Table 5 (Table 9) represent the possible values of the sub-pixel offset for the first row of the prediction sample according to the mode difference. The sub-pixel offset for other rows of the prediction sample is obtained by multiplying the sub-pixel offset by the difference between the position of the row of the prediction sample and the first row.
[0234] [Table 8]
[0235] [Table 9]
[0236] When Table 4 (Table 8) or Table 5 (Table 9) is used to determine the sub-pixel offset for the bottom-right prediction sample, as shown in FIG. 15A, it can be noted that the major reference side size is equal to the sum of the integer part of the greatest or maximum sub-pixel offset, the size of the side of the block of the prediction sample (i.e., the block side length), and half of the length of the interpolation filter (i.e., half of the interpolation filter length).
[0237] The following steps may be performed to obtain the size of the major reference side for the selected directional intra prediction mode that provides a positive value of the sub-pixel offset.
[0238] 1. Step 1 may consist of determining which side of the block should be taken as the major side based on the index of the selected intra prediction mode, and which adjacent samples should be used to generate the major reference side. The major reference side is the line of reference samples used in predicting the samples in the current block. The "major side" is the side of the block parallel to the major reference side. When the (intra prediction) mode is above the diagonal mode (e.g., mode 34 as shown in FIG. 10A), the adjacent samples above (on top of) the block being predicted (i.e., the current block) are used to generate the major reference side, and the upper side is selected as the major side; otherwise, the adjacent samples to the left of the block being predicted are used to generate the major reference side, and the left side is selected as the major side. In summary, in Step 1, based on the intra prediction mode of the current block, the major side for the current block is determined. Based on the major side, the major reference side including the reference samples (some or all of them) used for predicting the current block is determined. For example, as shown in FIG. 15A, the major reference side is parallel to the major (block) side, but may be longer than the block side, for example. In other words, for example, when an intra prediction mode is given, for each sample in the current block, the corresponding reference sample among a plurality of reference samples (e.g., major reference side samples) is also given.
[0239] 2. Step 2 may consist of determining the maximum sub-pixel offset, which is calculated by multiplying the length of the non-primary side by the maximum value from either Table 4 (Table 8) or Table 5 (Table 9) such that the result of this multiplication represents a non-integer sub-pixel offset. Tables 4 (Table 8) and 5 (Table 9) provide exemplary values of sub-pixel offsets for samples in the first line of samples of the current block (either the top row of samples corresponding to the upper side selected as the primary side, or the leftmost column of samples corresponding to the left side selected as the primary side). Thus, the values shown in Tables 4 (Table 8) and 5 (Table 9) correspond to the sub-pixel offset per line of samples. Thus, the maximum offset occurring in the prediction of the overall block is obtained by multiplying this per-line value by the length of the non-primary side. In particular, in this example, since the fixed-point resolution is 1 / 32 samples, the result should not be a multiple of 32. If the multiplication of the length of the non-primary reference side by any per-line value from, for example, Table 4 (Table 8) or Table 5 (Table 9) results in a multiple of 32 corresponding to an integer sum value of sub-pixel offsets (i.e., an integer number of samples), this multiplication result is discarded. The non-primary side is the side of the block (either the upper side or the left side) not selected in Step 1. Thus, when the upper side is selected as the primary side, the length of the non-primary side is the width of the current block, and when the left side is selected as the primary side, the length of the non-primary side is the height of the current block.
[0240] 3. Step 3 consists of taking the integer part of the sub-pixel offset obtained in Step 2 corresponding to the multiplication result described above (i.e., by right-shifting only 5 times in binary representation), and adding it to the length of the main side (either the block width or the block length) and half of the length of the interpolation filter, which results in the total value of the main reference side. Thus, the main reference side comprises a line of samples that is parallel to and of equal length to the main reference side, extended by adjacent samples within the non-integer part of the sub-pixel offset and further adjacent samples within half of the length of the interpolation filter. Interpolation is performed over the samples within the length of the fractional part of the sub-pixel offset and the same amount of samples located beyond the length of the sub-pixel offset, so only half of the length of the interpolation filter is required.
[0241] According to another embodiment of the present disclosure, the reference samples used to obtain the value of the predicted pixel are not adjacent to the block of prediction samples. The encoder may signal the offset value within the bitstream, and as a result, this offset value indicates the distance between the adjacent line of reference samples and the line of reference samples from which the value of the prediction samples is derived.
[0242] FIG. 24 represents the possible positions of the lines of reference samples and the corresponding values of the ref_offset variable.
[0243] Examples of the values of the offset used in a particular implementation of a video codec (e.g., a video encoder or a video decoder) are as follows. - Use the adjacent line of reference samples (ref_offset = 0, indicated by "Reference Line 0" in FIG. 24), - Use the first line (ref_offset = 1, indicated by "Reference Line 1" in FIG. 24), which is the closest to the adjacent line, - Use the third line (ref_offset = 3, indicated by "Reference Line 3" in FIG. 24).
[0244] The variable "ref_offset" has the same meaning as the further used variable "refIdx". In other words, the variable "ref_offset" or the variable "refIdx" indicates a reference line. For example, when ref_offset = 0, it represents that "reference line 0" (as shown in FIG. 24) is used.
[0245] The directional intra prediction mode specifies the value (deltaPos) of the sub-pixel offset between two adjacent lines of the prediction samples. This value is represented by a fixed-point integer value with 5-bit precision. For example, deltaPos = 32 means that the offset between two adjacent lines of the prediction samples is exactly 1 sample.
[0246] When the intra prediction mode is greater than DIA_IDX (mode #34), for the example described above, the value of the major reference side size is calculated as follows. Among the set of intra prediction modes available (i.e., those that the encoder may indicate for the block of prediction samples), the mode that is greater than DIA_IDX and provides the maximum deltaPos value is considered. The value of the desired sub-pixel offset between the reference sample or the interpolated sample position and the sample to be predicted is derived as follows. That is, the block height is summed with ref_offset and multiplied by the deltaPos value. If the result is divisible by 32 with a remainder of 0, another maximum value of deltaPos as described above is used, provided that when obtaining the mode from the set of available intra prediction modes, the previously considered prediction mode is skipped. Otherwise, the result of this multiplication is considered as the maximum non-integer sub-pixel offset. The integer part of this offset is taken by shifting it 5 bits to the right.
[0247] The size of the major reference side is obtained by summing the integer part of the maximum non-integer sub-pixel offset, the width of the block of prediction samples, and half of the length of the interpolation filter (as shown in FIG. 15A).
[0248] Instead, when the intra prediction mode is smaller than DIA_IDX (mode #34), for the example described above, the value of the major reference side size is calculated as follows. Among the set of available intra prediction modes (i.e., those that the encoder may indicate for the block of prediction samples), the mode that is smaller than DIA_IDX and provides the largest deltaPos value is considered. The value of the desired sub-pixel offset is derived as follows. That is, the block width is summed with ref_offset and multiplied by the deltaPos value. If the result of this is divisible by 32 with a remainder of 0, another maximum value of deltaPos as described above, provided that when obtaining the mode from the set of available intra prediction modes, the previously considered prediction mode is skipped. Otherwise, the result of this multiplication is considered to be the maximum non-integer sub-pixel offset. The integer part of this offset is taken by shifting it 5 bits to the right. The size of the major reference side is obtained by summing the integer part of the maximum non-integer sub-pixel offset, the height of the block of prediction samples, and half the length of the interpolation filter.
[0249] FIGS. 15A, 15B, 15C to 18 show some examples of intra prediction of a block from the reference samples of the major reference side. For each row of samples of the block of prediction samples 1120, a fractional sub-pixel offset 1150 is determined. This offset may have an integer or non-integer value depending on the difference between the selected directional intra prediction mode M and the orthogonal intra prediction mode M o (either HOR_IDX or VER_IDX, depending on which of them is closer to the selected intra prediction mode).
[0250] State-of-the-art video coding methods and existing implementations of these methods exploit the fact that, in the case of intra angle prediction, the size of the primary reference side is determined as twice the length of the corresponding block side. For example, in HEVC, when the intra prediction mode is 34 or higher (see FIGS. 10A or 10B), the primary reference side samples are taken from the upper and upper-right adjacent blocks if these blocks are available, i.e., not from within a slice that has already been reconstructed and processed. The total number of adjacent samples used is set equal to twice the width of the block. Similarly, when the intra prediction mode is less than 34 (see FIG. 10), the primary reference side samples are taken from the left and lower-left adjacent blocks, and the total number of adjacent samples used is set equal to twice the height of the block.
[0251] However, when applying the sub-pixel interpolation filter, additional samples at the left and right ends of the primary reference side are used. To maintain compatibility with existing solutions, it is proposed that these additional samples be obtained by padding the primary reference side to the left and right. The padding is performed by repeating the first and last samples of the primary reference side to the left and right sides, respectively. Denoting the primary reference side as ref and its size as refS, the padding can be represented as the following assignment operation. ref[-1]=p[0] ref[refS+1]=p[refS]
[0252] In practice, the use of negative indices can be avoided by applying a positive integer offset when referring to elements of an array. Specifically, this offset can be set equal to the number of elements padded to the left of the primary reference side.
[0253] Specific examples of how to perform right and left padding are given in the following two cases illustrated by FIG. 15B.
[0254] The right padding case occurs when specifying a sub-pixel offset equal to, for example, 22 for the wide-angle modes 72 and -6 (Figure 10B), i.e., |M - M o | (see Table 4 (Table 8)).
[0255]
Number
[0256] When the aspect ratio of the block is 2 (i.e., when the dimensions of the predicted block are equal to 4×8, 8×16, 16×32, 32×64, 8×4, 16×8, 32×16, 64×32), the corresponding maximum sub-pixel offset value is calculated as
[0257]
Number
[0258] where S is the smaller side of the block.
[0259] Therefore, for an 8×4 block, the maximum sub-pixel offset is
[0260]
Number
[0261] i.e., the maximum value of the integer sub-pixel part of this offset is equal to 7. When applying a 4-tap intra interpolation filter to obtain the value of the bottom-right sample with coordinates x = 7, y = 3, the reference samples with indices x + 7 - 1, x + 7, x + 7 + 1, and x + 7 + 2 will be used. Since the primary reference side has 16 adjacent samples with indices 0..15, it means that one sample at the end of the primary reference side is padded by repeating the reference sample with position x + 7 + 1, so the rightmost sample position x + 7 + 2 = 16.
[0262] The same steps are performed when Table 5 is in use for modes 71 and -5. The subpixel offsets for this case are
[0263]
number
[0264] is equal to
[0265]
number
[0266] The maximum value is obtained.
[0267] For example, the left padding case occurs for angle modes 35..65 and 19..33 when the subpixel offset is fractional and less than one sample. For the top-left predicted sample, the corresponding subpixel offset value is calculated. According to Table 4 and Table 5, this offset corresponds to an integer subsample offset of 0.
[0268]
number
[0269] Applying a 4-tap interpolation filter to compute a predicted sample with coordinates x=0, y=0 would require reference samples with indices x-1, x, x+1, and x+2. The leftmost sample position x-1=-1. Since the main reference side has 16 adjacent samples with indices 0..15, the sample at this position is padded by repeating the reference sample with position x.
[0270] From the above example, for a block having an aspect ratio, the major reference side portion continues to be padded by half of the 4-tap filter length, i.e., 2 samples, one of which is added to the beginning (left end) of the major reference side portion and the other is added to the end (right end) of the major reference side portion. In the case of the 6-tap interpolation filter processing example, following the steps described above, 2 samples should be added to the beginning and the end of the major reference side portion. Generally, when an N-tap intra interpolation filter is used, the major reference side portion is
[0271]
Number
[0272] padded using
[0273]
Number
[0274] samples, of which
[0275]
Number
[0276] are padded to the left side portion and
[0277] When the steps described above are repeated for other block aspect ratios, the following offsets are obtained (see Table 6 (Table 10)).
[0278]
Table 10
[0279] From the values given in Table 6 (Table 10), for the wide-angle intra prediction mode, the following continues. When Table 4 (Table 8) is in use, in the case of 4-tap interpolation filter processing, left padding operation and right padding operation are required for block sizes 4×8, 8×4, 8×16, and 16×8. When Table 5 (Table 9) is in use, in the case of 4-tap interpolation filter processing, left padding operation and right padding operation are required only for block sizes 4×8 and 8×4.
[0280] Details of the proposed method are described in Table 7 (Table 11) in the specification format. The padding embodiments described above can be expressed as the following modifications to the VVC draft (Section 8.2.4.2.7).
[0281]
Table 11
[0282] Table 4 (Table 8) and Table 5 (Table 9) as described above represent the possible values of the sub-pixel offset between two adjacent lines of prediction samples according to the intra prediction mode.
[0283] State-of-the-art video coding solutions use different interpolation filters in intra prediction. Specifically, FIGS. 19 to 21 show various examples of interpolation filters.
[0284] In the present invention, as shown in FIG. 22 or FIG. 23, an intra prediction process of a block is executed. During the intra prediction process of the block, a sub-pixel interpolation filter is applied to luminance and chrominance reference samples. The sub-pixel interpolation filter (such as a 4-tap filter) is selected based on the sub-pixel offset between the position of the reference sample and the position of the sample to be interpolated, and the size of the main reference side used in the intra prediction process is determined according to the length of the sub-pixel interpolation filter and the intra prediction mode that brings the maximum value of the sub-pixel offset. The memory requirement is determined by the maximum value of the sub-pixel offset. The memory requirement is determined by the maximum value of the sub-pixel offset.
[0285] FIG. 15B shows a case where the top-left sample is not included in the main reference side, but instead is padded using the leftmost sample belonging to the main reference side. However, when the predicted sample is calculated by applying a 2-tap sub-pixel interpolation filter (for example, a linear interpolation filter), the top-left sample is not referenced, and thus padding is not required in this case.
[0286] FIG. 15C shows a case where a 4-tap sub-pixel interpolation filter (for example, a Gaussian filter, a DCT-IF filter, or a cubic filter) is used. In this case, it can be noted that four reference samples, namely the top-left sample (marked as "B") and the next three samples (each marked as "C", "D", and "E"), are required to calculate at least the top-left predicted sample (marked as "A").
[0287] In this case, two alternative methods are disclosed.
[0288] Padding the value of B using the value of C.
[0289] Use the reconstructed samples of the adjacent blocks in exactly the same way as other samples of the main reference side (including "B", "C", and "D") are obtained. In this case, the size of the main reference side is the block main side length (i.e., the block side length of the prediction sample or the size of the side of the block), and half of the interpolation filter length - 1, and the following two values M, namely the block main side length, the integer part of the maximum sub-pixel offset + half of the interpolation filter length, or the integer part of the maximum sub-pixel offset + half of the interpolation filter length + 1 (in view of memory issues, this addition of 1 to the total may or may not be included) of which the maximum and is determined as the sum of.
[0290] Note that "block main side", "block side length", "block main side length", and "size of the side of the block of the prediction sample" are the same concept throughout this disclosure.
[0291] It can be understood that half of the interpolation filter length - 1 is used to determine the size of the main reference side, and thus it is allowed to expand the main reference side at that time to the left.
[0292] It can be understood that the maximum of the two values M is used to determine the size of the main reference side, and thus it is allowed to expand the main reference side at that time to the right.
[0293] In the above description, the block main side length is determined according to the intra prediction mode (Figure 10B). When the intra prediction mode is greater than or equal to the diagonal intra prediction mode (#34), the block main side length is the width of the block of the prediction sample (i.e., the block to be predicted). Otherwise, the block main side length is the height of the block of the prediction sample.
[0294] Sub-pixel offset values can be defined for a wider range of angles (see Table 8 (Table 12)).
[0295]
Table 12
[0296] Depending on the aspect ratio, different maximum and minimum values of the intra prediction mode index (Figure 10B) are allowed. Table 9 (Table 13) gives an example of this mapping.
[0297]
Table 13
[0298] According to Table 9 (Table 13), for the maximum mode difference value max(|M - M o |), an integer sub-pixel offset is used for interpolation (the maximum sub-pixel offset per line is a multiple of 32), which means that the predicted samples of the prediction block are calculated by copying the values of the corresponding reference samples and no sub-sample interpolation filter is applied.
[0299] Considering the constraints for max(|M - M o |) in Table 9 (Table 13) and the values in Table 8 (Table 12), the maximum sub-pixel offset per line that does not require interpolation is defined as follows (see Table 10 (Table 14)).
[0300]
Table 14
[0301] Using Table 10 (Table 14), the value of the integer part of the maximum sub-pixel offset + half of the interpolation filter length for a square 4×4 block can be calculated using the following steps.
[0302] Step 1. The block main side length (equal to 4) is multiplied by 29 and the result is divided by 32, thus giving a value of 3.
[0303] Step 2. Half of the 4-tap interpolation filter length is 2, which is added to the value obtained in Step 1, giving a value of 5.
[0304] From the above example, it can be observed that the obtained value is larger than the block main side length. In this example, the size of the main reference side is set to 10, which is the block main side length (equal to 4), and half of the interpolation filter length - 1 (equal to 1), and the following two values M, namely, the block main side length (equal to 4), the integer part of the maximum sub-pixel offset + half of the interpolation filter length (equal to 5), or the integer part of the maximum sub-pixel offset + half of the interpolation filter length + 1 (equal to 6) (in view of memory issues, the addition of 1 to this sum may or may not be included) of which the maximum and is determined as the sum.
[0305] The total number of reference samples included in the main reference side is larger than twice the block main side length.
[0306] When the maximum of the two values M is equal to the block main side length, right padding is not performed. Otherwise, right padding is applied to reference samples having a position horizontally or vertically more than 2*nTbS (where nTbS indicates the block main side length) away from the position of the top-left prediction sample (denoted as "A" in Fig. 15C). Right padding is performed by assigning the value of the padded sample to the last reference sample value on the main block side having a position within the range of 2*nTbS.
[0307] When half of the interpolation filter length minus 1 is greater than 0, the value of sample "B" (shown in Figure 15C) is obtained by left padding, or the corresponding reference sample can be obtained in exactly the same way as reference samples "C", "D", and "E" are obtained.
[0308] Details of the proposed method are described in Table 11 (Table 15) in the specification format. Instead of right or left padding, the corresponding reconstructed adjacent reference samples can be used. The case when left padding is not used can be represented by the following part of the VVC specification (Section 8.2).
[0309]
Table 15
[0310] Similarly, using Table 10 (Table 14), the value of the integer part of the maximum sub-pixel offset plus half of the interpolation filter length for a non-square block with a width of 4 samples and a height of 2 samples can be calculated using the following steps (when the major side length of the block is the width).
[0311] Step 1. The block height (equal to 2) is multiplied by 57, and the result is divided by 32, thus giving a value of 3.
[0312] Step 2. Half of the 4-tap interpolation filter length is 2, which is added to the value obtained in Step 1, giving a value of 5.
[0313] The remaining steps for calculating the total number of reference samples included in the major reference side are the same as in the case of a square block.
[0314] Using the block dimensions from Table 10 (Table 14) and Table 6 (Table 10), it can be noted that the maximum number of reference samples receiving left or right padding is 2.
[0315] When the block to be predicted is not adjacent to the adjacent reconstructed reference samples used in the intra prediction process (the reference lines can be selected as shown in FIG. 24), the embodiments described below are applicable.
[0316] The first step is to define the aspect ratio of the block according to the major side of the prediction block in the intra prediction mode. When the upper side of the block is selected as the major side, the aspect ratio R a (denoted as "whRatio" in the VVC specification) is set equal to the result of the integer division of the width of the block (denoted as "nTbW" in the VVC specification) by the height of the block (denoted as "nTbH" in the VVC specification). Otherwise, in the case where the major side is the left side of the prediction block, the aspect ratio R a (denoted as "hwRatio" in the VVC specification) is set equal to the result of the integer division of the height of the block by the width of the block. In either case, if the value of Ra is less than 1 (i.e., the numerator value of the integer division operator is less than the denominator value), it is set equal to 1.
[0317] The second step is to add a portion of the reference sample (denoted as "p" in the VVC specification) to the primary reference side. Depending on the value of refIdx, either an adjacent reference sample or a non - adjacent reference sample is used. The reference sample added to the primary reference side is selected using an offset with respect to the primary block side in the direction of the primary side orientation. Specifically, when the primary side is the upper side of the prediction block, the offset is horizontal and is defined as the -refIdx sample. When the primary side is the left side of the prediction block, the offset is vertical and is defined as the -refIdx sample. In this step, starting from the top - left reference sample (denoted as the "B" sample in FIG. 15C) + the value of the offset explained above, nTbS + 1 samples are added (nTbS indicates the primary side length). Note that the explanation or definition of RefIdx is presented in this disclosure in combination with FIG. 24.
[0318] The next step to be executed depends on whether the sub - pixel offset (denoted as "intraPredAngle" in the VVC specification) is positive or negative. A value of 0 for the sub - pixel offset corresponds to the horizontal intra - prediction mode (in the case where the primary side of the block is the left block side) or the vertical intra - prediction mode (in the case where the primary side of the block is the upper block side).
[0319] When the sub-pixel offset is negative (e.g., in step 3, negative sub-pixel offset), in the third step, the primary reference side is extended to the left using the reference samples corresponding to the non-primary side. The non-primary side is the side not selected as the primary side. That is, when the intra prediction mode is 34 or more (Figure 10B), the non-primary side is the left side of the block to be predicted; otherwise, the non-primary side is the left side of the block. The extension is performed as shown in Figure 7, and the description of this process can be found in the related description of Figure 7. The reference samples corresponding to the non-primary side are selected according to the process disclosed in the second step, with the difference that the non-primary side rather than the primary side is used. When this step is completed, the primary reference sides are each extended from the beginning to the end using their first and last samples. In other words, in step 3, negative sub-pixel offset padding is performed.
[0320] When the sub-pixel offset is positive (e.g., in step 3, positive sub-pixel offset), in the third step, the primary reference side is extended to the right only by the additional nTbS samples in the same way as described in step 2. When the value of refIdx is greater than 0 (the reference sample is not adjacent to the block to be predicted), right padding is performed. The number of samples to be right-padded is equal to the aspect ratio Ra calculated in the first step multiplied by the refIdx value. When a 4-tap filter is in use, the number of samples to be right-padded increases by 1.
[0321] Details of the proposed method are described in Table 12 (Table 16) in the specification format. The VVC specification modifications for this embodiment can be as follows (refW is set to nTbS - 1).
[0322] [Table 16]
[0323] The portion described above in the VVC specification is also applicable to the case where the primary reference side part is left-padded by only 1 sample in the third step for positive values of the sub-pixel offset. The details of the proposed method are described in Table 13 (Table 17) in the specification format.
[0324] [Table 17]
[0325] The present disclosure provides an intra prediction method for predicting a current block included in a picture such as a video frame. The method steps of the intra prediction method are shown in FIG. 25. The current block is the above-described block including samples to be predicted (or "predicted sample" or "prediction sample"), for example, luminance samples or chrominance samples.
[0326] The method includes a step (S2510) of determining the size of the primary reference side part based on the intra prediction mode that results in the maximum non-integer value of the sub-pixel offset and the size (i.e., length) of the interpolation filter among a plurality of available intra prediction modes (for example, shown in FIGS. 10 to 11).
[0327] The sub-pixel offset is the offset between a sample (or "target sample") in the current block to be predicted and a reference sample (or reference sample position) on which the sample in the current block is predicted. If the reference sample includes samples that are not directly or linearly above (e.g., a mode having a number greater than the diagonal mode) or to the left (e.g., a mode having a number less than the diagonal mode) of the current block, but includes samples that are offset or shifted with respect to the position of the current block, the offset may be related to an angular prediction mode. Since not all modes indicate an integer reference sample position, the offset has a sub-pixel resolution, and this sub-pixel offset may take a non-integer value and may have an integer part + a non-integer part. In the case of a non-integer value sub-pixel offset, interpolation between the reference samples is performed. Thus, the offset is the offset between the position of the sample to be predicted and the position of the reference sample to be interpolated. The maximum non-integer value may be the maximum non-integer value (integer part + non-integer part) for any sample in the current block. For example, as shown in FIGS. 15A to 15C, the target sample related to the maximum non-integer sub-pixel offset may be the lower right sample in the current block. Note that intra prediction modes that result in an offset with an integer value greater than the maximum non-integer value of the sub-pixel offset are ignored.
[0328] Possible sizes (i.e., lengths) of the interpolation filter include 4 (e.g., the filter is a 4-tap filter) or 6 (e.g., the filter is a 6-tap filter).
[0329] The method further includes applying an interpolation filter to a reference sample included in the primary reference side (S2520) and predicting a target sample included in the current block based on the filtered reference sample (S2530).
[0330] In correspondence with the method shown in FIG. 26, an apparatus 2600 for intra prediction of a current block included in a picture is also provided. The apparatus 2600 is shown in FIG. 26 and may be included in the video encoder shown in FIG. 2 or the video decoder shown in FIG. 3. In one example, the apparatus 2600 may correspond to the intra prediction unit 254 in FIG. 2. In another example, the apparatus 2600 may correspond to the intra prediction unit 354 in FIG. 3.
[0331] The apparatus 2600 includes an intra prediction unit 2610 configured to predict a target sample included in a current block based on filtered reference samples. The intra prediction unit 2610 may be the intra prediction unit 254 shown in FIG. 2 or the intra prediction unit 354 shown in FIG. 3.
[0332] The intra prediction unit 2610 includes a determination unit 2620 (or "primary reference size determination unit") configured to determine a size of a primary reference side used in intra prediction. Specifically, based on an intra prediction mode (of a plurality of available intra prediction modes) that results in a maximum non-integer value of a sub-pixel offset between a target sample (among a plurality of target samples in the current block) and a reference sample (hereinafter referred to as "target reference sample") used to predict the target sample in the current block, and based on a size of an interpolation filter to be applied to the reference samples included in the primary reference side, the size is determined. The target sample is any sample of a block to be predicted. The target reference sample is one of the reference samples of the primary reference side.
[0333] The intra prediction unit 2610 further includes a filtering unit configured to apply an interpolation filter to the reference samples included in the primary reference side to obtain filtered reference samples.
[0334] In summary, the memory requirement is determined by the maximum value of the sub-pixel offset. Therefore, by determining the size of the primary reference side according to the present disclosure, the present disclosure facilitates bringing about memory efficiency in video coding using intra prediction. Specifically, the memory (buffer) used by the encoder and / or decoder to perform intra prediction can be allocated in an efficient manner according to the determined size of the primary reference side. This is because, firstly, the size of the primary reference side determined according to the present disclosure includes all the reference samples to be used for predicting the current block. Therefore, no access to additional samples is required to perform intra prediction. Secondly, this is not necessary for all the already processed samples of adjacent blocks. Rather, the memory size may be specifically allocated to the primary reference side, i.e., to those reference samples belonging to the determined size.
[0335] The following is an explanation of an encoding method, a decoding method as shown in the above-described embodiments, and application examples of systems using them.
[0336] FIG. 27 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types.
[0337] The capture device 3102 generates data and can encode the data by an encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0338] In the content supply system 3100, the terminal device 310 receives and plays back encoded data. The terminal device 3106 can be a device having data reception and restoration capabilities, such as a smartphone or a pad 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0339] In the case of a terminal device having its display, such as a smartphone or a pad 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. In the case of a terminal device not equipped with a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is contacted to them to receive and display the decoded data.
[0340] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device as shown in the above-described embodiments may be used.
[0341] FIG. 28 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progress unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination of these types.
[0342] After the protocol progress unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0343] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. The video decoder 3206 including the video decoder 30 as described in the above embodiments decodes the video ES by the decoding method shown in the above embodiments to generate video frames, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames can be stored in a buffer (not shown in FIG. Y) before supplying them to the synchronization unit 3212. Similarly, the audio frames can be stored in a buffer (not shown in FIG. Y) before supplying them to the synchronization unit 3212.
[0344] The synchronization unit 3212 synchronizes the video frame and the audio frame and supplies the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of the video and the audio information. The information can be coded in the syntax using time stamps related to the presentation of the coded audio data and visual data, as well as time stamps related to the delivery of the data stream itself.
[0345] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frame and the audio frame, and supplies the video / audio / subtitle to the video / audio / subtitle display 3216.
[0346] The present invention is not limited to the above-described system, and any of the picture encoding device or the picture decoding device in the above-described embodiments can be incorporated into other systems, for example, an automotive system.
[0347] Although embodiments of the present invention are mainly described based on video coding, embodiments of the coding system 10, the encoder 20, and the decoder 30 (and correspondingly, the system 10), as well as other embodiments described herein, can also be configured for still image processing or still image coding, that is, the processing or coding of individual pictures independent of any preceding or consecutive pictures as in video coding. Note that generally, if picture processing coding is limited to a single picture 17, only the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functionality (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can be equally used for still image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra prediction 254 / 354 and / or loop filter processing 220, 320, as well as entropy coding 270 and entropy decoding 304.
[0348] For example, the functions described herein with reference to, for example, embodiments of encoder 20 and decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on a computer-readable medium as one or more instructions or code, or may be transmitted via a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include a tangible medium such as a data storage medium, or a computer-readable storage medium corresponding to a communication medium that facilitates transfer of a computer program from one place to another, for example, in accordance with a communication protocol. In this way, the computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0349] By way of example, and not limitation, such a computer-readable storage medium can comprise RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can be properly termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. It should be understood, however, that the computer-readable storage medium and data storage medium do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0350] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated logic circuit configurations or discrete logic circuit configurations. Thus, as used herein, the term "processor" may refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within and / or by dedicated hardware configured to encode and decode, or within a combined codec. Also, the techniques may be implemented entirely within one or more circuits or logic elements.
[0351] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to execute the disclosed techniques, but do not necessarily require implementation by various hardware units. Rather, as described above, the various units may be combined within codec hardware units or provided by a set of interoperable hardware units including one or more processors as described above, along with suitable software and / or firmware.
Description of the Signs
[0352] 10 Video coding system, coding system 12 Source device 13 Communication channel 14 Destination device 16 Picture source 17 Picture, picture data, raw picture, raw picture data 18 Preprocessor, preprocessing unit, picture preprocessor 19 Preprocessed picture, preprocessed picture data 20 Video encoder, encoder 21 Encoded picture data, bitstream, encoded bitstream 22, 28 Communication interface, communication unit 30 Video decoder, decoder 31 Decoded picture, decoded picture data 32 Postprocessor, postprocessing unit 33 Postprocessed picture, postprocessed picture data 34 Display device 40 Video coding system 41 Imaging device 42 Antenna 43 Processor 44 Memory store 45 Display device 46 Processing unit 47 Logic circuit configuration 201 Input section, input interface 203 Picture block 204 Residual calculation unit 205 Residual block, residual 206 Transformation processing unit 207 Transformation coefficient 208 Quantization unit 209 Quantization coefficient, quantization transformation coefficient, quantization residual coefficient 210 Inverse quantization unit 211 Inverse quantization coefficient, inverse quantization residual coefficient 212 Inverse transformation processing unit 213 Reconstructed residual block, transformation block 214 Reconstruction unit 215 Reconstruction block 216 Buffer 220 Loop filter, loop filter unit 221 Filter-processed block, filter-processed reconstruction block 230 Decoded picture buffer (DPB) 231 Decoded picture 244 Inter prediction unit 254 Intra prediction unit 260 Mode selection unit 262 Partitioning unit 265 Prediction block 266 Syntax element 270 Entropy encoding unit 272 Output unit, output interface 304 Entropy decoding unit 309 Quantization coefficient 310 Inverse quantization unit 311 Transform coefficient, inverse quantization coefficient 312 Inverse transform processing unit 313 Transform block, reconstructed residual block 314 Reconstruction unit 315 Reconstruction block 320 Loop filter, loop filter unit 321 Filter-processed block 330 Decoded picture buffer (DBP) 331 Decoded picture 344 Inter prediction unit 354 Intra prediction unit 360 Mode selection unit 365 Prediction block 400 Video coding device 410 Inlet port, input port 420 Receiver unit (Rx) 430 Processor, logic unit, central processing unit (CPU) 440 Transmitter unit (Tx) 450 Outlet port, output port 460 Memory 470 Coding module 500 Device 502 Processor 504 Memory 506 Data 508 Operating System 510 Application Program 512 Bus 514 Secondary Storage 518 Display 520 Image Sensing Device 522 Sound Sensing Device 2600 Device 2610 Intra Prediction Unit 2620 Decision Unit 3100 Content Supply System 3102 Capture Device 3104 Communication Link 3106 Terminal Device 3108 Smart Phone, Pad 3110 Computer, Laptop 3112 Network Video Recorder (NVR), Digital Video Recorder (DVR) 3114 TV 3116 Set Top Box (STB) 3118 Video Conference System 3120 Video Surveillance System 3122 Personal Digital Assistant (PDA) 3124 In-Vehicle Device 3126 Display 3202 Protocol Progression Unit 3204 Demultiplexing Unit 3206 Video Decoder 3208 Audio Decoder 3210 Subtitle Decoder 3212 Synchronization Unit 3214 Video / Audio Display 3216 Video / Audio / Subtitle Display
Claims
1. A method of video encoding performed by an encoding device, comprising: executing an intra prediction process of a block to obtain a predicted value of samples of the block, wherein an interpolation filter is applied to reference samples of the block during the intra prediction process of the block; obtaining residual information according to values of the samples of the block and the predicted value of the samples of the block; executing transformation, quantization, and entropy coding on the residual information to obtain an encoded bitstream. The interpolation filter is selected based on a sub-pixel offset between the reference sample and the sample. The size of a main reference side used in the intra prediction process is an integer part of a maximum non-integer value of the sub-pixel offset, and an intra prediction mode is an integer part of the maximum non-integer value of the sub-pixel offset that results in the maximum non-integer value of the sub-pixel offset among a set of available intra prediction modes, a size of a side of the block, half of a length of the interpolation filter, and is determined as a sum thereof. A method.
2. When the intra prediction mode is greater than a vertical intra prediction mode VER_IDX, a side of the block of predicted samples is a width of the block, or when the intra prediction mode is less than a horizontal intra prediction mode HOR_IDX, the side of the block is a height of the block. The method according to claim 1.
3. The method according to claim 1 or 2, wherein in the main reference side, values of reference samples having positions greater than a size doubled of the side of the block are set to be equal to values of samples located at a size doubled of the side of the block.
4. Padding is performed by repeating a first reference sample and / or a last reference sample of the main reference side to a left side and / or a right side, specifically as follows: That is, when the main reference side is denoted as ref and the size of the main reference side is denoted as refS, ref[-1]=p[0] and / or ref[refS + 1]=p[refS] represents the padding. ref[-1] represents the value to the left of the main reference side portion, p[0] represents the value of the first reference sample of the main reference side portion, ref[refS + 1] represents the value to the right of the main reference side portion, p[refS] represents the value of the last reference sample of the main reference side portion, The method according to any one of claims 1 to 3.
5. The method according to any one of claims 1 to 4, wherein the interpolation filter used in the intra prediction process is a finite impulse response filter, and the coefficients of the interpolation filter are fetched from a look-up table.
6. The method according to any one of claims 1 to 5, wherein the interpolation filter used in the intra prediction process is a 4-tap filter.
7. The coefficient c of the interpolation filter 0 , c 1 , c 2 , and c 3 are as follows, that is 【Table 1】 As such, depending on the non-integer part of the sub-pixel offset, the "non-integer part of the sub-pixel offset" column is defined at a 1 / 32 sub-pixel resolution, The method according to claim 6.
8. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 are as follows, that is, 【Table 2】 As such, depending on the non-integer part of the sub-pixel offset, the "non-integer part of the sub-pixel offset" column is defined at a 1 / 32 sub-pixel resolution, The method according to claim 6.
9. The coefficients c of the interpolation filter 0 , c 1 , c 2 , and c 3 are as follows, that is 【Table 3】 As such, depending on the non-integer part of the sub-pixel offset, the "non-integer part of the sub-pixel offset" column is defined at a 1 / 32 sub-pixel resolution, The method according to claim 6.
10. The method according to any one of claims 1 to 9, wherein the interpolation filter is selected from a set of filters used for the intra prediction process for a given sub-pixel offset.
11. The method according to claim 10, wherein the set of filters comprises a Gaussian filter and a cubic filter.
12. The method according to any one of claims 1 to 11, wherein the number of the interpolation filters is N, and the N interpolation filters are used for intra reference sample interpolation, and N >= 1 and is a positive integer.
13. The method according to any one of claims 1 to 12, wherein the reference sample includes samples not adjacent to the block.
14. An encoder comprising a processing circuit configuration configured to execute the method according to any one of claims 1 to 13.
15. A computer program including instructions that enable a computer to execute the following processes when executed on the computer, the processes comprising: Executing an intra prediction process of a block to obtain a predicted value of a sample of the block, wherein an interpolation filter is applied to a reference sample of the block during the intra prediction process of the block; Obtaining residual information according to the value of the sample of the block and the predicted value of the sample of the block; Executing transformation, quantization, and entropy coding on the residual information to obtain an encoded bitstream; The interpolation filter is selected based on a sub-pixel offset between the reference sample and the sample; The size of a main reference side used in the intra prediction process is The integer part of the maximum non-integer value of the sub-pixel offset, and the intra prediction mode is the integer part of the maximum non-integer value of the sub-pixel offset that results in the maximum non-integer value of the sub-pixel offset among a set of available intra prediction modes, The size of the side of the block, Half of the length of the interpolation filter, Determined as the sum of; Computer program.
Citation Information
Patent Citations
Method and apparatus for image interpolation using smoothing interpolation filter
JP2013542666A
Low-complexity interpolation filtering with adaptive tap size
JP2014502822A
Interpolation filters for intra prediction in video coding
US20180091825A1