Quantization matrix prediction for video encoding and decoding
By refining quantization matrix predictions with variable-length residuals and scale factors, and simplifying the encoding syntax, the challenges of complex matrix encoding in existing video compression standards are addressed, resulting in improved compression efficiency and reduced bit cost.
Patent Information
- Application Number
- JP2025025952
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-26
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-09
AI Technical Summary
Existing video compression standards, such as HEVC and VVC, face challenges in efficiently encoding and decoding video data due to the complexity of quantization matrices and the high bit cost associated with custom matrix encoding.
The proposed solution involves improving the prediction of quantization matrices by allowing predictions to be refined using variable-length residuals and scale factors, while also simplifying the syntax and reducing the bit cost of encoding custom matrices.
This approach enhances compression efficiency and reduces the complexity of the video encoding and decoding process, allowing for user-defined trade-offs between accuracy and bit cost while maintaining concise specifications.
Smart Images

Figure 2025072672000001_ABST
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates to video compression, and more particularly, to the quantization stage of video compression. [Background technology]
[0002] To achieve high compression efficiency, image and video coding schemes typically utilize prediction and transformation to leverage spatial and temporal redundancy in video content. In general, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original and predicted picture blocks, often denoted as prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to prediction, transformation, quantization, and entropy coding.
[0003] High Efficiency Video Coding (HEVC) is an example of a compression standard. It was developed by the Joint Collaborative Team on Video Coding (JCT-VC) (see, for example, ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265). Another example of a compression standard is the one under development by the Joint Video Experts Team (JVET) and associated with the development effort named Versatile Video Coding (VVC). VVC is aimed at providing improvements over HEVC. Summary of the Invention
[0004] In general, at least one of the aspects may include a method that includes obtaining, from a bitstream including encoded video information, information representing at least one coefficient of a quantization matrix and a syntax element; determining, based on the syntax element, that the information representing the at least one coefficient is to be interpreted as a residual; and decoding at least a portion of the encoded video information based on a combination of a prediction of the quantization matrix and the residual.
[0005] In general, at least one example of an embodiment may include an apparatus that includes one or more processors configured to obtain information representing at least one coefficient of a quantization matrix and a syntax element from a bitstream that includes encoded video information, determine based on the syntax element that the information representing the at least one coefficient is to be interpreted as a residual, and decode at least a portion of the encoded video information based on a combination of a prediction of the quantization matrix and the residual.
[0006] In general, at least one example of an aspect may include a method that includes obtaining video information and information representing at least one coefficient of a predicted quantization matrix associated with at least a portion of the video information, determining that the at least one coefficient is to be interpreted as a residual, encoding at least a portion of the video information based on a combination of the predicted quantization matrix and the residual, and encoding a syntax element indicating that the at least one coefficient is to be interpreted as a residual.
[0007] In general, at least one example of an aspect may include an apparatus that includes one or more processors configured to obtain video information and information representing at least one coefficient of a predicted quantization matrix associated with at least a portion of the video information, determine that at least one coefficient is to be interpreted as a residual, encode at least a portion of the video information based on a combination of the predicted quantization matrix and the residual, and encode a syntax element indicating that the at least one coefficient is to be interpreted as a residual.
[0008] In general, another example aspect may include a bitstream formatted to include picture information, the picture information being encoded by processing the picture information based on any one or more of the example aspects of the method in accordance with this disclosure.
[0009] In general, one or more other examples of the aspects may also provide a computer-readable recording medium, e.g., a non-transitory computer-readable recording medium, having instructions stored thereon for encoding or decoding picture information, e.g., video data, etc., in accordance with a method or apparatus described herein. Additionally, one or more aspects may also provide a computer-readable recording medium having a bitstream stored thereon generated in accordance with a method or apparatus described herein. Additionally, one or more aspects may provide methods and apparatus for transmitting or receiving a bitstream generated in accordance with a method or apparatus described herein.
[0010] Various modifications and embodiments are envisioned that may provide improvements to video encoding and / or decoding systems including, but not limited to, one or more of increased compression efficiency and / or coding efficiency and / or processing efficiency and / or reduced complexity, as described below.
[0011] The foregoing presents a simplified summary of the subject matter in order to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the subject matter. It is not intended to identify key / critical elements of aspects or to delineate the scope of the subject matter. Its sole purpose is to present some concepts of the subject matter in a simplified form as a prelude to the more detailed description provided below.
[0012] The present disclosure may be better understood from consideration of the following detailed description in conjunction with the accompanying drawings. [Brief description of the drawings]
[0013] [Figure 1] 13 shows an example of transform / quantization coefficients that are inferred to zero for block sizes greater than 32 in VVC. [Diagram 2] A comparison of intra and inter default quantization matrices in H.264 (top: inter (solid) and intra (lattice); bottom: inter (solid) and intra scaling (lattice)). [Diagram 3] A comparison of intra and inter default quantization matrices in HEVC (top: inter (solid) and intra (lattice); bottom: inter (solid) and intra scaling (lattice)). [Figure 4] 1 illustrates an example of a quantization matrix (QM) decoding workflow. [Diagram 5] 1 illustrates an example embodiment of a modified QM decoding workflow. [Figure 6] 13 illustrates another example of an embodiment of a modified QM decoding workflow. [Figure 7] 1 illustrates another example of an embodiment of a QM decoding workflow that includes a prediction (eg, copy) and a variable-length residual. [Figure 8] 1 illustrates another example of an embodiment of a QM decoding workflow that includes prediction (e.g., scaling) and variable-length residuals. [Figure 9] 13 illustrates another example of an embodiment of a QM decoding workflow that involves the consistent use of prediction (eg, scaling) and variable length residuals. [Figure 10] An example embodiment of an encoder, e.g., a video encoder, suitable for implementing various aspects, features, and aspects described herein is illustrated in block diagram form. [Figure 11]An example embodiment of a decoder, e.g., a video decoder, suitable for implementing various aspects, features, and aspects described herein is illustrated in block diagram form. [Figure 12] An example of a system suitable for implementing various aspects, features, and aspects described herein is illustrated in block diagram form. [Figure 13] 4 illustrates another example of an embodiment of a decoder according to the present description. [Figure 14] 1 illustrates another example embodiment of an encoder according to the present description. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] It should be understood that the drawings are for purposes of illustrating examples of the various aspects and embodiments and not necessarily the only possible configurations. Throughout the various drawings, the same reference characters refer to the same or similar features.
[0015] The HEVC specification allows the use of a quantization matrix in the inverse quantization process, where the frequency transformed coefficients of a coded block are scaled by the current quantization stage and then further scaled by a quantization matrix (QM) as follows: d[ x ][ y ]=Clip3( coeffMin, coeffMax, ( (TransCoeffLevel[ xTbY ][ yTbY ][ cIdx ][ x ][ y ] ) * m[ x ][ y ] * levelScale[ qP%6 ] << (qP / 6 ) )+ ( 1 << ( bdShift - 1 ) ) ) >> bdShift ) however, ● TransCoeffLevel[...] is the absolute value of the coefficient being transformed for the current block identified by the spatial coordinates xTbY, yTbY and the component index cIdx. ●x and y are the horizontal / vertical frequency indexes. ●qP is the current quantization parameter. A multiplication by levelScale[qP%6] and a left shift by (qP / 6) is equivalent to a multiplication by the quantization step qStep = (levelScale[qP%6] << (qP / 6)). ●m[...][...] is the two-dimensional quantization matrix. bdShift is an additional scaling factor that takes into account the bit depth of the image samples. The term (1 << (bdShift - 1)) serves the purpose of rounding to the nearest integer. ●d[...] are the absolute values of the resulting dequantized transformed coefficients.
[0016] The syntax used by HEVC to transmit the quantization matrices is as follows:
[0017] [Table 1]
[0018] The following can be noted: ●A different matrix is specified for each transform size (sizeId). For a given transform size, six matrices are specified for intra / inter coding and Y / Cb / Cr components. The matrix can be one of the following: If scaling_list_pred_mode_flag is zero, it is copied from a previously sent matrix of the same size (the reference matrixId is obtained as matrixId - scaling_list_pred_matrix_id_delta), Copied from the default value specified in the standard (if both scaling_list_pred_mode_flag and scaling_list_pred_matrix_id_delta are zero), ○ Fully specified in DPCM coding mode with exp-Golomb entropy coding in the order of up-right diagonal scanning. For block sizes larger than 8x8, to save coding bits, only the 8x8 coefficients are transmitted to signal the quantization matrix. The coefficients are then interpolated using zero-hold (i.e., iterations), except for the DC coefficient, which is explicitly transmitted.
[0019] The use of quantization matrices similar to those in HEVC has been adopted in VVC draft 5 based on contribution JVET-N0847 (see O. Chubach, T. Toma, S. C. Lim, et al., "CE7-related: Support of quantization matrices for VVC", JVET-N0847, Geneva, CH, March 2019). The syntax of scaling_list_data has been adapted for the VVC codec as shown below.
[0020] [Table 2]
[0021] Compared to HEVC, VVC requires more quantization matrices due to a higher number of block sizes.
[0022] Regarding VVC draft 5 (adopted by JVET-N0847), like HEVC, QM is identified by two parameters matrixId and sizeId. What has just been said is illustrated in the following two tables.
[0023] [Table 3]
[0024] [Table 4]
[0025] The combinations of both identifiers are shown in the following table:
[0026] [Table 5]
[0027] Like HEVC, for block sizes larger than 8x8, only the 8x8 coefficients + DC are transmitted. The QM of the correct size is reconstructed using zero-hold interpolation. For example, for a 16x16 block, every coefficient is repeated twice in both directions, and then the DC coefficient is replaced with the one that was transmitted.
[0028] For rectangular blocks, the size (sizeId) retained for QM selection is the maximum of the larger dimension, i.e., width and height. For example, for a 4x16 block, the QM for a block size of 16x16 is selected. The reconstructed 16x16 matrix is then vertically decimated by a factor of 4 to obtain the final 4x16 quantization matrix (i.e., 3 lines out of 4 are skipped).
[0029] In the following, for sizeId and the square block size used, the QM for a given family for block size (square or rectangular) is referred to as size-N, e.g., for block sizes 16x16 or 16x4, the QM is identified as size-16 (sizeId 4 in VVC draft 5). The size-N notation is used to distinguish the exact block shape and the number of QM coefficients signaled, which is limited to 8x8 as shown in Table 3.
[0030] Furthermore, in VVC draft 5, for size 64, the QM coefficients to the bottom right are not transmitted (inferred to 0, referred to as "zero-out" below). What has just been stated is implemented by the "x>=4 && y>=4" condition in the scaling_list_data syntax. What has just been stated avoids transmitting QM coefficients that are never used by the transform / quantization process. Indeed, in VVC, for transform block sizes larger than 32 in either dimension (64xN, Nx64 if N<=64), transformed coefficients with x / y frequency coordinates greater than or equal to 32 are not transmitted and are inferred to zero. As a result, no quantization matrix coefficients need to be quantized. What has just been stated is illustrated in FIG. 1, where the hatched areas correspond to transform coefficients that are inferred to zero.
[0031] In general, an aspect of at least one embodiment described herein can include saving bits in the transmission of a custom QM while keeping the specification as concise as possible. Another aspect of the present disclosure in general can include refining the QM prediction so that simple adjustments are possible with low bit cost.
[0032] For example, two common adjustments applied to QM are: - control of the dynamic range (ratio of lowest value - typically, low frequencies, to highest value - typically, high frequencies), which is easily achieved by a global scaling of the matrix coefficients based on a neutral value of choice; and - General offsets to either tweak bitrate, expand QP range, or balance quality between different transform sizes One example is the 8x8 "intra" and "inter" default matrices for h264 illustrated in Figure 2. Figure 2 is a comparison of intra and inter default h264 matrices. The top example in Figure 2 shows inter (solid) and intra (lattice). The bottom example in Figure 2 shows inter (solid) and scaled intra (lattice). In Figure 2, the "inter" matrix is simply a scaled version of the "intra" matrix, where "scaled intra" is (intra_QM-16)*11 / 16+17.
[0033] A similar trend is seen for the HEVC 8x8 default matrix, illustrated in Figure 3, which is a comparison of intra and inter default h264 matrices. The top example in Figure 3 shows inter (solid) and intra (lattice). The bottom example in Figure 3 shows inter (solid) and scaled intra (lattice), where "scaled intra" is (intra_QM-16)*12 / 16+16.
[0034] The above simple adjustments require a full re-encoding of the matrix in HEVC and VVC, which can be costly (e.g., 240 bits to encode the scaled HEVC intra matrix shown above). Furthermore, in both HEVC and VVC, prediction (=copy) and explicit QM coefficients are mutually exclusive. In other words, it is not possible to refine the prediction.
[0035] In the VVC context, changes have been proposed for QM syntax and prediction. For example, in one proposal (JVET-O0223), QMs are transmitted in order from larger to smaller, allowing prediction from all previously transmitted QMs, including from QMs intended for larger block sizes. "Prediction" here means copying, or decimation if the reference is large. What has just been said exploits the similarity between QMs intended for different block sizes of the same type (Y / Cb / Cr, intra / inter). The QM index matrixId included in the described proposal is a mixture of the size identifier sizeId (shown in Table 4) and the type identifier matrixTypeId (shown in Table 5) combined using the following formula, resulting in a unique identifier matrixId shown in Table 6. matrixId=6 * sizeId+matrixTypeId (3-1)
[0036] [Table 6]
[0037] [Table 7]
[0038] [Table 8]
[0039] The syntax is then changed as shown by the highlighting below (underlining instead of highlighting in this translation):
[0040] [Table 9]
[0041] The QM decoding workflow is illustrated in FIG.
[0042] In Figure 4, -The input is the encoded bitstream. -The output is an array of ScalingMatrix. - "Decode QM prediction modes": Get prediction flags from the bitstream. - "Predicted?": Determines whether QM is inferred (predicted) or signaled in the bitstream, depending on the flags mentioned above. - "Decode QM prediction data": Prediction data needed to infer the QM when not signaled, e.g. the QM index difference scaling_list_pred_matrix_id_delta, is obtained from the bitstream. - "Default?": Determines whether the QM is predicted from a default value (e.g. scaling_list_pred_matrix_id_delta is zero) or from a previously decoded QM. - "The reference QM is the default QM": Select a default QM as the reference QM. There can be multiple default QMs, selected for example depending on the parity of the matrixId. - "Get Reference QM": Select a previously decoded QM as the reference QM. The index of the reference QM is derived from the difference between the matrixId and the previous index. - "Copy or downscale a reference QM": Predict QM from the reference QM. Prediction involves a simple copy if the reference QM is the correct size, or decimation if it is larger than expected. The result is stored in ScalingMatrix[matrixId]. - " Get the number of coefficients: Determines the number of QM coefficients that will be decoded from the bitstream depending on matrixId, e.g., 64 if matrixId is less than 20, 16 if matrixId is between 20 and 25 inclusive, and 4 otherwise. - "Decode QM coefficients": The number of relevant QM coefficients is decoded from the bitstream. - " Diagonal Scanning: Arrange the decoded QM coefficients in a 2D matrix. The result is stored in ScalingMatrix[matrixId]. - "The last QM": Whether to loop or stop when all QMs are decoded from the bitstream. Details regarding the -DC value are omitted for clarity. The proposal described above (JVET-O0223) attempts to reduce the bit cost of encoding custom QMs in the context of VVC by allowing prediction from QMs with different sizeIds. However, prediction is limited to copying (or decimation). In the HEVC context, JCTVC-A114 (Response to the HEVC Call for Proposal) proposed an Inter-QM prediction technique with scaling capabilities. The following syntax and semantics can be found in Annex A of JCTVC-A114, with the most relevant parts shown in bold.
[0043] [Table 10]
[0044] however, - qscalingmatrix_update_type_flag : indicates the method for updating the scaling matrix. A value of 0 means that the new matrix is altered from the existing matrix, a value of 1 means that the new matrix is retrieved from the header. - delta_scale : indicates the delta scale value. The delta_scale value is decoded in Z-scanning order. If not present, the default value for delta_scale is equal to zero. - matrix_scale : Denotes a parameter that scales all matrix values. - matrix_slope : This parameter adjusts the gradient of a non-flat matrix. - num_delta_scale : indicates the number of delta_scale values that will be decoded when qscalingmatrix_update_type_flag is equal to 0. Depending on -qscalingmatrix_update_type_flag, the value of the scaling matrix ScaleMatrix is defined as follows: If -qscalingmatrix_update_type_flag is equal to 0, ScaleMatrix[ i ] = (ScaleMatrix[ i ] * (matrix_scale + 16) + matrix_slope * (ScaleMatrix[ i ] - 16) + 8 )>>4) + delta_scale, where i = 0 to MaxCoeff - Otherwise (qscalingmatrix_update_type_flag is equal to 1), ScaleMatrix[ i ] is derived as shown in Section 5.10. Note: caleMatrix[ i ] is equal to 16 by default. The JCTVC-A114 approach allows for QM prediction with a scaling factor and a variable-length residual, but there are two redundant scaling factors and no easy method to specify a global offset (one would have to code all delta_scale values, since not all delta_scale values are DPCM coded).
[0045] In general, at least one aspect of the present disclosure may include one or more of extending the prediction of QM by applying a scale factor in addition to copying and downsampling, providing a simple way to specify a global offset, or allowing the prediction to be refined by a variable length residual.
[0046] In general, at least one aspect of the present disclosure can include combining an improved prediction with a residual to further refine the QM, with potentially lower cost than full encoding.
[0047] In general, at least one example of an embodiment according to the present disclosure can include QM prediction with offset in addition to copy / downscaling.
[0048] In general, at least one example of an embodiment according to the present disclosure can include QM prediction plus a (variable length) residual.
[0049] In general, at least one example of an embodiment according to the present disclosure can include QM prediction with a scale factor in addition to copying / downscaling.
[0050] In general, at least one example of an aspect according to the present disclosure can include a combination of a scale factor with either an offset or a residual.
[0051] In general, at least one example embodiment according to the present disclosure provides for reducing the number of bits required to transmit a QM, and can allow for a user-defined tradeoff between accuracy and bit cost while keeping specifications simple.
[0052] In general, at least one example of an embodiment according to the present disclosure may include one or more of the following features that provide a reduction in the number of bits required to transmit the QM and allow a user-defined tradeoff between accuracy and bit cost while keeping specifications simple. Add residuals to the QM predictions, ●Adding scaling factors to QM predictions, Combining residuals and scaling ● Always use prediction (remove prediction mode flag).
[0053] Examples of various aspects incorporating one or more of the described features and aspects are provided below. The following examples are given in the form of syntax and text description based on VVC draft 5, including proposal JVET-O0223. However, the context just described is merely an example selected for ease and clarity of explanation, and does not limit the scope of the present disclosure or the scope of application of the principles described herein. In other words, the same principles can be applied in many different situations, as will be apparent to those skilled in the art, for example to HEVC or VVC draft 5, to QM incorporated as QM offsets interpreted as QP offsets instead of scaling factors, to QM coefficients coded with 7 bits instead of 8 bits, etc.
[0054] In general, at least one example embodiment includes a change to a copy for a QM prediction that includes adding a global offset. As an example, the global offset can be specified explicitly as shown in the following changes to the syntax and specification text based on VVC draft 5, including JVET-O0223, where additions are highlighted with shading (underlining instead of highlighting) and deletions are lined out ( / * and * / instead of lined out).
[0055] [Table 11]
[0056] An example of a modified QM decoding workflow corresponding to the embodiment described above is shown in FIG. 5, with the modifications from JVET-O0223 being in bold text and thick outline.
[0057] Alternatively, the global offset can be specified as the first coefficient of the DPCM encoded residual (described below).
[0058] In general, at least one example of an embodiment includes adding a residual to a prediction. For example, it may be desirable to send the residual in addition to the QM prediction to refine the prediction or to provide an optimized coding scheme for QM (prediction plus residual instead of direct coding). The syntax used to send QM coefficients when not predicted can be reused to send the residual when QM is predicted.
[0059] In addition, the number of residual coefficients can be variable. When the coefficients just mentioned are DPCM coded as in HEVC, VVC, and JVET-O0223, and missing coefficients are inferred to zero, this means that the trailing coefficient for the residual is a repetition of the last transmitted residual coefficient. Then, transmitting a single residual coefficient is equivalent to applying a constant offset to the entire QM.
[0060] In general, at least one variant of the described embodiment including adding residuals may further include using a flag when in prediction mode indicating that the coefficients are still transmitted as in non-prediction mode but can be interpreted as residuals. When not in prediction mode, to unify the design, the coefficients can still be interpreted as residuals in addition to the default all-8 QM prediction (value 8 is to mimic the behavior of HEVC and VVC, but is not limited to that value). An example of syntax and semantics based on JVET-O0223 is illustrated below.
[0061] [Table 12]
[0062] scaling_list_pred_mode_flag[matrixId] equal to 0 specifies that the scaling matrix is derived from the values of a reference scaling matrix specified by scaling_list_pred_matrix_id_delta[matrixId]. scaling_list_pred_mode_flag[matrixId] equal to 1 specifies that the values of the scaling list are explicitly signaled. When scaling_list_pred_mode_flag[matrixId] is equal to 1, all values of the (matrixSize) x (matrixSize) array predScalingMatrix[matrixId] and the value of the variable predScalingDC[matrixId] are inferred to be equal to 8. scaling_list_pred_matrix_id_delta[ matrixId ] is as follows: Expected Specifies the reference scaling matrix used to derive the predicted scaling matrix. The value of scaling_list_pred_matrix_id_delta[matrixId] shall be in the range from 0 to matrixId, inclusive. When scaling_list_pred_mode_flag[ matrixId ] is equal to zero, First, the variable refMatrixSize and the array refScalingMatrix are derived as follows: If scaling_list_pred_matrix_id_delta[ matrixId ] is equal to zero, the following applies to set the default value: refMatrixSize is set equal to 8, If matrixId is even, refScalingMatrix = (6-1) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the default value of INTRA { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ● Otherwise, refScalingMatrix = (6-2) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for default value of INTER { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ○ Otherwise (if scaling_list_pred_matrix_id_delta[ matrixId ] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (6-3) refMatrixSize =(refMatrixId < 20) ?8 : (refMatrixId < 26) ?4 : 2 ) (6-4) refScalingMatrix = ScalingMatrix[refMatrixId] (6-5) - Next, the array predS The calingMatrix[ matrixId ] is derived as follows: predS calingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (6-6) where matrixSize = (matrixId < 20) ? 8 : (matrixId < 26) ? 4 : 2 ), x = 0...matrixSize - 1, y = 0...matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ) j = y << ( log2(refMatrixSize) - log2( matrixSize ) ) -If matrixId is less than 14, then if scaling_list_pred_matrix_id_delta[matrixId] is equal to zero then the variable predScalingDC[matrixId] is set equal to 16, otherwise it is set equal to ScalingDC[refMatrixId]. scaling_list_coef_present_flag equal to 1 specifies that the scaling list coefficients are to be transmitted and interpreted as residuals to be added to the predicted quantization matrix. scaling_list_coef_present_flag equal to 0 specifies that no residuals are added. When not present, scaling_list_coef_present_flag equal to 1 is inferred. scaling_list_dc_coef / *_minus8* / [ matrixId ] / *plus 8* / is Used to derive ScalingDC[ matrixId ], Specifies the first value for the scaling matrix when relevant, as described in section xxx. The value of scaling_list_dc_coef / *_minus8* / [ matrixId ] is inclusive. -128 to 127 The range is assumed to be within the range. When not present, the value of scaling_list_dc_coef[ matrixId ] is inferred to be equal to 0. When matrixId is less than 14, the variable ScalingDC[matrixId] is derived as follows: ScalingDC[ matrixId ] = ( predScalingDC[ matrixId ] + scaling_list_dc_coef[ matrixId ] + 256 ) % 256 (6-7) / * when scaling_list_pred_mode_flag[matrixId] is equal to zero, scaling_list_pred_matrix_id_delta[matrixId] is greater than zero, and refMatrixId<14, the following applies: When -matrixId<14, scaling_list_dc_coef_minus8[ matrixId ] is inferred to be equal to scaling_list_dc_coef_minus8[ refMatrixId ]. - Otherwise, ScalingMatrix[ matrixId ]
[0000]
[0000] is set equal to scaling_list_dc_coef_minus8[ refMatrixId ]+8. If scaling_list_pred_mode_flag[ matrixId ] is equal to zero, scaling_list_pred_matrix_id_delta[ matrixId ] is equal to zero (indicating the default value), and matrixId<14, then scaling_list_dc_coef_minus8[ matrixId ] is inferred to be equal to 8. * / scaling_list_delta_coef is scaling_list_coef_present_flagWhen / *scaling_list_pred_mode_flag[matrixId]* / is equal to 1, it specifies the difference between the current matrix coefficient ScalingList[matrixId][i] and the previous matrix coefficient ScalingList[matrixId][i - 1]. The value of scaling_list_delta_coef shall be in the range of -128 to 127, inclusive. / *The value of ScalingList[matrixId][i] shall be greater than 0. * / When scaling_list_coef_present_flag is equal to 0, all values of the array ScalingList[matrixId] are inferred to be equal to zero. The (matrixSize) x (matrixSize) array ScalingMatrix[ matrixId ] is derived as follows: ScalingMatrix [ i ][ j ] = ( predScalingMatrix[ i ][ j ] + ScalingList[ matrixId ][ k] + 256 ) % 256 (6-8) where k = 0 coefNum - 1. i = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0000] , j = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0001] Additionally, in the "Scaling Matrix Derivation Process" section, the text is changed to read as follows: If matrixId is less than 14, m
[0000]
[0000] is further changed as follows: m
[0000]
[0000] = ScalingDC[ matrixId ] / *scaling_list_dc_coef_minus8[ matrixId ] + 8* / (6-9) An example of a modified QM decoding workflow corresponding to the immediately preceding embodiment is shown in FIG. 6, where the changes from JVET-O0223 are indicated in bold text and thick outline.
[0063] In general, at least one example of an aspect can include using a variable length residual in addition to a prediction. A variable length residual can be useful, for example, for the following reasons. ● The low frequency coefficients, which are transmitted first for diagonal scanning, have the greatest impact. ● With the coefficients DPCM coded, if the missing delta coefficients are inferred to zero, this results in the last residual coefficient being repeated. This means that the last transmitted residual is reused for those parts of the QM where no residual is transmitted. The approach just described, transmitting a single residual coefficient, is equivalent to applying an offset to the entire QM, without the use of any specific additional syntax. The number of residual coefficients can be given, for example, by: Explicit number of coefficients by exp-Golomb coding or a fixed-length coding (with length depending on matrixId), or A flag indicating the presence of a residual, followed by the number of residual coefficients minus one, which can be restricted to a power of two; ○ or equivalently, 1 + (log2 of the number of residual coefficients), where zero indicates no residual. What has just been said can be coded as a fixed length (possibly depending on matrixId) or exp-Golomb coding, or The index into the list where the list contains the values of the coefficients. It is suggested that the list just mentioned is fixed by the standard (does not need to be transmitted), e.g. {0,1,2,4,8,16,32,64}. The number of residual coefficients does not need to count the potential extra DC residual coefficient, which can be transmitted at the same time that at least one normal residual coefficient is sent.
[0064] Examples of syntax and semantics based on the immediately preceding exemplary embodiment are provided below.
[0065] [Table 13]
[0066] scaling_list_pred_mode_flag[matrixId] equal to 0 specifies that the scaling matrix is derived from the values of a reference scaling matrix specified by scaling_list_pred_matrix_id_delta[matrixId]. scaling_list_pred_mode_flag[matrixId] equal to 1 specifies that the values of the scaling list are explicitly signaled. When scaling_list_pred_mode_flag[matrixId] is equal to 1, all values of the matrixSize × matrixSize array predScalingMatrix[matrixId] and the value of the variable predScalingDC[matrixId] are inferred to be equal to 8. scaling_list_pred_matrix_id_delta[matrixId] specifies the reference scaling matrix used to derive the predicted scaling matrix as follows: The value of scaling_list_pred_matrix_id_delta[matrixId] shall be in the range from 0 to matrixId, inclusive. When scaling_list_pred_mode_flag[ matrixId ] is equal to zero, First, the variable refMatrixSize and the array refScalingMatrix are derived as follows: If scaling_list_pred_matrix_id_delta[ matrixId ] is equal to zero, the following applies to set the default value: refMatrixSize is set equal to 8, If matrixId is even, refScalingMatrix = (6-10) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the default value of INTRA { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ● Otherwise, refScalingMatrix = (6-11) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for default value of INTER { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ○ Otherwise (if scaling_list_pred_matrix_id_delta[ matrixId ] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (6-12) refMatrixSize = (refMatrixId < 20) ?8 : (refMatrixId < 26) ?4 : 2 ) (6-13) refScalingMatrix = ScalingMatrix[ refMatrixId ] (6-14) -The array predScalingMatrix[matrixId] is then derived as follows: predScalingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (6-15) where / *matrixSize = (matrixId < 20) ? 8 : (matrixId < 26) ? 4 : 2 ),* / x = 0...matrixSize - 1, y = 0...matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ), and j = y << ( log2(refMatrixSize) - log2( matrixSize ) ) If -matrixId<14, then if scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, then the variable predScalingDC[matrixId] is set equal to 16, otherwise it is set equal to ScalingDC[refMatrixId]. coefNum specifies the number of residual coefficients to be transmitted. The value of coefNum shall be in the range from 0 to maxCoefNum, inclusive. / * scaling_list_coef_present_flag equal to 1 specifies that the scaling list coefficients are to be transmitted and interpreted as residuals to be added to the predicted quantization matrix. scaling_list_coef_present_flag equal to 0 specifies that no residuals are added. When not present, scaling_list_coef_present_flag is inferred to be equal to 1. * / scaling_list_dc_coef[matrixId] specifies the first value for the scaling matrix used to derive ScalingDC[matrixId], when relevant, as described in section xxx. The value of scaling_list_dc_coef[matrixId] shall be in the range of -128 to 127, inclusive. When not present, the value of scaling_list_dc_coef[matrixId] is inferred to be equal to 0. When matrixId is less than 14, the variable ScalingDC[matrixId] is derived as follows: ScalingDC[ matrixId ] = ( predScalingDC[ matrixId ] + scaling_list_dc_coef[ matrixId ] + 256 ) % 256 (6-16) scaling_list_delta_coef is For i between 0 and CoefNum-1 inclusive, Specifies the difference between the current matrix coefficient ScalingList[matrixId][i] and the previous matrix coefficient ScalingList[matrixId][i - 1]. For i between CoefNum and maxCoefNum-1 inclusive, scaling_list_delta_coef is / *When scaling_list_coef_present_flag is equal to 1,* / Inferred to be equal to 0 The value of scaling_list_delta_coef shall be in the range of -128 to 127, inclusive. / *When scaling_list_coef_present_flag is equal to 0, all values of the array ScalingList[matrixId] are inferred to be equal to zero. * / The (matrixSize) x (matrixSize) array ScalingMatrix[ matrixId ] is derived as follows: ScalingMatrix [ i ][ j ] = ( predScalingMatrix[ i ][ j ] + ScalingList[ matrixId ][ k ] + 256 ) % 256 (6-17) where k = 0... maxC oefNum- 1, i = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][k ]
[0000] j = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][k ]
[0001] An example of a modified QM decoding workflow is shown in Figure 7. The example embodiment shown in Figure 7 includes prediction = copy; variable length residual. In Figure 7, the changes from the example embodiment illustrated in Figure 6 are in bold text and thick outlines.
[0067] In general, at least one example of the embodiment includes adding a scale factor to the prediction. For example, the prediction of the QM can be enhanced by adding a scale factor to the copy / decimation. When combined with the preceding example of the embodiment, what has just been described allows two common adjustments, scaling and offset, to be made to the QM described above, and also allows for a user-defined tradeoff between accuracy and bit cost with a variable length residual. The scale factor can be an integer normalized to a number N (e.g., 16). The scaling formula can scale the reference QM based on a neutral value of choice (e.g., 16 for HEVC, VVC, or JVET-O0223, but would be zero if the QM is interpreted as a QP offset). An example formula with N = (1 << shift) = 16 is shown below. scaledQM = ( ( refQM - 16 ) * scale + rnd ) >> shift ) + 16 (6-18) however, shift = 4 and rnd = 8 means N = ( 1 << 4 ) = 16 refQM is the input QM to be scaled ●scale is the scaling factor ●scaledQM is the output scaled QM The resulting scaled QM should then be clipped to the QM range (e.g. 0...255 or 1...255).
[0068] The scaling factor scale can be transmitted as an offset relative to N, so that the value 0 represents a neutral value (no scaling) with the minimum number of bits cost, 1 bit when encoded as exp-Golomb (in the example above, the value transmitted would be scale - ( 1 << shift ) ). Furthermore, the values can be restricted to a practical range, e.g., a value between -32 and 32 means scaling between -1 and +3, or a value between -16 and 16 means scaling between 0 and +2 (thus restricting scaling to positive values).
[0069] An example syntax is shown below. Note that scale factors are not necessarily combined with offsets or residuals, but may be applied alone in addition to JVET-O0223, VVC, or HEVC.
[0070] [Table 14]
[0071] scaling_list_pred_mode_flag[matrixId] equal to 0 specifies that the scaling matrix is derived from the values of a reference scaling matrix specified by scaling_list_pred_matrix_id_delta[matrixId]. scaling_list_pred_mode_flag[matrixId] equal to 1 specifies that the values of the scaling list are explicitly signaled. When scaling_list_pred_mode_flag[matrixId] is equal to 1, all values of the matrixSize × matrixSize array predScalingMatrix[matrixId] and the value of the variable predScalingDC[matrixId] are inferred to be equal to 8. scaling_list_pred_matrix_id_delta[matrixId] specifies the reference scaling matrix used to derive the predicted scaling matrix as follows: The value of scaling_list_pred_matrix_id_delta[matrixId] shall be in the range from 0 to matrixId, inclusive. scaling_list_pred_scale_minus16[matrixId] plus 16 specifies the scale factor to be applied to the reference scaling matrix used to derive the predicted scaling matrix, as follows: The values of scaling_list_pred_scale_minus16[matrixId] shall be in the range of -16 to +16, inclusive. When scaling_list_pred_mode_flag[ matrixId ] is equal to zero, First, the variable refMatrixSize and the array refScalingMatrix are derived as follows: If scaling_list_pred_matrix_id_delta[ matrixId ] is equal to zero, the following applies to set the default value: refMatrixSize is set equal to 8, If matrixId is even, refScalingMatrix = (6-19) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the default value of INTRA { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ● Otherwise, refScalingMatrix = (6-20) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for default value of INTER { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ○ Otherwise (if scaling_list_pred_matrix_id_delta[ matrixId ] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (6-21) refMatrixSize = (refMatrixId < 20) ?8 : (refMatrixId < 26) ?4 : 2 ) (6-22) refScalingMatrix = ScalingMatrix[ refMatrixId ] (6-23) -The array predScalingMatrix[matrixId] is then derived as follows: predScalingMatrix[ matrixId ][ x ][ y ] = Clip3( 1, 255, ( ( ( refScalingMatrix[ i ][ j ] - 16 ) * scale + 8 ) >> 4 ) )+ 16 ) (6-24) where x = 0 matrixSize - 1, y = 0 matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ) / *, and* / j = y << ( log2(refMatrixSize) - log2( matrixSize ) ) , and scale = scaling_list_pred_scale_minus16[ matrixId ] + 16 -If matrixId<14, then if scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, then the variable predScalingDC[matrixId] is set equal to 16, otherwise. As follows: / * Set equal to ScalingDC[ refMatrixId ]. * / It is derived. predScalingDC[ matrixId ] = Clip3( 1, 255, ( ( ( ScalingDC[ refMatrixId ] - 16 ) * scale + 8 ) >> 4 ) )+ 16 ) (6-25) where scale = scaling_list_pred_scale_minus16[ matrixId ] + 16 coefNum specifies the number of residual coefficients to be transmitted. The value of coefNum shall be in the range from 0 to maxCoefNum, inclusive. scaling_list_dc_coef[matrixId] specifies the first value for the scaling matrix used to derive ScalingDC[matrixId], when relevant, as described in section xxx. The value of scaling_list_dc_coef[matrixId] shall be in the range of -128 to 127, inclusive. When not present, the value of scaling_list_dc_coef[matrixId] is inferred to be equal to 0. When matrixId is less than 14, the variable ScalingDC[matrixId] is derived as follows: ScalingDC[ matrixId ] = ( predScalingDC[ matrixId ] + scaling_list_dc_coef[ matrixId ] + 256 ) % 256 (6-26) scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[matrixId][i][i - 1] and the previous matrix coefficient ScalingList[matrixId][i - 1] for i between 0 and CoefNum-1, inclusive. For i between CoefNum and maxCoefNum-1, inclusive, scaling_list_delta_coef is inferred to be equal to 0. Values of scaling_list_delta_coef shall be in the range of -128 to 127, inclusive. The (matrixSize) x (matrixSize) array ScalingMatrix[ matrixId ] is derived as follows: ScalingMatrix [ i ][ j ] = ( predScalingMatrix[ i ][ j ] + ScalingList[ matrixId ][ k ] + 256 ) % 256 (6-27) where k = 0 maxCoefNum - 1. i = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][ k ]
[0000] , j = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][ k ]
[0001] An example of a modified QM decoding workflow is shown in Figure 8. The example embodiment shown in Figure 9 includes prediction = scaling; variable length residual. In Figure 8, the changes from the example embodiment illustrated in Figure 7 are in bold text and thick outlines.
[0072] In general, at least one example of the embodiment includes removing the prediction mode flag. For example, as long as the complete residual can be transmitted, such as in the embodiment described above, the full specification of QM is possible regardless of the scaling_list_pred_mode_flag flag, so this flag can be removed, simplifying the specification text with only minimal impact on the bit cost.
[0073] The following provides examples of syntax and semantics of the current embodiment, based on modifications to the previous embodiment.
[0074] [Table 15]
[0075] / * scaling_list_pred_mode_flag[ matrixId ] equal to 0 specifies that the scaling matrix is derived from the values of a reference scaling matrix. The reference scaling matrix is specified by scaling_list_pred_matrix_id_delta[ matrixId ]. scaling_list_pred_mode_flag[ matrixId ] equal to 1 specifies that the values of the scaling list are explicitly signaled. * / / *When scaling_list_pred_mode_flag[matrixId] is equal to 1, all values of the matrixSize × matrixSize array predScalingMatrix[matrixId] and the value of the variable predScalingDC[matrixId] are inferred to be equal to 8. * / scaling_list_pred_matrix_id_delta[matrixId] specifies the reference scaling matrix used to derive the predicted scaling matrix. The value of scaling_list_pred_matrix_id_delta[matrixId] shall be in the range from 0 to matrixId, inclusive. scaling_list_pred_scale_minus16[matrixId] plus 16 specifies the scale factor to be applied to the reference scaling matrix used to derive the predicted scaling matrix, as follows: The values of scaling_list_pred_scale_minus16[matrixId] shall be in the range of -16 to +16, inclusive. / *When scaling_list_pred_mode_flag[ matrixId ] is equal to zero, * / First, the variable refMatrixSize and the array refScalingMatrix are derived as follows: If scaling_list_pred_matrix_id_delta[ matrixId ] is equal to zero, the following applies to set the default value: refMatrixSize is set equal to 8, If matrixId is even, refScalingMatrix = (6-28) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the default value of INTRA { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ● Otherwise, refScalingMatrix = (6-29) { { 16, 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for default value of INTER { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, ○ Otherwise (if scaling_list_pred_matrix_id_delta[ matrixId ] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (6-30) refMatrixSize = (refMatrixId < 20) ?8 : (refMatrixId < 26) ?4 : 2 ) (6-31) refScalingMatrix = ScalingMatrix[ refMatrixId ] (6-32) -The array predScalingMatrix[matrixId] is then derived as follows: predScalingMatrix[ matrixId ][ x ][ y ] = Clip3( 1, 255, ( ( ( refScalingMatrix[ i ][ j ] - 16 ) * scale + 8 ) >> 4 ) )+ 16 ) (6-33) where x = 0 matrixSize - 1, y = 0 matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ) j = y << ( log2(refMatrixSize) - log2( matrixSize ) ), and scale = scaling_list_pred_scale_minus16[ matrixId ] + 16 If -matrixId<14, then if scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, then the variable predScalingDC[matrixId] is set equal to 16, otherwise it is derived as follows: predScalingDC[ matrixId ] = Clip3( 1, 255, ( ( ( ScalingDC[ refMatrixId ] - 16 ) * scale + 8 ) >> 4 ) )+ 16 ) (6-34) where scale = scaling_list_pred_scale_minus16[ matrixId ] + 16 coefNum specifies the number of residual coefficients to be transmitted. The value of coefNum shall be in the range from 0 to maxCoefNum, inclusive. scaling_list_dc_coef[matrixId] specifies the first value for the scaling matrix used to derive ScalingDC[matrixId], when relevant, as described in section xxx. The value of scaling_list_dc_coef[matrixId] shall be in the range of -128 to 127, inclusive. When not present, the value of scaling_list_dc_coef[matrixId] is inferred to be equal to 0. When matrixId is less than 14, the variable ScalingDC[matrixId] is derived as follows: ScalingDC[ matrixId ] = ( predScalingDC[ matrixId ] + scaling_list_dc_coef[ matrixId ] + 256 ) % 256 (6-35) scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[matrixId][i][i - 1] and the previous matrix coefficient ScalingList[matrixId][i - 1] for i between 0 and CoefNum-1, inclusive. For i between CoefNum and maxCoefNum-1, inclusive, scaling_list_delta_coef is inferred to be equal to 0. Values of scaling_list_delta_coef shall be in the range of -128 to 127, inclusive. The (matrixSize) x (matrixSize) array ScalingMatrix[ matrixId ] is derived as follows: ScalingMatrix [ i ][ j ] = ( predScalingMatrix[ i ][ j ] + ScalingList[ matrixId ][ k ] + 256 ) % 256 (6-36) where k = 0 maxCoefNum - 1. i = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][ k ]
[0000] , j = diagScanOrder[ log2(matrixSize) ][ log2(matrixSize) ][ k ]
[0001] An example of a modified QM decoding workflow is shown in Figure 9. The example embodiment shown in Figure 9 includes: Always use prediction; Prediction = Scaling; Variable length residual. In Figure 9, the changes from the example embodiment illustrated in Figure 8 include the removal of the features of Figure 8.
[0076] At least some examples of various other embodiments are taken from current VVC techniques described above. 回 Some contributions submitted for change to the JVET conference are described below. Most people could consider equivalent changes to the techniques proposed in this disclosure that impact the example syntax and semantics by affecting the number of QMs, the mapping of QM identifiers, the QM matrix size formulas, the QM selection and / or resizing process, but without changing the essence. Examples of aspects that provide accommodation for changes are detailed in the following subsections.
[0077] At least one example of an aspect may include the removal of some or all of the 2×2 chroma QM. Recently, VVC removed the 2×2 chroma transform block in intra mode, thus wasting sending the 2×2 QM intra QM, but the waste may be limited to 2×3=6 bits (3 bits are needed to signal the use of the default matrix with the example syntax described above).
[0078] Removing the luminance 2x2 QM will affect the number of QMs sent with a simple syntax change.
[0079] [Table 16]
[0080] And the identifier mapping will also be affected (ids 26 and 28 will be invalid from Table 6), which could be reflected by adjusting the matrixId selection formula in the "Quantization Matrix Derivation Process" from JVET-O0223 (or the update of JVET-P0110) as follows: matrixId = 6 * sizeId + matrixTypeId where log2TuWidth = log2(blkWidth) + (cIdx > 0)? log2(SubWidthC) : 0, log2TuHeight = log2( blkHeight ) + ( cIdx > 0 ) ?log2( SubHeightC ) : 0, sizeId = 6 - max( log2TuWidth, log2TuHeight ), and matrixTypeId = ( 2 * cIdx + ( predMode = MODE_INTRA ? 0 : 1 ) ) matrixId -= (matrixId > 25) ?(matrixId > 27 ? : 2 : 1) : 0 Removing all 2x2 QMs and disabling QM for 2x2 chroma inter blocks would make the complexity even less by reducing the number of QMs to 26, as shown in the table below.
[0081] [Table 17]
[0082] Then, do not call the “Scaling process for transform coefficients” in the VVC “Derivation process for scaling matrix”, but instead add that case to the exception list. - The intermediate (nTbW) x (nTbH) scaling factor array m is derived as follows: - m[ x ][ y ] is set equal to 16 if one or more of the following conditions are true: -sps_scaling_list_enabled_flag is equal to 0, -transform_skip_flag[xTbY][yTbY] is equal to 1, - Both nTbW and nTbH are equal to 2 , Otherwise, m is the output of the derivation process for matrix scaling specified in subclause 8.7.1, invoked with the prediction mode CuPredMode[xTbY][yTbY], the color component variables cIdx, the block width nTbW, and the block height nTbH as input. Alternatively, when the block size is 2x2 or similar, assign all-16 values to the m[][] array in the “Derivation process for scaling matrix”.
[0083] At least one other embodiment includes larger sizes for a luma QM of size 64. The matrix size formula can be modified to enable, for example, 16×16 coefficients for a luma QM of size 64, with an example syntax shown below.
[0084] [Table 18]
[0085] It will be apparent to one skilled in the art that the refMatrixSize in the exemplary semantics of the above described embodiments shall follow the same rules and be changed accordingly.
[0086] At least one other aspect includes adding a joint CbCr QM. It may be possible that a specific QM is desired for the joint CbCr residual coded block. This can be added after the regular chrominance one, resulting in the following new identity mapping table:
[0087] [Table 19]
[0088] As shown below, this affects the number of QMs, the matrixSize formula, the DC coefficient presence condition (changing from <14 to <18), and the matrixId formula in the QM selection process.
[0089] [Table 20]
[0090] matrixId = 8 / *6* / * sizeId + matrixTypeId where log2TuWidth = log2(blkWidth) + (cIdx > 0)? log2(SubWidthC) : 0, log2TuHeight = log2( blkHeight ) + ( cIdx > 0 ) ?log2( SubHeightC ) : 0, sizeId = 6 - max( log2TuWidth, log2TuHeight ), and matrixTypeId = ( 2 * cIdx + ( predMode = MODE_INTRA ? 0 : 1 ) ) Other relevant places in the semantics (such as refMatrixSize) are affected by corresponding changes that will be obvious to those skilled in the art.
[0091] At least one other embodiment includes adding a QM to LNFST. LNFST is a secondary transform used in VVC that changes the spatial frequency meaning of the final transform coefficients, thus affecting the meaning of the QM coefficients. A specific QM may be desired. Since the output of LNFST is limited to 4x4 and has two alternative transform cores, the QM just mentioned could be added as a spurious TU size inserted before or after the normal size 4 (keeping the QMs ordered in decreasing size), as shown in the table below. Since LNFST is for intra mode only, intra / interline would be used to distinguish the core transform.
[0092] [Table 21]
[0093] The syntax and semantics should be modified accordingly to reflect the new matrixId mapping, number of QMs, matrixSize, DC coefficient conditions, etc., as in the previous example.
[0094] The derivation of matrixId in the "Quantization matrix derivation process" that selects the correct QM should also be updated, either by the tabulated description as above, or by revising the formula / pseudo-code as in the following example. if (lfnst_idx[ xTbY ][ yTbY ] != 0) matrixId = 24 + 2*cIdx + lfnst_idx[ xTbY ][ yTbY ] else { matrixId = [...] if (matrixId > 23) matrixId += 6 } At least one other aspect includes handling QM with LNFST by forcing the use of 4x4QM. This can be implemented, for example, by forcing a specific matrixId to be used for the current block in the "quantization matrix derivation process" in the case of LNFST, as follows: if (lfnst_idx[ xTbY ][ yTbY ] != 0) matrixId = 20 + ((cIdx==0) ?10 : 2*cIdx) + lfnst_idx[ xTbY ][ yTbY ] else matrixId = [...] This application describes various aspects including tools, features, aspects, models, approaches, etc. For example, the various aspects include, but are not limited to, the following: - To save bits when sending custom QMs while keeping the specification as simple as possible, - Refine QM predictions so that simple adjustments are possible with low bit cost; -Providing an easy way to specify global offsets, One or more of extending QM prediction by applying a scale factor in addition to copying and downsampling, providing a simple way to specify a global offset, and allowing predictions to be refined by variable length residuals; Combining the improved prediction with the residual to further refine the QM, with a potentially lower cost than the full encoding; - Providing QM predictions that use offsets in addition to copying / downscaling, - Providing QM predictions with (variable length) residuals; - Providing QM predictions using scale factors in addition to copying / downscaling, - Providing a QM prediction that includes a combination of a scale factor with either an offset or a residual; -Providing a reduction in the number of bits required to transmit the QM and allowing a user-defined trade-off between accuracy and bit cost while keeping the specification concise; - Providing a reduction in the number of bits required to transmit the QM and allowing a user-defined trade-off between accuracy and bit cost while keeping specifications concise, based on including one or more of the following features: Add residuals to the QM predictions, Adding a scaling factor to the QM predictions, Combine residuals and scaling, ○ Always use prediction (remove prediction mode flag).
[0095] Many of the features just described are described with particularity and in a manner that may often give the impression of limitation, at least to indicate their individual characteristics. However, what is just described is for purposes of clarity in description and does not limit the application or scope of the features just described. Indeed, all of the different features can be combined and substituted to provide further features. Furthermore, embodiments can be combined and substituted with features described in earlier applications as well.
[0096] The aspects described and contemplated in this application can be implemented in many different ways. Figures 10, 11, and 12 provide some aspects below, but other aspects are contemplated, and the discussion of Figures 10, 11, and 12 does not limit the breadth of implementations. In general, at least one of the aspects relates to video encoding and decoding, and in general, at least one other aspect relates to transmitting the generated or encoded bitstream. The just-described aspects and other aspects can be implemented as a method, an apparatus, a computer-readable medium having instructions stored thereon, and / or a computer-readable medium having a bitstream generated according to any of the described methods, for encoding or decoding video data according to any of the described methods.
[0097] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably.
[0098] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the inherent operation of the method, the order and / or use of specific steps and / or actions may be varied or combined.
[0099] Various methods and other aspects described in this application can be used to modify modules, such as the quantization and dequantization modules (130, 140, 240) of the video encoder 100 and decoder 200 shown in Figures 10 and 11. Furthermore, the aspects are not limited to VVC or HEVC, but can be applied to other standards and recommendations, such as those previously existing or developed in the future, and to extensions of any of the above standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically prohibited, the aspects described in this application can be used individually or in combination.
[0100] Various numerical values, e.g., the size of the maximum quantization matrix, the number of block sizes considered, are used in this application, the specific values are for illustrative purposes, and the aspects described are not limited to the specific values just mentioned.
[0101] 10 illustrates an encoder 100. Although variations of encoder 100 are envisioned, encoder 100 is described below for purposes of clarity without describing all contemplated variations. Before being encoded, a video sequence may go through a pre-encoding process (101), for example, applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or remapping the input picture components to make the signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata is associated with the pre-processing and can be tied to the bitstream.
[0102] In the encoder 100, a picture is encoded by the encoder elements described below. The picture to be encoded is divided (102) and processed, for example, in units of CUs. Each unit is encoded, for example, using either intra or inter mode. When a unit is encoded in intra mode, intra prediction (160) is performed. In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) which one of intra or inter mode to use to encode the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. A prediction residual is calculated, for example, by subtracting (110) the predicted block from the original image block.
[0103] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can also bypass both the transform and quantization, i.e., the residual is coded directly without applying either the transform or quantization process.
[0104] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined (155) with the predicted block to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed picture, for example to perform deblocking / Sample Adaptive Offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).
[0105] Figure 11 illustrates a block diagram of a video decoder 200. In the decoder 200, the bitstream is decoded by the decoder elements described below. In general, the video decoder 200 performs a decoding path that is the reverse of the encoding path described in Figure 10. Additionally, the encoder 100 also generally performs video decoding as part of encoding the video data.
[0106] In particular, the decoder input includes a video bitstream that may be generated by the video encoder 100. First, the bitstream is entropy decoded (230) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. The decoder may then partition (235) the picture according to the decoded picture partition information. The transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined (255) with the predicted block to reconstruct an image block. The predicted block may be obtained (270) from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0107] Additionally, the decoded pictures may go through a post-decoding process (285), such as an inverse color transform (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that reverses the remapping process performed in the pre-encoding process (101). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0108] FIG. 12 illustrates a block diagram of an example of a system in which various aspects and aspects are implemented. System 1000 can be embodied as a device including various components described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as, for example, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television broadcast receivers, personal video recording systems, connected consumer electronics appliances, and servers. The elements of system 1000, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various aspects, the system 1000 is configured to implement one or more aspects described in this document. The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 can include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which can include non-volatile and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Flash, magnetic disk drives, and / or optical disk drives. The storage devices 1040 may include, by way of non-limiting example, internal storage devices, attached storage devices (including detachable and non-detachable storage devices), and / or network-accessible storage devices.
[0109] The system 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, encoded or decoded video, which may include its own processor and memory. The encoder / decoder module 1030 represents a module or modules that may be included in a device that performs encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000 or may be incorporated within the processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0110] Program code loaded onto the processor 1010 or encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various aspects, one or more of the processors 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, formulas, operations, and operational logic.
[0111] In some embodiments, memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of the just-mentioned functions. The external memory can be the memory 1020 and / or the storage device 1040, for example, a dynamic volatile memory and / or a non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store, for example, the operating system of the television. In at least one embodiment, a high speed external dynamic volatile memory such as a RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, which is further called ISO / IEC 13818, which is further known as H.222, and which is further known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, which is further known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a standard being developed by JVET, the Joint Video Experts Team).
[0112] Inputs to the elements of system 1000 can be provided via various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives a radio frequency (RF) signal transmitted, for example, over the air by a broadcast station, (ii) a Component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples not shown in FIG. 10 include composite video.
[0113] In various aspects, the input devices of block 1130 are associated with respective input processing elements known in the art. For example, the RF section can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies that selects a signal frequency band, which in some aspects may be referred to as a channel (e.g.), (iv) demodulating the down-converted band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various aspects includes one or more elements that perform the functions just described, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include, for example, a tuner that performs various functions, including down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one set-top box embodiment, the RF section and associated input processing elements receive, filter, downconvert, and perform frequency selection by refiltering RF signals transmitted over a wired medium (e.g., cable) to the desired frequency band. Various embodiments rearrange the order of the above-mentioned (and other) elements, remove some of the elements just mentioned, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0114] Additionally, the USB and / or HDMI terminals can include respective interface processors to couple the system 1000 to other electronic devices over USB and / or HDMI connections. It is understood that various aspects of the input processing, e.g., Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or within the processor 1010, as appropriate. Similarly, aspects of the USB or HDMI interface processing can be implemented in a separate input processing IC or within the processor 1010, as appropriate. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, the processor 1010 and an encoder / decoder 1030, which operates in cooperation with memory and storage elements to process the data stream for presentation at an output device, as appropriate.
[0115] The various elements of the system 1000 can be provided in an integrated housing in which the various elements can be interconnected and transmit data therebetween using a suitable coupling arrangement 1140, such as an internal bus known in the art, including an I2C (Inter-IC) bus, wiring, and printed circuit boards.
[0116] The system 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented within a wired and / or wireless medium, for example.
[0117] Data is streamed or otherwise provided to the system 1000 in various embodiments using a wireless network, such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE citing the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in the embodiment just described is received via a communication channel 1060 and communication interface 1050 that are tailored for Wi-Fi communication. Typically, the communication channel 1060 in the embodiment just described is coupled to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 1000 using a set-top box that delivers data via an HDMI connection in the input block 1130. Still other embodiments provide streamed data to the system 1000 using an RF connection in the input block 1130. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0118] The system 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a cellular (mobile) phone, or other device. Furthermore, the display 1100 can be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). In various example embodiments, the other peripheral devices 1120 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments employ one or more peripheral devices 1120 that provide functionality based on the output of the system 1000. For example, a disc player is responsible for playing the output of the system 1000.
[0119] In various aspects, control signals are communicated between the system 1000 and the display 1100, speaker 1110, or other peripheral device 1120 using signaling such as AV.Link, CEC (Consumer Electronics Control), or other communication protocols that allow device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be coupled to the system 1000 via the communication interface 1050 using the communication channel 1060. The display 1100 and speaker 1110 can be integrated with other components of the system 1000 in a single unit, such as in an electronic device such as a television. In various aspects, the display interface 1070 includes a display driver such as, for example, a timing controller (T Con) chip.
[0120] Alternatively, the display 1100 and speakers 1110 can be separate from one or more other components, for example if the RF portion of the input 1130 is part of a separate set-top box. In various aspects where the display 1100 and speakers 1110 are external components, the output signal can be submitted via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0121] The embodiments can be performed by computer software implemented by the processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories, and removable memories. The processor 1010 can be of any type suitable for the technical environment and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0122] Another example of the aspect is illustrated in FIG. 13. In FIG. 13, input information, e.g., a bitstream, includes encoded video information and other information, e.g., information associated with the encoded video information, e.g., a quantization matrix or matrices, and control information, e.g., one or more syntax elements. At 1310, information representing at least one coefficient of the quantization matrix and a syntax element is obtained from an input, e.g., an input bitstream. At 1320, based on the syntax element, it is determined that the information representing the at least one coefficient is to be interpreted as a residual. Then, at 1330, at least a portion of the encoded video information is decoded based on a combination of a prediction of the quantization matrix and the residual. One or more examples of the features shown and described with respect to FIG. 13 have been illustrated and described previously herein, e.g., with respect to FIG. 6.
[0123] Another example of an aspect is illustrated in FIG. 14. In FIG. 14, an input including video information is processed at 1410 to obtain the video information and information representing at least one coefficient of a predicted quantization matrix associated with at least a portion of the video information. At 1420, it is determined that at least one coefficient is to be interpreted as a residual. For example, the processing of the video information may indicate that use of the residual provides advantageous processing of some video information, such as providing improved compression efficiency when encoding the video information, or improved video quality when decoding the encoded video information. Then, at 1430, at least a portion of the video information is encoded based on a combination of the predicted quantization matrix and the residual. The encoding at 1430 further encodes a syntax element indicating that at least one coefficient is to be interpreted as a residual.
[0124] Various implementations and examples of the aspects described herein include decoding. As used herein, "decoding" can include all or a portion of the processing performed on a received encoded sequence to generate a final output suitable for, for example, a display. In various aspects, the processing includes one or more of the processing typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various aspects, the processing also or alternatively includes the processing performed by a decoder in the various implementations described herein.
[0125] As a further example, in one aspect "decoding" refers only to "entropy decoding," in another aspect "decoding" refers only to differential decoding, and in another aspect "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding" is intended to refer specifically to a subset of operations or more generally, it is believed that the decoding process will be clear based on the context of the particular description and will be well understood by those of ordinary skill in the art.
[0126] Various implementations include encoding. In a manner similar to the above discussion of "decoding," "encoding" as used herein can include all or part of the processing performed on an input video sequence, for example, to generate an encoded bitstream. In various aspects, such processing includes one or more of the processing typically performed by an encoder, such as partitioning, differential encoding, transforming, quantizing, and entropy encoding.
[0127] As a further example, in one embodiment "encoding" refers only to "entropy decoding," while in another embodiment "encoding" refers only to differential decoding, while in another embodiment "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding" is intended to refer specifically to a subset of operations or more generally, it is believed that the encoding process will be clear based on the context of the particular description and will be well understood by those of ordinary skill in the art.
[0128] It is noted that the syntax elements used herein, e.g., scaling_list_pred_mode_flag, scaling_list_pred_matrix_id_delta, scaling_list_dc_coef_minus8, are descriptive terms. As such, they do not preclude the use of other syntax element names.
[0129] It should be understood that when a diagram is provided as a flow diagram, it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is provided as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0130] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Additionally, a processor may also include a communication device, such as, for example, a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information between end users.
[0131] Reference to "one embodiment" or "embodiment" or "one implementation" or "implementation" means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment, as well as other variations. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" appearing in various places throughout this application are not necessarily all referring to the same embodiment, as well as any other variations.
[0132] Additionally, the application may refer to "determining" various portions of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0133] Additionally, the application may refer to "accessing" various portions of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0134] Additionally, the application may refer to "receiving" various portions of information. Receiving, as with "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically included among operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information in some way or another.
[0135] It should be understood that the use of any of the following " / ", "and / or" and "at least one of" is intended to include, for example, in the case of "A / B", "A and / or B" and "at least one of A and B", the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", the phrase is intended to include the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of the second and third listed alternatives (B and C), or the selection of all three alternatives (A and B and C). What has just been said may be expanded beyond those items merely listed, as would be apparent to one of ordinary skill in the art and related arts.
[0136] As will be apparent to one skilled in the art, the implementations can generate a variety of signals formatted to carry information that can be, for example, stored or transmitted. For example, the information can include instructions for performing a method or data generated by one of the described implementations. For example, the signals can be formatted to carry a bit stream of the described manner. For example, the signals can be formatted as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as baseband signals. For example, the formatting can include encoding a data stream and modulating a carrier wave with the encoded data stream. For example, the information the signals carry can be analog or digital information. The signals can be transmitted over a variety of separate wired or wireless links, as is known. The signals can be stored in a processor-readable medium.
[0137] Various generalized and specific aspects are also supported and contemplated throughout this disclosure. Examples of aspects according to this disclosure include, but are not limited to, the following.
[0138] In general, at least one of the aspects may include a method that includes obtaining, from a bitstream including encoded video information, information representing at least one coefficient of a quantization matrix and a syntax element; determining, based on the syntax element, that the information representing the at least one coefficient is to be interpreted as a residual; and decoding at least a portion of the encoded video information based on a combination of a prediction of the quantization matrix and the residual.
[0139] In general, at least one example of an embodiment may include an apparatus that includes one or more processors configured to obtain information representing at least one coefficient of a quantization matrix and a syntax element from a bitstream that includes encoded video information, determine based on the syntax element that the information representing the at least one coefficient is to be interpreted as a residual, and decode at least a portion of the encoded video information based on a combination of a prediction of the quantization matrix and the residual.
[0140] In general, at least one example of an aspect may include a method that includes obtaining video information and information representing at least one coefficient of a predicted quantization matrix associated with at least a portion of the video information, determining that the at least one coefficient is to be interpreted as a residual, encoding at least a portion of the video information based on a combination of the predicted quantization matrix and the residual, and encoding a syntax element indicating that the at least one coefficient is to be interpreted as a residual.
[0141] In general, at least one example of an aspect may include an apparatus that includes one or more processors configured to obtain video information and information representing at least one coefficient of a predicted quantization matrix associated with at least a portion of the video information, determine that at least one coefficient is to be interpreted as a residual, encode at least a portion of the video information based on a combination of the predicted quantization matrix and the residual, and encode a syntax element indicating that the at least one coefficient is to be interpreted as a residual.
[0142] In general, at least one example of an embodiment may include a method or apparatus comprising a residual as described herein, where the residual may comprise a variable length residual.
[0143] In general, at least one example of an aspect may include a method or apparatus described herein, wherein the at least one coefficient includes a plurality of predicted quantization matrix coefficients, and the combining includes adding a residual to the plurality of predicted quantization matrix coefficients.
[0144] In general, at least one example of the aspects can include a method or apparatus described herein, where the combination further includes applying a scale factor.
[0145] In general, at least one example of an aspect may include a method or apparatus described herein, where the combining includes applying a scale factor to a plurality of predicted quantization matrix coefficients and then adding a residual.
[0146] In general, at least one example of the aspects can include a method or apparatus described herein, where the combination further includes applying an offset.
[0147] In general, at least one example of an aspect may include a method or apparatus described herein, wherein the bitstream includes one or more syntax elements used to transmit at least one coefficient when the quantization matrix is not predicted that can be reused to transmit a residual when the quantization matrix is predicted.
[0148] In general, at least one example of an aspect may include a method or apparatus described herein that generates a bitstream including one or more syntax elements used to transmit at least one coefficient when a quantization matrix is not predicted that can be reused to transmit a residual when a quantization matrix is predicted.
[0149] In general, one or more other examples of the aspects may also provide a computer-readable recording medium, e.g., a non-transitory computer-readable recording medium, having instructions stored thereon for encoding or decoding picture information, e.g., video data, etc., in accordance with a method or apparatus described herein. Additionally, one or more aspects may also provide a computer-readable recording medium having a bitstream stored thereon generated in accordance with a method or apparatus described herein. Additionally, one or more aspects may provide methods and apparatus for transmitting or receiving a bitstream generated in accordance with a method or apparatus described herein.
[0150] In general, another example of an embodiment can include a signal including data generated according to any of the methods described herein. In general, another example of an aspect may include a signal or bitstream that includes at least one syntax element that is used to represent at least one coefficient of a quantization matrix when the quantization matrix is not predicted, and the at least one syntax element is used to represent a residual when the quantization matrix is predicted.
[0151] In general, another example of an embodiment can include a device including an apparatus according to any example of the embodiments described herein and at least one of: (i) an antenna configured to receive a signal, the signal including data representing image information; (ii) a band limiter configured to limit the received signal to a band of frequencies including the data representing the image information; and (iii) a display configured to display an image from the image information.
[0152] In general, another example of an embodiment may include a device described herein, such as a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a mobile phone, a tablet, or other electronic device.
[0153] Various aspects are described herein. The features of the just described aspects may be provided alone or in any combination across the various claim categories and types. In addition, the aspects may include one or more of the following features, devices, or aspects alone or in any combination across the various claim categories and types. ● Providing video encoding and / or decoding, including saving bits in the transmission of custom QMs, while keeping the specifications as concise as possible; Providing video encoding and / or decoding that includes improving QM prediction so that simple adjustments are possible at low bit cost; Providing a video encoding and / or decoding that includes a simple way to specify a global offset; Providing video encoding and / or decoding that includes one or more of applying a scale factor in addition to copying and downsampling, providing a simple way to specify a global offset, and extending the prediction of QM by allowing the prediction to be refined by a variable length residual; Providing a video encoding and / or decoding that includes combining an improved prediction with a residual to further refine the QM, with a potentially lower cost than a full encoding; Providing video encoding and / or decoding including QM prediction with offset in addition to copy / downscaling; Providing video encoding and / or decoding including QM prediction with (variable length) residuals; Providing video encoding and / or decoding including QM prediction using scale factors in addition to copy / downscaling; Providing video encoding and / or decoding including QM prediction including a combination of a scale factor with either an offset or a residual; Providing video encoding and / or decoding that includes reducing the number of bits required to transmit the QM and allowing a user-defined tradeoff between accuracy and bit cost while keeping specifications concise; Providing video encoding and / or decoding that includes reducing the number of bits required to transmit QM and to allow a user-defined tradeoff between accuracy and bit cost while keeping specifications concise, based on including one or more of the following features: Add residuals to the QM predictions, Adding a scaling factor to the QM predictions, Combine residuals and scaling, ○ Always use prediction (remove prediction mode flag), providing video encoding and / or decoding including modifications to a copy of the QM prediction including adding a global offset, the global offset being able to be explicitly specified; Providing a video encoding and / or decoding including a global offset that can be specified as a first coefficient of a DPCM coded residual; Providing video encoding and / or decoding including adding residual to prediction; Providing video encoding and / or decoding including residual on top of QM prediction, refining prediction or providing an optimized coding scheme for QM (prediction on top of residual instead of direct coding); Providing a video encoding and / or decoding method comprising: providing ... Providing a video encoding and / or decoding including adding a residual, the number of coefficients of the residual being variable; providing a video encoding and / or decoding method including adding a residual, which may include using a flag when in a predictive mode to indicate that the coefficients are still transmitted as a non-predictive mode but can be interpreted as a residual; Providing a video encoding and / or decoding including, when not in prediction mode, coefficients can still be interpreted as residuals in addition to the default QM prediction; Providing video encoding and / or decoding that includes using variable length residuals in addition to prediction; providing video encoding and / or decoding including using a variable length residual in addition to prediction, the number of residual coefficients being capable of being indicated by one or more of the following: ○ An explicit number of coefficients, or a flag indicating the presence of a residual followed by the number of residual coefficients minus one, which can be restricted to a power of two, or ○ 1 + (log2 of the number of residual coefficients), where zero indicates no residual, which can be coded as a fixed length (with length possibly depending on matrixId) or an exp-Golomb coding, or o an index in a list, where the list contains the number of coefficients, for example this list can be fixed by the standard and does not need to be transmitted, e.g. {0,1,2,4,8,16,32,64}, or The number of residual coefficients does not need to count the potentially extra DC residual coefficient that can be transmitted at the same time as at least one normal residual coefficient is sent. Providing video encoding and / or decoding including adding a scale factor to prediction, e.g. adding a scale factor to copying / decimation, Providing video encoding and / or decoding including QM prediction that includes a variable length residual and a scale factor, thereby allowing adjustments to the QM based on scaling and / or offsets, and allowing a user-defined tradeoff between accuracy and bit cost based on the variable length residual; providing video encoding and / or decoding including a QM prediction that allows for adjustments to the QM based on scaling and / or offsets by including a variable length residual and a scale factor, and allows for a user-defined tradeoff between accuracy and bit cost based on the variable length residual, where the scale factor can be an integer normalized to a number N (e.g., 16), and the scaling equation scales a reference QM based on a neutral value; providing a video encoding and / or decoding including adding a scale factor to the prediction, the scale factor being an integer normalized to a number N, the scale factor being transmitted as an offset to N; Providing video encoding and / or decoding including removing prediction mode flags; A bitstream or signal containing one or more of the described syntax elements or variations thereof; A bitstream or signal including a syntax conveying information generated according to any of the described aspects; Inserting into the signaling syntax elements syntax elements that enable the decoder to operate in a manner corresponding to the syntax elements used by the encoder; Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations; Creating and / or transmitting and / or receiving and / or decoding according to any of the described aspects; A method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described aspects; A television, set-top box, mobile phone, tablet, or other electronic device provided to implement any of the described aspects. ● A television, set-top box, mobile phone, tablet, or other electronic device provided for implementing any of the described aspects and displaying the resulting images (e.g., using a monitor, screen, or other type of display). ●A television, set-top box, mobile phone, tablet, or other electronic device that is provided for selecting a channel (e.g., using a tuner) to receive a signal containing an encoded image and implementing any of the described aspects. ●A television, set-top box, mobile phone, tablet, or other electronic device configured to receive a signal over the airwaves (e.g., using an antenna) containing an encoded image and implement any of the described aspects.
Claims
1. obtaining a syntax element from a bitstream containing encoded video information, and obtaining information representative of at least one coefficient of a quantization matrix based on the syntax element; decoding a flag, the flag having a first value specifying that coefficients of the quantization matrix are explicitly signaled in the bitstream or a second value specifying that the quantization matrix is derived from a reference quantization matrix, and when the flag has the first value, all elements of the reference quantization matrix are inferred to be equal to 8; interpreting the information representing the at least one coefficient as a residual based on the syntax element and the flag; decoding at least a portion of the encoded video information based on a combination of the quantization matrix prediction and the residual; and 23. A method comprising:
2. obtaining at least one coefficient of a quantization matrix and a syntax element from a bitstream containing encoded video information; interpreting the information representing the at least one coefficient as a residual based on the syntax element; Decoding at least a portion of the encoded video information based on a combination of the prediction of the quantization matrix and the residual. One or more processors configured to An apparatus comprising:
3. obtaining information representative of at least one coefficient of a quantization matrix associated with at least a portion of the video information; encoding a syntax element indicating that the at least one coefficient is to be interpreted as a residual; encoding the at least a portion of the video information based on a combination of the prediction of the quantization matrix and the residual; 23. A method comprising:
4. obtaining information representative of at least one coefficient of a quantization matrix associated with at least a portion of the video information; encoding a syntax element indicating that the at least one coefficient is to be interpreted as a residual; encoding the at least a portion of the video information based on a combination of the prediction of the quantization matrix and the residual; One or more processors configured to An apparatus comprising:
5. 5. A method according to claim 1 or 3 or an apparatus according to claim 2 or 4, wherein the residual comprises a variable length residual.
6. 6. The method of claim 5, wherein when the quantization matrix is in a prediction mode in which the quantization matrix is predicted from a reference quantization matrix, the variable length of the residual depends on the syntax element.
7. The apparatus of claim 5 , wherein when the quantization matrix is in a prediction mode in which the quantization matrix is predicted from a reference quantization matrix, the variable length of the residual depends on the syntax element.
8. The method of claim 1, characterized in that when the quantization matrix is in a prediction mode in which the quantization matrix is predicted from a reference quantization matrix, the number of pieces of information representing at least one coefficient obtained from the bitstream depends on the syntax element.
9. 2. The method of claim 1, wherein the quantization matrix is obtained by adding elements of a reference quantization matrix to elements of the residual.
10. 10. The method of claim 9, wherein when the quantization matrix is in a prediction mode in which the quantization matrix is predicted from the reference quantization matrix, for a given value for the syntax element, all elements of the residual are inferred to be equal to 0.
11. The method of claim 1 or the device of claim 2, characterized in that the bitstream is obtained by decoding one or more syntax elements, the information representing at least one coefficient of the quantization matrix being obtained by decoding one or more syntax elements, the same one or more syntax elements being used to transmit the at least one coefficient when the quantization matrix is not predicted and to transmit the residual when the quantization matrix is predicted.
12. 5. The method of claim 3 or the apparatus of claim 4, wherein the encoding or at least one of the processors configured to encode generates a bitstream including one or more syntax elements signaling the information representing at least one coefficient of the quantization matrix, the same one or more syntax elements being used to transmit the at least one coefficient when the quantization matrix is not predicted and to transmit the residual when the quantization matrix is predicted.
13. 2. The method of claim 1, comprising obtaining another syntax element that indicates the reference quantization matrix.
14. 14. The method of claim 13, wherein if the other syntax element is 0, then all elements of the reference quantization matrix are inferred to be equal to 16.
15. 14. The method of claim 13, wherein if the other syntax element is non-zero, the reference quantization matrix is derived from a previously transmitted matrix identified by the other syntax element.
16. A non-transitory computer readable medium storing executable program instructions that cause a computer executing said instructions to perform the method according to claim 1 or 3.
17. 4. A non-transitory computer readable medium storing a bitstream formatted to include syntax elements and encoded image information according to the method of claim 3.
18. 20. The non-transitory computer-readable medium of claim 17, further comprising one or more syntax elements that are used to transmit the at least one coefficient of the quantization matrix when the quantization matrix is not predicted, the same one or more syntax elements being used to transmit the residual when the quantization matrix is predicted.
Citation Information
Patent Citations
Image encoder, image encoding method, program, image decoder, image decoding method and program
JP2013038758A