Secondary conversion of joint components
JCST addresses the high correlation between chroma channels in AV1 by applying a joint component second transform to reduce redundancy in residual coding, thereby improving encoding and decoding efficiency.
Patent Information
- Application Number
- JP2024131004
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-16
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-11-02
AI Technical Summary
The high correlation between chroma channels in AV1 video encoding leads to statistical redundancy in residual coding, which can be improved by reducing the redundancy between prediction residuals of Cb and Cr components.
Implementing a joint component second transform (JCST) technique for encoding residuals from multiple color components, which includes obtaining an encoded video bitstream, performing entropy analysis, dequantizing color components, applying JCST to the transform coefficients, and obtaining residual components for decoding.
JCST reduces statistical redundancy, enhancing the efficiency of video encoding and decoding by improving compression performance.
Smart Images

Figure 0007701527000002 
Figure 0007701527000003 
Figure 0007701527000004
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 020,280, filed May 5, 2020, and U.S. Patent Application No. 17 / 072,606, filed Oct. 16, 2020, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure generally relates to the field of data processing, and more specifically, to video encoding and decoding. Even more specifically, embodiments of the present disclosure relate to a joint components secondary transform (JCST) technique for encoding residuals from multiple color components, e.g., residuals from two chroma components.
Background Art
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. As a successor to VP9, AV1 was established in 2015 and developed by the Alliance for Open Media (AOMedia), which consists of semiconductor companies, video - on - demand providers, video content producers, software development companies, and web browser vendors. Many components of the AV1 project were supplied from previous research activities by AOMedia members. And individual contributors had started experimental technology platforms years ago. For example, Xiph’s / Mozilla’s Daala made its code public in 2010. Google’s experimental VP9 evolution project VP10 was announced on September 12, 2014. Cisco’s Thor was made public on August 11, 2015.
[0004] AV1, which is built based on the VP9 codebase, incorporates other technologies, some of which were developed in these experimental forms. The first version (0.1.0) of the AV1 reference codec was released on April 7, 2016. AOMedia announced the release of the AV1 bitstream specification on March 28, 2018, along with reference software-based encoders and decoders. On June 25, 2018, the verified version 1.0.0 of the AV1 specification was released. On January 8, 2019, the verified version 1.0.0 of the AV1 specification was released along with Errata 1. The AV1 bitstream specification includes a reference video codec. Summary of the Invention Means for Solving the Problems
[0005] In AV1, since the prediction residual signals generated for the chroma channels, such as Cb and Cr, have a high correlation with each other, it is expected that the coding of the residuals can be further improved by reducing the statistical redundancy between the prediction residuals of Cb and Cr.
[0006] Embodiments of the present disclosure provide solutions to the above problems.
[0007] For example, a method of decoding an encoded video bitstream using at least one processor, the method including: obtaining an encoded video bitstream including encoded color components; performing entropy analysis on the encoded color components; dequantizing the color components to obtain transform coefficients of the color components; generating a JCST output by applying joint component second transform (JCST) to the transform coefficients of the color components; obtaining a residual component of the color components by performing an inverse transform on the JCST output; and decoding the encoded video bitstream based on the residual component of the color components. Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0009] The embodiments described in this specification provide methods and apparatuses for encoding and / or decoding image data.
[0010] Residual encoding by AV1
[0011] For each defined transformation unit, the AV1 coefficient encoder starts with a skip symbol, followed by the transformation kernel type, and if transform coding is not skipped, encodes the end-of-block (EOB) position of all non-zero coefficient blocks. Next, each coefficient value can be mapped to a plurality of hierarchical diagrams and symbols. Here, the symbol layer covers the symbol of the coefficient and three hierarchies, and each coefficient value corresponds to different hierarchical ranges of the coefficient, namely the lower layer, the middle layer, and the upper layer. The lower layer corresponds to the range of 0 to 2, the middle layer corresponds to the range of 3 to 14, and the upper layer corresponds to the range of 15 or more.
[0012] After the EOB position is encoded, the lower layer indicating whether the magnitude of the coefficient is between 0 and 2 and the middle layer indicating whether its magnitude is between 3 and 14 are encoded together in reverse scan order, and then the symbol layer and the upper layer indicating the residual value with a magnitude of 14 or more are encoded together in forward scan order. The remaining part is entropy-coded with Exp-Golomb code. In AV1, the conventional zigzag scan order is adopted.
[0013] Such separation enables the assignment of rich context models to the lower layer. Among them, up to five adjacent coefficients for improving compression efficiency in conversion directions such as bidirectional, horizontal, and vertical, conversion sizes, and appropriate context model sizes are considered. In the middle layer, a context model similar to the lower layer is used, and the number of context adjacent coefficients therein decreases from 5 to 2. In the upper layer, it is encoded using Exp-Golomb code without using a context model. In the symbol layer, symbols except for the DC symbol are encoded using the DC symbol of the adjacent transformation unit as context information. Other symbol bits are encoded as they are without using a context model.
[0014] In Versatile Video Coding (VVC), a coding block is first divided into sub-blocks of size 4×4, and the sub-blocks within the coding block and the transform coefficients within the sub-blocks are encoded according to a predefined scan order. In the case of a sub-block having at least one non-zero transform coefficient, the coding of the transform coefficients is divided into four scan paths.
[0015] For example, assume that absLevel is the absolute value of the current transform coefficient. In the first path, the syntax elements sig_coeff_flag (indicating that absLevel is greater than 0), par_level_flag (indicating the parity of absLevel), and rem_abs_gt1_flag (indicating that (absLevel - 1) >> 1 is greater than 0) are encoded; in the second path, the syntax element rem_abs_gt2_flag (indicating that absLevel is greater than 4) is encoded; in the third path, the remaining value of the coefficient level (referred to as abs_remainder) is called; and if necessary, in the fourth path, the sign information is encoded.
[0016] To utilize the correlation between transform coefficients, the encoded coefficients covered by the local template shown in FIG. 1 are used for the context selection of the current coefficient. In the figure, the position (101) shown in black indicates the position of the current transform coefficient. The positions (102) shown in shaded areas indicate its five adjacent coefficients. Here, absLevel1[x][y] represents the partially reconstructed absolute level of the coefficient at position (x, y) after the first path, d represents the diagonal position of the current coefficient (d = x + y), numSig represents the number of non-zero coefficients in the local template, and sumAbs1 represents the sum of the partially reconstructed absolute levels absLevel1[x][y] of the coefficients covered by the local template.
[0017] When encoding the current coefficient sig_coeff_flag, the context model index is selected according to sumAbs1 and the diagonal position d. More specifically, for the Luma component, the context model index is determined according to the following formula 1: ctxSig = 18 * max(0, state - 1) + min(sumAbs1, 5 + (d < 2? 12 : (d < 5? 6 : 0))). This is equivalent to the following formulas 2 and 3. Formula 2: ctxIdBase = 18 * max(0, state - 1 + (d < 2? 12 : (d < 5? 6 : 0)). Formula 3: ctxSig = ctxIdSigTable[min(sumAbs1, 5)] + ctxIdBase.
[0018] For Chroma, the context model index is determined according to the following formula 4: ctxSig = 12 * max(0, state - 1 + min(sumAbs1, 5 + (d < 2? 6 : 0)). This is equivalent to the following formulas 5 and 6. Formula 5: ctxIdBase = 12 * max(0, state - 1 + (d < 2? 6 : 0)). Formula 6: ctxSig = ctxIdSigTable[min(sumAbs1, 5)] + ctxIdBase.
[0019] Here, when dependent quantization is used and the state is derived using the state conversion process, a scalar quantizer is used. The table ctxIdSigTable stores the context model index offsets, ctxIdSigTable[0~5] = {0, 1, 2, 3, 4, 5}.
[0020] When encoding the par_level_flag which is the current coefficient, the context model index is selected according to sumAbs1, numSig and the diagonal position d. More specifically, for the Luma component, the context model index is determined according to the following formula 7: ctxPar = 1 + min(sumAbs1 - numSig, 4) + (d == 0? 15 : (d < 3? 10 : (d < 10? 5 : 0))). This is equivalent to the following formulas 8 and 9. Formula 8: ctxIdBase = (d == 0? 15 : (d < 3? 10 : (d < 10? 5 : 0))). Formula 9: ctxPar = 1 + ctxIdTable[min(sumAbs1 - numSig, 4)] + ctxIdBase. For chroma, the context model index is determined according to the following formula 10: ctxPar = 1 + min(sumAbs1 - numSig, 4) + (d == 0? 5 : 0). This is equivalent to the following formulas 11 and 12. Formula 11: ctxIdBase = (d == 0? 5 : 0). Formula 12: ctxPar = 1 + ctxIdTable[min(sumAbs1 - numSig, 4)] + ctxIdBase.
[0021] Here, the table ctxIdTable stores the context model index offsets, ctxIdTable[0~4] = {0, 1, 2, 3, 4}. When encoding the rem_abs_gt1_flag and rem_abs_gt2_flag which are the current coefficients, the context model index is determined in the same way as in the case of par_level_flag. ctxGt1 = ctxPar and ctxGt2 = ctxPar (Formula 13).
[0022] Different context model sets are used for rem_abs_gt1_flag and rem_abs_gt2_flag. Therefore, even when ctxGt1 is equal to ctxGt2, the context model used for rem_abs_gt1_flag is different from the context model of rem_abs_gt2_flag.
[0023] Residual Encoding for Transform Skip Mode (TSM) and Differential Pulse-Code Modulation (DPCM)
[0024] In addition to the residual encoding scheme described in the residual encoding part of AV1 above, in order to adapt the residual encoding to the statistics and signal characteristics of the residual levels of transform skip and Block Differential Pulse-Code Modulation (BDPCM) representing the quantized prediction residuals (spatial domain), it is proposed to modify the following residual encoding process and apply it to the TSM and BDPCM modes.
[0025] Three encoding paths are described below. In the first encoding path, sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag are first encoded in one path. In the second encoding path, abs_level_gtX_flag is encoded. Here, X can be 3, 5, 7,.... In the third path, the remaining coefficient levels are encoded. The encoding paths operate at the Coefficient Group (CG) level, that is, for each CG, three encoding paths are executed.
[0026] There is no valid scan position. Since the residual signal reflects the spatial residual after prediction and no energy compression by transformation is performed on the TS, there is no high probability of trailing zeros or meaningless levels occurring at the lower right corner of the transform block. Therefore, in this case, the signaling of the last valid scan position is omitted. Instead, the first sub-block to be processed is the sub-block at the lower right corner within the transform block.
[0027] Next, the Coding Block Flag (CBF) of the sub-block will be described. When there is no last valid scan position signaling, the sub-block CBF signaling with coded_sub_block_flag in the TS needs to be changed as follows.
[0028] Due to quantization, the aforementioned meaningless sequences may partially occur within the transform block. Therefore, the last valid scan position is deleted as described above, and coded_sub_block_flag is encoded for all sub-blocks.
[0029] The coded_sub_block_flag of the sub-block covering the DC frequency position (the top-left sub-block) is a special case. In VVC Draft 3, the coded_sub_block_flag of this sub-block is not signaled and is assumed to be equal to 1. When the last valid scan position is in another sub-block, there is at least one valid level other than the DC sub-block. Therefore, the coded_sub_block_flag of this sub-block is assumed to be equal to 1, but the DC sub-block may only contain zero / invalid levels. Since there is no last scan position information in the TS, the coded_sub_block_flag of each sub-block is signaled. This includes the coded_sub_block_flag of the DC sub-block, except when all other coded_sub_block_flag syntax elements are already equal to 0. In this case, the DC coded_sub_block_flag is assumed to be equal to 1 (inferDcSbCbf = 1). Since there is at least one valid level in this DC sub-block, the sig_coeff_flag syntax element at the first position at (0,0) is not signaled, and it is derived that when all other sig_coeff_flag syntax elements within the DC sub-block are equal to 0, it becomes equal to 1 (inferSbDcSigCoeffFlag = 1).
[0030] The context modeling of coded_sub_block_flag can be changed. The context model index can be calculated as the sum of the coded_sub_block_flag on the right side and the coded_sub_block_flag below the current sub-block, but it is not the logical OR of the two.
[0031] The sig_coeff_flag context modeling is described below. The local template within the sig_coeff_flag context modeling can be changed to include only the neighbor (NB0) on the right side of the current scan position and the neighbor (NB1) below it. The offset of the context model represents the number of valid adjacent positions sig_coeff_flag[NB0] + sig_coeff_flag[NB1]. Therefore, the selection of different contexts is set according to the diagonal d within the current transform block (d is removed). This creates three context models and a single context model set for coding the sig_coeff_flag flag.
[0032] The abs_level_gt1_flag and par_level_flag context modeling are described below. A single context model is adopted for abs_level_gt1_flag and par_level_flag.
[0033] The abs_remainder coding is described below. The empirical distribution of the residual absolute level of the transform skip usually conforms to the Laplacian distribution or the geometric distribution, but there may be greater instability than the absolute level of the transform coefficient. In particular, for the residual absolute level, the variance within the window of consecutive realizations is high. This allows the binarization and context modeling of the abs_remainder syntax to be changed as follows.
[0034] Use a higher cutoff value for binarization. That is, the transition point from the coding using sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag to the Rice code of abs_remainder, and the use of a dedicated context model for each bin position result in higher compression efficiency. Increasing the cutoff will generate more "greater than X" flags, such as introducing abs_level_gt5_flag, abs_level_gt7_flag, etc. until the cutoff is reached. The cutoff itself is fixed at 5 (numGtFlags = 5).
[0035] Similar to the local template of the sig_coeff_flag modeling context, the template for deriving the rice parameter can be changed. That is, only the neighbors on the left and below the current scan position are considered.
[0036] The following explains the context modeling of the coeff_sign_flag. Due to the instability within the sign sequence and the fact that the prediction residual is often biased, even when the global empirical distribution is almost uniformly distributed, symbols can be encoded using a context model. A single dedicated context model can be used for symbol encoding, and after sig_coeff_flag, the symbols can be parsed and all the binary numbers encoded in the context can be grouped together.
[0037] The following describes the limit on the number of context - encoded binary numbers. The total number of context - encoded binary numbers for each transform unit (TU) is limited to twice the TU area size. For example, for a 16×8 TU, the maximum number of context - encoded binary numbers is 16×8×2 = 256. The budget for the context - encoded binary numbers is consumed at the TU level. That is, instead of individual budgets for the number of context - encoded binary numbers for each CG, all the CGs within the current TU share one budget for the number of context - encoded binary numbers.
[0038] Joint encoding of chroma residuals
[0039] VVC Draft 6 supports a mode that jointly encodes chroma residuals. The use (activation) of the joint chroma encoding mode is indicated by the TU-level flag tu_joint_cbcr_residual_flag, and the selected mode is implicitly indicated by the chroma CBF. The flag tu_joint_cbcr_residual_flag is considered to exist if either or both of the chroma CBFs of the TU are equal to 1. In the PPS and slice headers, a chroma quantization parameter (QP) offset value is signaled for the joint chroma residual encoding mode, to be distinguished from the chroma QP offset value signaled for the normal chroma residual encoding mode. These chroma QP offset values are used to derive the chroma QP values of the blocks encoded by the joint chroma residual encoding mode. When the corresponding joint chroma encoding mode (mode 2 in Table 1) is active in a TU, during the quantization and decoding periods of that TU, a chroma QP offset is added to the applied luma to obtain the chroma QP. For the other modes (modes 1 and 3 in Table 1), the chroma QP is derived in the same way as for conventional Cb or Cr blocks. Table 1 shows the reconstruction process of the chroma residuals (resCb and resCr) from the transmitted transform blocks. When this mode is active, a single joint chroma residual block (resJointC[x][y] in Table 1) is signaled, and the residual blocks for Cb (resCb) and Cr (resCr) are derived considering information such as tu_cbf_cb, tu_cbf_cr, and CSign (the sign value specified in the slice header).
[0040] The above three joint chroma coding modes are supported only in the intra-coded coding unit (CU). In the inter-coded CU, only mode 2 is supported. Therefore, for the inter-coded CU, the syntax element tu_joint_cbcr_residual_flag exists only when both chroma cbfs are equal to 1.
[0041]
Table 1
[0042] Here, the value CSign is the symbol value (+1 or -1) specified in the slice header, and resJointC[][] is the transmitted residual.
[0043] Referring now to FIG. 2, FIG. 2 is a block diagram of a communication system 200 according to an embodiment. The communication system 200 may include at least two terminals 210 and 220 interconnected via a network 250. When transmitting data in one direction, the first terminal 210 can encode the data at a local location for transmission to the second terminal 220 via the network 250. The second terminal 220 can receive the encoded data of the first terminal 210 from the network 250, decode the encoded data, and display the decoded data. One-way transmission of data may be common in media serving applications and the like.
[0044] FIG. 2 further shows terminals 230 and 240 provided to support, for example, two-way transmission of encoded data that may occur during a video conference. In the case of two-way transmission of data, each terminal 230 or 240 can encode data captured at a local location and transmit it to other terminals via network 250. Each terminal 230 or 240 can also receive encoded data transmitted by other terminals, decode the encoded data, and display the decoded data on a local display device.
[0045] In FIG. 2, terminals 210-240 can be shown as servers, personal computers, and smartphones, but the principles of the embodiments are not limited thereto. The embodiments are suitable for laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 250 represents any number of networks that transmit encoded data among terminals 210-240, including, for example, wired and / or wireless communication networks. Communication network 250 can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network 250 may not be important for the operation of the embodiments unless otherwise described herein below.
[0046] FIG. 3 is a diagram showing an arrangement of a G-PCC compressor 303 and a G-PCC decompressor 310 in an environment according to an embodiment. The disclosed subject matter is equally applicable to other mountable applications, such as, for example, video conferencing, digital TV, storage of compressed data on digital media such as CDs, DVDs, memory sticks, etc.
[0047] The streaming system 300 can include a source 310, such as a digital camera, and may include a capture subsystem 313 that can create, for example, uncompressed data 302. The data 302 with a larger amount of data can be processed by a G-PCC compressor 303 coupled to the source 301. The G-PCC compressor 303 can include hardware, software, or a combination thereof so as to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded data 304 with a smaller amount of data can be stored in a streaming server 305 for subsequent use. One or more streaming clients 306 and 308 can access the streaming server 305 to retrieve copies 307 and 309 of the encoded data 304. The client terminal 306 can include a G-PCC decompressor 310 that decodes the incoming copy 307 of the encoded data and creates outgoing data 311 that can be displayed on a display 312 or other display device (not shown). In some streaming systems, the encoded data 304, 307, and 309 can be encoded according to video coding / compression standards. Examples of these standards include those developed by MPEG for G-PCC.
[0048] Embodiments of the present disclosure can jointly apply a second-order transformation to the transformation coefficients of a plurality of color components. The joint transformation scheme proposed here can be called joint component second-order transformation (JCST). In FIG. 4, an encoder scheme for applying JCST to two color components is shown. Here, JCST is performed after the forward transformation and before quantization.
[0049] Embodiments of the present disclosure can perform JCST after inverse quantization transformation and before inverse (reverse) transformation, as shown in FIG. 5.
[0050] Referring to FIG. 9, LRRX is in the first block 901. Method 900 includes the step of obtaining an encoded video bitstream including encoded color components.
[0051] In the second block 902, method 900 includes the step of performing entropy analysis on the encoded color components.
[0052] In the third block 903, method 900 includes the steps of dequantizing the color components and obtaining the transform coefficients of the color components.
[0053] In the fourth block 904, method 900 includes the step of generating a JCST output by applying joint component second transform (JCST) to the transform coefficients of the color components.
[0054] According to an embodiment, a fifth block 905 may be provided. In the fifth block 905, method 900 may include the step of obtaining the residual components of the color components by performing an inverse transform on the JCST output.
[0055] According to an embodiment, a sixth block 906 may be provided. In the sixth block 906, method 900 may include the step of decoding the encoded video bitstream based on the residual components of the color components.
[0056] According to an embodiment, this method may be run in reverse as an encoding method. In fact, although this specification may refer to specific encoding or decoding schemes, these descriptions are not limited to specific encoding or decoding schemes. That is, it is equally applicable to both the encoding scheme and the decoding scheme.
[0057] In one embodiment, the input of JCST may be the Cb and Cr transform coefficients.
[0058] In another embodiment, the input of JCST may be the Y, Cb, and Cr transform coefficients.
[0059] In one embodiment, JCST may be performed on a per-element basis such that JCST is performed for each pair of Cb and Cr conversion coefficients located at the same coordinates. FIG. 6 illustrates pairs of Cb and Cr conversion coefficients derived from two 4×2-sized blocks.
[0060] In one embodiment, JCST can be a two-point transform, and the input can be a pair of Cb and Cr conversion coefficients located at the same coordinates.
[0061] In one embodiment, JCST can be a two-point transform, and the output can be a pair of secondary conversion coefficients that replace the pair of Cb and Cr conversion coefficients.
[0062] In one embodiment, the output pair of Cb and Cr conversion coefficients can be arranged at the same position as the pair of Cb and Cr conversion coefficients used as the input to JCST. FIG. 7 illustrates JCST applied to two 4×2-sized Cb and Cr blocks. The output of JCST that constructs yet another two 4×2-sized Cb and Cr blocks is further quantized / dequantized and entropy encoded / analyzed in an encoder / decoder.
[0063] In one embodiment, the output of JCST may be less than the input. For example, the input can be a pair of Cb and Cr coefficients, while the output can be only one secondary conversion coefficient.
[0064] In one embodiment, the transform applied to JCST can include, but is not necessarily limited to, Hadamard transform, discrete cosine / sine transform, and data-driven transforms such as KLT and LGT (line graph transform).
[0065] In one embodiment, the input to JCST can be derived from one or more (e.g., three pairs) of different color components located at different coordinates.
[0066] In one embodiment, the JCST can be a four-point transform, and the input can be two pairs of Cb and Cr transform coefficients located at the same coordinates. An illustration is shown in FIG. 8.
[0067] In one embodiment, the output can be one or more pairs of secondary transform coefficients that replace one or more pairs of Cb and Cr transform coefficients.
[0068] In one embodiment, the output of the JCST can be less than the input. For example, the input can be one or more pairs of Cb and Cr coefficients, while the output can be only two secondary transform coefficients.
[0069] In one embodiment, the transform applied to the JCST can include, but is not necessarily limited to, Hadamard transform, discrete cosine / sine transform, and data-driven transforms such as KLT and LGT (line graph transform).
[0070] In one embodiment, the JCST can be applied to block sizes within a limited range.
[0071] In one example, the JCST can be applied to block sizes below a predetermined threshold. Here, the block size can refer to the block width, block height, block area size, block width and height, and the maximum (or minimum) value of the block width and height.
[0072] In one example, the JCST can be applied to block sizes above a predetermined threshold. Here, the block size can refer to the block width, block height, block area size, block width and height, and the maximum (or minimum) value of the block width and height.
[0073] In one embodiment, whether the JCST is applied can be notified by a JCST flag at the transform block level.
[0074] In one embodiment, the JCST flag can be signaled after the transform coefficients.
[0075] In one embodiment, the JCST flag may be signaled only when at least one color component to which JCST is applied has at least one non-zero coefficient.
[0076] In one embodiment, the JCST flag is signaled only when each color component to which JCST is applied has at least one non-zero coefficient.
[0077] In one embodiment, the JCST flag is signaled only when the total number of non-zero coefficients of the color components to which JCST is applied is greater than a predetermined threshold, for example, 0, 1, 2, 3, or 4.
[0078] In one embodiment, the JCST flag is signaled only when the last non-zero coefficient of the color components to which JCST is applied is located at a position greater than a predetermined threshold, for example, 0, 1, 2, 3, or 4, along the scanning order.
[0079] In one embodiment, whether JCST is applied or not is signaled by the JCST flag at the CU level (or CB level).
[0080] In one embodiment, whether JCST can be applied or not is signaled via high-level syntax. The high-level syntax includes, but is not limited to, the Video Parameter Set (VPS), Sequence Parameter Setting (SPS), Picture Parameter Set (PPS), slice header, picture header, tile, tile group, or Coding Tree Unit (CTU) level.
[0081] In one embodiment, when JCST is applied, the primary transformation is a predetermined transformation type.
[0082] In one example, the predetermined conversion type is DCT-2. In another example, the predetermined conversion type is DST-7.
[0083] In one embodiment, the selection of the transform applied to JCST is determined by the coding information. Here, the coding information includes, but is not limited to, the primary transform type, the intra prediction mode such as the intra prediction direction / angle, the inter prediction mode, whether Intra Block Copy (IntraBC) is applied, whether the palette mode is applied, whether the DPCM mode is applied, the motion vector information (direction, magnitude), whether sub-block motion is applied, and whether warp motion (affine motion) is applied.
[0084] In one embodiment, the context used for entropy coding of the flag indicating whether JCST is applied depends on the adjacent block information from adjacent blocks. Here, the adjacent block information includes, but is not limited to, the information listed above.
[0085] The above techniques can be implemented by a video encoder and / or decoder adapted for compression / decompression. The encoder and / or decoder can be implemented in hardware, software, or any combination thereof, and if there is software, it may be stored on one or more non-transitory computer-readable media. For example, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable media.
[0086] The above technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 10 shows a computer system 900 suitable for implementing some embodiments of the present disclosure.
[0087] The computer software can be coded using any suitable machine code or computer language, and for any suitable machine code or computer language, mechanisms such as assembly, compilation, and linking can be performed to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed via interpretation, microcode execution, etc.
[0088] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0089] The components for the computer system 900 shown in FIG. 10 are exemplary and are not intended to limit the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of the components should not be construed as having dependencies or requirements related to any one or combination of the components shown in the non-limiting embodiments of the computer system 900.
[0090] The computer system 900 may include a specific human interface input device. Such a human interface input device can respond, for example, to tactile inputs (such as keystrokes, swipes, movements of a data glove, etc.), voice inputs (such as speech, clapping, etc.), visual inputs (such as gestures, etc.), and olfactory inputs (not shown) by one or more human users. Using the human interface device, it is also possible to capture some media that are not necessarily directly related to conscious human input, such as audio (speech, music, ambient sounds, etc.), images (scanned images, photographic images captured by a still image camera, etc.), and video (2D video, 3D video including stereoscopic video, etc.).
[0091] The input human interface device may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, a data glove, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of each item is depicted).
[0092] Computer system 900 may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices can include tactile output devices (e.g., a touch screen 910, a data glove, or a joystick 905 that may have tactile feedback, but there can also be tactile feedback devices that do not function as input devices). For example, such devices can include audio output devices (such as speakers 909, headphones (not shown), etc.), visual output devices (screens 910 including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, etc., regardless of whether they have a touch screen input function, and regardless of whether they have a tactile feedback function, some of which may be able to output two-dimensional visual output or three-dimensional or more output by means of a stereographic output method), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown).
[0093] Computer system 900 may also include human-accessible storage devices and associated media, such as optical media having ROM / RW 920 with media 921 such as CD / DVDs, thumb drives 922, removable hard drives or solid state drives 923, legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0094] Also, it should be understood by those skilled in the art that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.
[0095] Computer system 900 can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, terrestrial broadcast television, vehicle and industrial such as CANBus, etc. A particular network typically requires an external network interface adapter connected to a particular general-purpose data port or peripheral bus 949 (such as a USB port of computer system 900). Other networks are generally integrated into the core of computer system 900 by connecting to a system bus (such as an Ethernet interface to a PC computer system or a cellular network interface) as described below. Using any of these networks, computer system 900 can communicate with other entities. Such communication can be only unidirectional reception (such as broadcast television), only unidirectional transmission (such as to a CANbus of a particular CANbus device), or bidirectional (such as to another computer system using a local or wide area digital network). Such communication includes communication to a cloud computing environment 955. Specific protocols and protocol stacks can be used for each of the networks or network interfaces as described above.
[0096] The aforementioned human interface device, human-accessible storage device, and network interface 954 can be connected to the core 940 of the computer system 900.
[0097] The core 940 may include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, a dedicated program processing device in the form of a field programmable gate array (FPGA) 943, a hardware accelerator 944 for specific tasks, and the like. These devices can be connected via a system bus 948 together with a read-only memory (ROM) 945, a random access memory 946, an internal hard drive that is not accessible to the user, and an internal mass storage 947 such as an SSD. In some computer systems, access to the system bus 948 is in the form of one or more physical plugs, and it can be expanded by an attached CPU, GPU, etc. Peripheral devices can be connected directly to the system bus 948 of the core or via a peripheral bus 949. The architecture of the peripheral bus includes PCI, USB, etc. The graphics adapter 950 can be included in the core 940.
[0098] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute specific instructions that can configure the aforementioned computer code in combination. This computer code can be stored in the ROM945 or RAM946. The transition data can also be stored in the RAM 946. Also, the permanent data can be stored, for example, in the internal mass storage device 947. By using a cache memory that can be closely associated with one or more CPUs 941, GPUs 942, mass storage 947, ROM 945, RAM946, etc., fast storage and retrieval to any memory device can be enabled.
[0099] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be of the kind specially designed and constructed for the purposes of this disclosure, or of the kind well-known and available to those having skill in the computer software arts.
[0100] As an example, and not by way of limitation, a computer system 900 having an architecture, and in particular a core 940, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be, as introduced above, a mass storage device accessible to a user, as well as media associated with a particular storage device having a non-transitory nature such as core 940, for example, a mass storage device 947 inside the core or a ROM 945. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 940. The computer-readable media can include one or more storage devices or chips depending on particular requirements. The software can cause core 940, and in particular a processor (including a CPU, GPU, FPGA, etc.) therein, to execute a particular process or a particular portion of a particular process described herein, including limitations to data structures stored in RAM 946 and changes to these data structures in accordance with processes limited by the software. Further, or alternatively, the computer system can provide functionality as a result of logic incorporated in wired logic or circuitry (e.g., accelerator 944), which logic can operate instead of or in conjunction with software to execute a particular process or a particular portion of a particular process described herein. Optionally, such software may include logic, and conversely, such logic may include software. Optionally, such computer-readable media may include a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0101] Although the present disclosure describes several non-limiting embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art can embody the principles of the disclosure and thus devise many systems and methods that are within its spirit and scope, even though they are not explicitly shown or described herein.
[0102] ALF: Adaptive Loop Filter APS: Adaptive Parameter Set AV1: AOMedia Video 1 AV2: AOMedia Video 2 CB: Coding Block CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CU: Coding Unit, CTU: Coding Tree Unit DPCM: Differential Pulse Code Modulation DPS: Decoding Parameter Set HDR: High Dynamic Range HEVC: High Efficiency Video Coding ISP: Intra Sub-Partition JCCT: Joint Chroma Component Transform JVET: Joint Video Exploration Team LR: Loop Restoration Filter PDPC: Position-Dependent Prediction Combination PPS: Picture Parameter Set PU: Prediction Unit SDR: Standard Dynamic Range SPS: Sequence Parameter Set TSM: Transform Skip Mode TU: Transform Unit VVC: Versatile Video Coding WAIP: Wide-Angle Intra Prediction VPS: Video Parameter Set
Explanation of Signs
[0103] 944 Accelerator 948 System Bus 950 Graphics Adapter 954 Network Interface
Claims
1. 1. A method for encoding a video bitstream using at least one processor, comprising: obtaining video data comprising a plurality of video frames, including a current video frame including Cb and Cr color components; obtaining a residual block for each of the Cb and Cr color components; performing a linear transform on each of the residual blocks to obtain Cb and Cr transform blocks; applying a joint components secondary transform (JCST) element-wise to the Cb and Cr transform block to generate a JCST output, the JCST being a two-point transform performed on each pair of Cb and Cr transform values that are located at the same coordinate in the Cb and Cr transform block, such that for each pair of Cb and Cr transform values that are located at the same coordinate, when the Cr transform value is in the Cr transform block, the Cb transform value is located at the same coordinate as the Cb transform block; performing quantization on the JCST output to obtain transform coefficients; encoding the transform coefficients; transmitting the encoded transform coefficients in a video bitstream.
2. The method of claim 1, wherein the current video frame includes Y, Cb and Cr color components.
3. The method of claim 1, wherein the JCST is applied to block sizes within a limited range.
4. The method of claim 1, further comprising the step of signaling a JCST flag indicating whether the JCST should be applied to the Cb and Cr transform blocks.
5. The method described in claim 4, wherein the JCST flag is signaled at a transform block level, a coding unit level, or a coding block level.
6. The method of claim 4, wherein the JCST flags are signaled in a high-level syntax.
7. The method of claim 1, wherein the JCST includes a second transformation determined via encoding information.
8. The method of claim 1, wherein each of the Cb and Cr transform blocks is a 4x2 block.
9. The method of claim 1, wherein the JCST includes at least one of a Hadamard transform, a discrete cosine transform, a discrete sine transform, and a data-driven transform.
10. The method of claim 1, wherein the JCST is conditionally applied based on the number of non-zero values in the Cb transform block and / or the Cr transform block.
11. A computing system configured to perform a method according to any one of claims 1 to 10.
12. A computer program product for causing at least one processor to carry out a method according to any one of claims 1 to 10.