Method and device for executing joint component quadratic transformation

By introducing the frequency-dependent joint component secondary transformation (FD-JCST) method in the video compression technology, only the low-frequency and non-zero transformation coefficients in the transform coefficient block are processed, and the low encoding efficiency caused by the concentration of chrominance component energy at low frequencies is solved, and more efficient video compression is achieved.

CN120034652APending Publication Date: 2025-05-23TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264592.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-07-14
Filing Date
2021-05-12
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In video compression techniques, especially in the transformation domain, the energy transformation coefficients for chromaticity components are usually concentrated at low frequencies, resulting in unnecessary processing of all frequencies when applying joint component secondary transformation, affecting coding efficiency.

Method used

A frequency-dependent joint component quadratic transformation (FD-JCST) method is proposed. By determining whether the transform coefficient in the transform coefficient block is a low-frequency coefficient and determining whether it is a non-zero value, only the low-frequency and non-zero transformation coefficients are performed, and the execution of JCST is indicated using the correlation syntax.

Benefits of technology

Improve the encoding efficiency, and by performing secondary transformation coefficients only on low-frequency and non-zero transformation coefficients, unnecessary calculations and storage are reduced, and the performance of video compression is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034652A_ABST
    Figure CN120034652A_ABST
Patent Text Reader

Abstract

A method and apparatus for performing frequency dependent joint component quadratic transform (FD-JCST). The method comprises: obtaining a plurality of transform coefficients in a transform coefficient block; determining whether at least one of the plurality of transform coefficients is a low frequency coefficient; determining whether at least one of the plurality of transform coefficients is a low frequency coefficient, based on determining that the low frequency coefficient is a non-zero value; and based on determining that the low frequency coefficient is a non-zero value, performing a joint component quadratic transform (JCST) for the low frequency coefficient and signaling a related syntax to indicate that the JCST is performed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority information

[0002] This application claims the benefit of priority to U.S. Application No. 16 / 928,760, filed on July 14, 2020, which is hereby incorporated by reference in its entirety.

[0003] This application files a divisional application for the Chinese patent application with application number 202180034137.7, application date May 12, 2021, and invention name “Method and device for secondary transformation of frequency-dependent joint components”. Technical Field

[0004] The present disclosure relates to the field of advanced video coding, and more particularly, to a method of performing a frequency-dependent joint component secondary transform (FD-JCST) on transform coefficients of multiple color components. Background Art

[0005] In video compression technology, devices such as computers, televisions, mobile terminals, digital cameras, etc. are configured to encode and decode image and video information to compress video data and transmit the compressed video data at a faster rate through the network. Specifically, in video compression, spatial prediction (intra-frame prediction) and temporal prediction (inter-frame prediction) are used to reduce or remove redundancy in video frames. A video frame or a portion of a video frame can be divided into video blocks, which are called coding units (CU). Here, each CU can be encoded using spatial prediction relative to reference samples in adjacent blocks in the same video frame or using temporal prediction relative to reference samples in other reference video frames. Spatial prediction and temporal prediction produce a prediction block for the block to be decoded. The residual data represents the pixel difference between the original block to be decoded and the predicted block. The residual data can be transformed from the pixel domain to the transform domain to obtain the residual transform coefficient. In Versatile Video Coding (VVC) and other predecessor standards, block-based video coding uses a transform stage that skips the prediction residual due to different residual signal characteristics. An example of transform skipping in residual coding in AV1 and VVC is described below.

[0006] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. AV 1 was developed by the Alliance for Open Media (AOMedia, a consortium of semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors formed in 2015) as the successor to VP9. Many components of the AV1 project derive from previous research work by members of the alliance. Individual contributors began experimenting with the technology platform several years ago: Xiph's / Mozilla's Daala code was made public in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor was made public on August 11, 2015. Building on the VP9 code base, AV1 combines other technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was made public on April 7, 2016. The consortium announced the release of the AV1 bitstream specification along with software-based reference encoders and decoders on March 28, 2018. Verified version 1.0.0 of the specification was released on June 25, 2018. Verified version 1.0.0 of the specification with Errata 1 was released on January 8, 2019.

[0007] The AV1 bitstream specification includes the reference video codec. The specification of the AV1 standard is available at https: / / aomediacodec.github.io / av1-spec / av1-spec.pdf.

[0008] In AV1 residual decoding, for each transform unit, the AV1 coefficient decoder first decodes the skip symbol, followed by the transform kernel type and end-of-block (EOB) position of all non-zero coefficients if transform decoding is not skipped. Then, each coefficient value is mapped to multiple level maps and symbols, where the symbol plane covers the sign of the coefficient and the three level planes correspond to different ranges of coefficient values, namely the low-level plane, the mid-level plane and the high-level plane. The low-level plane corresponds to the range of 0 to 2, the mid-level plane corresponds to the range of 3 to 14, and the high-level plane covers the range of 15 and above.

[0009] After the EOB position is decoded, the low-level plane and the mid-level plane are decoded together in a reverse scanning order, the former indicating whether the coefficient value is between 0 and 2, and the latter indicating whether the range is between 3 and 14. Then, the symbol plane and the high-level plane are decoded together in a forward scanning order, the high-level plane indicates the residual value greater than 14, and the remaining part is entropy decoded using Exp-Golomb code. AV1 adopts the traditional zigzag scanning order.

[0010] Such separation enables the assignment of a rich context model to the low-level plane, which accounts for the transform directions: bidirectional, horizontal, and vertical; the transform size; and, at a modest context model size, up to five neighboring coefficients are used to improve compression efficiency. The mid-level plane uses a similar context model to the low-level plane, where the number of context neighboring coefficients is reduced from 5 to 2. The high-level plane is decoded by Exp-Golomb codes without using a context model. In the symbol plane, except for the direct current (DC) symbol, which is decoded using the DC symbol of its neighboring transform unit as context information, other symbol bits are decoded directly without using a context model.

[0011] In VVC residual decoding, the decoding block is first divided into 4×4 sub-blocks, and the sub-blocks inside the decoding block and the transform coefficients within the sub-blocks are decoded according to a predefined scanning order. For sub-blocks with at least one non-zero transform coefficient, the decoding of the transform coefficient is divided into four scanning processes (pass). Assuming that absLevel is the absolute value of the current transform coefficient, in the first process, the syntax elements sig_coeff_flag (which indicates that absLevel is greater than 0), par_level_flag (which indicates the parity of absLevel) and rem_abs_gt1_flag (which indicates that (absLevel-1)>>1 is greater than 0) are decoded. In the second process, the syntax element rem_abs_gt2_flag (which indicates that absLevel is greater than 4) is decoded. In the third process, if necessary, the residual value of the coefficient level (called abs_remainder) is decoded. In the fourth process, the symbol information is decoded.

[0012] In order to exploit the correlation between transform coefficients, the context selection of the current coefficient is used by Figure 1 , where the positions marked in black indicate the position of the current transform coefficient and the positions marked in light grey indicate its five neighbors. Here, absLevel1[x][y] represents the partially reconstructed absolute level of the coefficient at position (x, y) after the first process, d represents the diagonal position of the current coefficient (where d=x+y), numSig represents the number of non-zero coefficients in the local template, and sumAbs1 represents the sum of the partially reconstructed absolute levels absLevel1[x][y] of the coefficients covered by the local template.

[0013] When decoding the sig_coeff_flag of the current coefficient, the context model index is selected according to sumAbs1 and the diagonal position d. More specifically, for the luminance component, the context model index is determined according to the following formula:

[0014] ctxSig=18*max(0, state-)+min(sumAbs1, 5)+(d<2?12:(d<5?6:0)), which is equivalent to the following:

[0015] ctxIdBase=18*max(0,state-1)+(d<2?12:(d<5?6:0)); and

[0016] ctxSig=ctxIdSigTable[min(sumAbs1,5)]+ctxIdBase.

[0017] For chrominance, the context model index is determined according to the following formula:

[0018] ctxSig=12*max(0, state-1)+min(sumAbs1, 5)+(d<2?6:0), which is equivalent to the following:

[0019] ctxIdBase=12*max(0,state-1)+(d<2?6:0); and

[0020] ctxSig=ctxIdSigTable[min(sumAbs1,5)]+ctxIdBase.

[0021] Here, "state" specifies the scalar quantizer used when correlated quantization is enabled, and the state is obtained using a state conversion process. Table ctxIdSigTable stores context model index offsets, ctxIdSigTable[0-5]={0, 1, 2, 3, 4, 5}.

[0022] When decoding the par_level_flag of the current coefficient, the context model index is selected according to sumAbs1, numSig and the diagonal position d. More specifically, for the luminance component, the context model index is determined according to the following formula:

[0023] ctxPar=1+min(sumAbs1-numSig, 4)+(d==0?15: (d<3?10: (d<10?5:0))),

[0024] is equivalent to the following:

[0025] ctxIdBase=(d==0?15:(d<3?10:(d<10?5:0))); and

[0026] ctxPar=1+ctxIdTable[min(sumAbs1-numSig, 4)]+ctxIdBase.

[0027] For chrominance, the context model index is determined according to the following formula:

[0028] ctxPar=1+min(sumAbs1-numSig, 4)+(d==0?5:0), which is equivalent to the following:

[0029] ctxIdBase=(d==0?5:0); and

[0030] ctxPar=1+ctxIdTable[min(sumAbs1-numSig, 4)]+ctxIdBase.

[0031] Here, the table ctxIdTable stores context model index offsets, ctxIdTable[0~4]={0, 1, 2, 3, 4}.

[0032] When decoding rem_abs_gt1_flag and rem_abs_gt2_flag of the current coefficient, their context model indexes are determined in a similar way to par_level_flag:

[0033] ctxGtl = ctxPar; and

[0034] ctxGt2=ctxPar.

[0035] Different context model sets are used for rem_abs_gtl_flag and rem_abs_gt2_flag. This means that even if ctxGt1 is equal to ctxGt2, the context model used for rem_abs_gtl_flag is different from the context model used for rem_abs_gt2_flag.

[0036] The joint component secondary transform (JCST) method is to jointly perform a secondary transform on the transform coefficients of multiple color components (e.g., Cb and Cr color components). However, in the transform domain, especially for chrominance components, the energy transform coefficients are usually concentrated in low frequencies. Therefore, when applying JCST in the transform domain, it may not be necessary to apply JCST for all frequencies. Summary of the invention

[0037] According to some embodiments, a method for performing a frequency-dependent joint component secondary transform (FD-JCST) is provided, the method comprising: obtaining multiple transform coefficients in a transform coefficient block; determining whether at least one of the multiple transform coefficients is a low-frequency coefficient; if at least one of the multiple transform coefficients is determined to be a low-frequency coefficient, determining whether the low-frequency coefficient is a non-zero value; and if the low-frequency coefficient is determined to be a non-zero value, performing a joint component secondary transform (JCST) on the low-frequency coefficient and indicating with relevant syntax that the JCST is performed.

[0038] According to some embodiments, a device for performing a frequency-dependent joint component secondary transform (FD-JCST) is provided, comprising: an acquisition unit configured to enable at least one processor to obtain multiple transform coefficients in a transform coefficient block; a first determination unit configured to determine whether at least one transform coefficient among the multiple transform coefficients is a low-frequency coefficient; a second determination unit configured to determine whether the low-frequency coefficient is a non-zero value based on determining that at least one transform coefficient among the multiple transform coefficients is a low-frequency coefficient; and a processing unit configured to perform a joint component secondary transform (JCST) on the low-frequency coefficient based on determining that the low-frequency coefficient is a non-zero value and use relevant syntax to indicate that the JCST is performed.

[0039] According to some embodiments, there is provided an apparatus for performing a frequency-dependent joint component secondary transform (FD-JCST), comprising: at least one memory storing computer program code; and at least one processor configured to access the at least one memory and operate according to instructions of the computer program code to perform the above method.

[0040] According to some embodiments, a non-transitory computer-readable storage medium is provided, which stores at least one instruction, and when the at least one instruction is loaded and executed by a processor, the processor performs the above method.

[0041] According to the above-mentioned method and device disclosed in the present invention, because in the transform domain, especially for the chrominance component, the energy transform coefficients are usually concentrated in the low frequency, therefore when in the transform domain, if multiple transform coefficients in the transform coefficient block are low-frequency coefficients and have non-zero values, JCST is performed on the low-frequency coefficients, thereby improving the coding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The following description briefly introduces the accompanying drawings that illustrate example embodiments of the present disclosure. These and other aspects, features and advantages will become apparent from the following detailed description of example embodiments to be read in conjunction with the accompanying drawings, in which:

[0043] Figure 1 is a diagram showing residual decoding for transform coefficients in a local template;

[0044] Figure 2 is a schematic diagram illustrating a communication system according to some embodiments;

[0045] Figure 3 is a schematic diagram illustrating a video encoder and a video decoder in a streaming environment according to some embodiments;

[0046] Figure 4 is a block diagram illustrating a video decoder according to some embodiments;

[0047] Figure 5 is a block diagram illustrating a video encoder according to some embodiments;

[0048] Figure 6 is a flow chart illustrating a method of residual decoding using Transform Skip Mode (TSM) and Differential Pulse Code Modulation (DPCM) according to some embodiments;

[0049] Fig. 7A is a block diagram illustrating an encoder using a Joint Component Secondary Transform (JCST) for two color components according to some embodiments;

[0050] Figure 7B is a block diagram illustrating a decoder using a Joint Component Secondary Transform (JCST) for two color components according to some embodiments;

[0051] Figure 8is a flow chart illustrating a method of applying a frequency-dependent joint component secondary transform (FD-JCST) according to some embodiments;

[0052] Fig. 9 is a block diagram showing an apparatus configured to apply FD-JCST according to some embodiments; and

[0053] Fig.10 is a block diagram of a computer suitable for implementing some embodiments. DETAILED DESCRIPTION

[0054] Example embodiments are described herein in detail with reference to the accompanying drawings.

[0055] Figure 2 is a schematic diagram illustrating a communication system according to some embodiments.

[0056] The communication system 200 may include at least two terminal devices 210 and 220 interconnected via a network 250, such as a laptop computer and a desktop computer. For the one-way transmission of data, the first terminal device 210 may encode the video data at the local location to transmit to the second terminal device 220 via the network 250. The second terminal device 220 may receive the encoded video data of the first terminal via the network 250, decode the video data and display the decoded video data. One-way data transmission may be common in media service applications, etc. However, embodiments are not limited thereto. For example, the at least two terminals may include a television, a personal digital assistant (PDA), a tablet computer, an e-book reader, a digital camera, a digital recording device, a digital media player, a video game device, a video game console, a video teleconference device, a video streaming device, etc.

[0057] The communication system 200 may also include other terminal devices, such as mobile devices 230 and 240, which may also be connected via the network 250. These terminal devices 230 and 240 may be provided to support bidirectional transmission of decoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal 230 and 240 may encode video data captured at a local location for transmission to another terminal via the network 250. Each terminal may also receive encoded video data transmitted by another terminal, decode the encoded data, and display the recovered video data on a local display device.

[0058] In addition, terminals 210 to 240 may be shown as servers, personal computers, and smart phones, but embodiments are not limited thereto. Embodiments may include applications with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 250 may include any number of networks for sending and receiving encoded video data, including, for example, wired and / or wireless communication networks. Communication network 250 may exchange data in circuit switching channels and / or packet switching channels. For example, a network may include a telecommunications network, a local area network, a wide area network, and / or the Internet.

[0059] Figure 3 is a schematic diagram illustrating a video encoder and a video decoder in a streaming environment according to some embodiments.

[0060] The streaming system may include a capture device 313, which includes a video source 301, such as a digital camera. The capture device 313 can capture images or videos to create an uncompressed video sample stream 302. The sample stream 302 is depicted as a thick line to emphasize the high amount of data compared to the encoded video bitstream, and can be processed by an encoder 303 coupled to the video source 301. The encoder 303 may include hardware, software, or a combination thereof to implement or implement various aspects of the embodiments described in more detail below. The encoded video bitstream 304, which is depicted as a thin line to emphasize the lower amount of data compared to the sample stream 302, can be stored in the streaming server 305 for later use. One or more streaming clients 306 and 308 can access the streaming server 305 to retrieve copies 307 and 309 of the encoded video bitstream. The client 306 may include a video decoder 310 configured to decode an incoming copy 307 of an encoded video bitstream and create an outgoing video sample stream 311 that may be presented on a display 312 or other presentation device. In some embodiments, the video bitstreams 304, 307, and 309 may be encoded according to a particular video coding / compression standard, such as ITU-T H.265 (also known as HEVC) and the currently under development ITU-T H.266 (also known as Future Video Coding (FVC)).

[0061] Figure 4 is a block diagram illustrating a video decoder according to some embodiments.

[0062] The video decoder 310 may include a receiver 410 , a buffer memory 415 , a parser 420 , a loop filter unit 454 , an intra prediction unit 452 , a scaler / inverse transform unit 451 , a reference picture buffer 457 , and a motion compensated prediction unit 453 .

[0063] Receiver 410 receives one or more codec video sequences to be decoded. Here, receiver 410 can receive one coded video sequence at a time, wherein the decoding of each coded video sequence is independent of other coded video sequences. Coded video sequences can be received from channel 412, which can be a hardware / software link to a storage device storing coded video data. Receiver 410 can receive coded video data and other data that can be forwarded to its corresponding use entity, such as coded audio data and / or auxiliary data stream. Receiver 410 can separate coded video sequences from other data. In order to combat network jitter, a buffer memory 415 can be coupled between receiver 410 and entropy decoder / parser 420 (hereinafter referred to as "parser"). When receiver 410 receives data from a storage / forwarding device or from an isochronous network with stable bandwidth and controllability, buffer memory 415 may not be necessary. However, for optimal use in a packet network such as the Internet, buffer memory 415 may be required and may be configured to reserve a relatively large storage space or may be configured so that the size of the memory may be varied depending on the load on the network.

[0064] In addition, the receiver 410 can receive additional (redundant) data of the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder 310 to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, time, space or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0065] Parser 420 may be configured to reconstruct symbols 421 from an entropy coded video sequence. The categories of these symbols may include information for managing the operation of decoder 310 and may potentially include information for controlling rendering devices such as ( Figure 3 The control information for the rendering device may be in the form of a supplementary enhancement information (SEI (Supplementary Enhancement Information) message) or a video usability information (VUI) parameter set fragment.

[0066] The parser 420 may receive an encoded video sequence from the buffer memory 415 and parse or entropy decode the received encoded video sequence. The decoding of the encoded video sequence may be performed according to a video decoding technique or standard, and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 420 may extract a subgroup parameter set for at least one pixel subgroup in the pixel subgroup in the video decoder from the received encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (Group of Picture, GOP), a picture, a tile, a slice, a macroblock, a coding unit (Coding Unit, CU), a block, a transform unit (Transform Unit, TU), a prediction unit (Prediction Unit, PU), etc. The parser 420 may also extract information such as transform coefficients, quantizer parameter (quantizer parameter, QP) values, motion vectors, etc. from the encoded video sequence.

[0067] The parser 420 may perform a parsing operation on the encoded video sequence received from the buffer 415 to create the symbol 421. The parser 420 may receive the encoded data and selectively decode the specific symbol 421. In addition, the parser 420 may determine where to provide the symbol 421. That is, the parser 420 may determine whether to provide the symbol 421 to the loop filter unit 454, the motion compensation prediction unit 453, the scaler / inverse transform unit 451, and / or the intra prediction unit 452.

[0068] The reconstruction of the symbol 421 may involve multiple different units depending on the type of the coded video picture or a portion thereof (such as inter-frame pictures and intra-frame pictures, inter-frame blocks and intra-frame blocks, and other factors). The parser 420 determines which units are involved and controls based on the sub-group control information parsed from the coded video sequence. Such sub-group control information flow between the parser 420 and the multiple units is not drawn.

[0069] Furthermore, the decoder 310 may be conceptually subdivided into a number of functional units. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other.

[0070] The sealer / inverse transform unit 451 may be configured to receive quantized transform coefficients as symbols 421 from the parser 420 and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit 451 may output a block including sample values ​​to an aggregator 455 .

[0071] In some cases, the output samples of the sealer / inverse transform 451 may be intra-coded blocks. That is, blocks that do not use predictive information from a previously reconstructed picture but use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra prediction unit 452. In some cases, the intra picture prediction unit 452 may use the surrounding reconstructed information obtained from the current (partially reconstructed) picture 456 to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator 455 adds the predictive information generated by the intra prediction unit 452 to the output sample information provided by the sealer / inverse transform unit 451 on a sample-by-sample basis.

[0072] In other cases, the output samples of the sealer / inverse transform unit 451 may be inter-coded blocks and may be motion compensated blocks. In such cases, the motion compensated prediction unit 453 may access the reference picture memory 457 to obtain samples for prediction. After the obtained samples are motion compensated according to the symbols 421 belonging to the block, these samples may be added to the output from the sealer / inverse transform unit (referred to as residual samples or residual signals in this case) by the aggregator 455 to generate output sample information. The address within the reference picture memory from which the motion compensation unit obtains the predictive samples may be controlled by a motion vector, which can be used by the motion compensation unit in the form of a symbol 421, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​obtained from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0073] The output samples of aggregator 455 may be subjected to various loop filtering techniques in loop filter unit 454. The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video bitstream and available to loop filter unit 454 as symbols 421 from parser 420, although the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of a coded picture or coded video sequence, as well as to previously reconstructed and loop filtered sample values.

[0074] The output of the loop filter unit 454 may be a sample stream that may be output to the display 212 or may be output to the reference picture memory 457 for later use in inter picture prediction.

[0075] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 420), current reference picture 456 may become part of reference picture buffer 457, and new current picture memory may be reallocated before starting reconstruction of subsequent coded pictures.

[0076] The video decoder 310 may perform decoding operations according to a predetermined video compression technology or standard such as ITU-T H.265 or H.266 recommendations. The decoded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the decoded video sequence follows the syntax of the video compression technology or standard as specified in the video compression technology or standard and specifically as specified by the profile file therein. For compliance, it is also required that the complexity of the decoded video sequence is within the limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstructed sample rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the decoded video sequence.

[0077] Figure 5 is a block diagram illustrating a video encoder according to some embodiments.

[0078] Encoder 303 can receive video samples from video source 301, which is configured to capture video frames to be encoded by encoder 303. Here, video source 301 is not a part of the encoder. Video source 301 can provide source video sequences to encoder 303 in the form of digital video sample streams, which can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601Y CrCB and RGB...) and any sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, video source 301 can be a storage device that stores previously captured videos. In a video conferencing system, video source 301 can be a camera that captures local image information as a video sequence. Video data can be provided as multiple pictures or frames that constitute motion when viewed in sequence. Frames can be organized as a spatial array of pixels, wherein each pixel can include one or more samples depending on the sampling structure, color space, etc. Those skilled in the art will easily understand the relationship between pixels and samples.

[0079] The encoder 303 may include a controller 550 that controls the overall operation of the decoding loop. For example, in addition to other functions of the encoder, the controller 550 may control the decoding speed in the decoding loop. Specifically, the decoding loop may include a source encoder 530 and a decoder 533. The source encoder 530 may be configured to create symbols based on an input frame to be encoded and a reference frame. The decoder 533 may be configured to reconstruct the symbols to create sample data that another decoder in a remote device can decode. The reconstructed sample stream is then input to a reference picture memory 534. Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (or buffer) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference frame samples that the prediction part of the encoder "sees" are exactly the same as the sample values ​​that the decoder will "see" when using prediction during decoding. The operation of the "local" decoder 533 may be the same as the operation of the "remote" decoder 310.

[0080] The source encoder 530 may perform motion compensated predictive encoding, in which the source encoder 530 predictively encodes an input frame with reference to one or more previously encoded frames from a video sequence designated as “reference frames.” In this manner, the encoding engine 532 encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame.

[0081] The local video decoder 533 can decode the encoded video data of the frame that can be designated as the reference frame based on the symbol created by the source encoder 530. The operation of the encoding engine 532 can advantageously be a lossy process. When the encoded video data is decoded at the remote decoder, the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 533 replicates the decoding process that can be performed on the reference frame by the remote video decoder, and can cause the reconstructed reference frame to be stored in the reference picture memory 534. In this way, the encoder 303 can store copies of the reconstructed reference frames locally, which have common content (no transmission errors) with the reconstructed reference frames to be obtained by the remote video decoder.

[0082] The predictor 535 may perform a predictive search for the encoding engine 532. That is, for a new frame to be encoded, the predictor 535 may search the reference picture memory 534 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor 535 may operate on a pixel block-by-pixel block basis on a sample block basis to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor 535, the input picture may have prediction references extracted from multiple reference pictures stored in the reference picture memory 534.

[0083] The outputs of all the above functional units may be entropy encoded by an entropy encoder 545. The entropy encoder may be configured to convert the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0084] The transmitter 540 can buffer the encoded video sequence created by the entropy encoder 545 in preparation for transmission via the communication channel 560, which can be hardware / software linked to a storage device that can store the encoded video data. The transmitter 540 can merge the encoded video data from the source encoder 530 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams. The transmitter 540 can transmit additional data with the encoded video. The source encoder 530 can include such data as part of the encoded video sequence. The additional data can include time / space / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0085] The controller 550 can control the entire decoding operation of the source encoder 530, including, for example, setting parameters and subgroup parameters for encoding video data. During decoding, the controller 550 can assign a specific decoded picture type to each decoded picture, which can be applied to the corresponding picture. For example, a picture can generally be assigned one of the following frame types:

[0086] Intra picture (I picture), which may be a frame type that is encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures.

[0087] Predictive picture (P picture), which may be a frame type that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0088] Bidirectional predictive pictures (B pictures), which can be frame types that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0089] A source picture may typically be spatially subdivided into a plurality of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or a block of an I picture may be predictively coded (spatial prediction or intra prediction) with reference to a coded block of the same picture. A pixel block of a P picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A block of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0090] The video encoder 303 may perform decoding operations according to a predetermined video decoding technique or standard such as the developing ITU-T H.265 and H.266 recommendations. In the operation of the video encoder 303, the video encoder 303 may perform various compression operations, including predictive decoding operations that exploit temporal redundancy and spatial redundancy in the input video sequence. Thus, the decoded video data may conform to the syntax specified by the video decoding technique or standard used.

[0091] Figure 6is a flowchart illustrating residual decoding using Transform Skip Mode (TSM) and Differential Pulse Code Modulation (DPCM) according to some embodiments. Here, in order to adapt the residual decoding to the statistical characteristics and signal characteristics of TSM and Block DPCM (BDPCM) residual levels representing quantized prediction residuals (spatial domain), the above-mentioned Figure 1 The residual decoding process is used to apply TSM and BDPCM.

[0092] According to some embodiments, there are three decoding processes in residual decoding using TSM and BDPCM. In the first decoding process, sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag are decoded in one process (S601). In the second process, abs_level_gtX_flag is decoded, where X can be 3, 5, 7, ... N (S602). In the third process, the remaining part of the coefficient level is decoded (S603). The decoding process operates at the coefficient group (CG) level, that is, three decoding processes are performed for each CG.

[0093] In this case, there is no last significant scanning position. Since the residual signal reflects the spatial residual after prediction and no energy compression by transform is performed for the transport stream (TS), a trailing zero or an insignificant level at the bottom right corner of the transform block is no longer given a high probability. Therefore, the last significant scanning position signaling is omitted in this case. Instead, the first subblock to be processed is the bottom rightmost subblock within the transform block.

[0094] However, the absence of the last valid scan position signaling requires modification of the sub-block constant rate factor (CBF) signaling with coded_sub_block_flag of the TS. Due to quantization, the above insignificant sequences may still appear locally inside the transform block. Therefore, the last valid scan position is removed as described above, and coded_sub_block_flag is decoded for all sub-blocks.

[0095] The coded_sub_block_flag of the sub-block covering the DC frequency position (upper left sub-block) represents a special case. In VVC Draft 3, the coded_sub_block_flag of this sub-block is never signaled or always equal to 1. In the case where the last valid scan position is in another sub-block, this means that there is at least one valid level outside the DC sub-block. Therefore, the DC sub-block may contain only zero / non-valid levels, although the coded_sub_block_flag of this sub-block is equal to 1. In the case where the last scan position information is not present in the TS, the coded_sub_block_flag of each sub-block is signaled. This also includes the coded_sub_block_flag of the DC sub-block, unless all other coded_sub_block_flag syntax elements are already equal to 0. In this case, the coded_sub_block_flag of the DC sub-block is inferred to be equal to 1 (inferDcSbCbf=1). Since there must be at least one valid level in this DC sub-block, the sig_coeff_flag syntax element at the first position (0,0) is not signaled and is derived to be equal to 1 (inferSbDcSigCoeffFlag=1), and all other sig_coeff_flag syntax elements in this DC sub-block are equal to 0 otherwise.

[0096] In addition, the context modeling of coded_sub_block_flag may be changed. The context model index may be calculated as the sum of the coded_sub_block_flag to the right of the current sub-block and the coded_sub_block_flag below the current sub-block instead of a logical separation of the two.

[0097] For example, in sig_coeff_flag context modeling, the local template in sig_coeff_flag context modeling is modified to include only neighbors to the right of the current scan position (NB0) and neighbors below the current scan position (NB1). The context model offset is the number of valid neighboring positions sig_coeff_flag[NB0]+sig_coeff_flag[NB1]. Therefore, the selection of different context sets depending on the diagonal d within the current transform block is removed. This results in three context models and a single context model set for decoding the sig_coeff_flag flag.

[0098] Regarding abs_level_gt1_flag and par_level_flag context models, a single context model may be employed.

[0099] Regarding abs_remainder decoding, although the empirical distribution of the absolute levels of transform skip residuals generally still conforms to the Laplace distribution or geometric distribution, there are still instances that are larger than the absolute levels of transform coefficients. In particular, for the absolute levels of the residuals, the variance within a window of consecutive realizations is high. Therefore, the following modifications can be made to the abs_remainder syntax binarization and context modeling.

[0100] According to some embodiments, abs_remainder decoding uses a higher cutoff value in binarization, i.e., the transition point from decoding with sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag to Rice code for abs_remainder, and a dedicated context model for each bin position produces higher compression efficiency. Increasing the cutoff value may result in more "greater than X" flags, e.g., introducing abs_level_gt5_flag, abs_level_gt7_flag, etc. until the cutoff value is reached. The cutoff value itself can be fixed to 5 (e.g., numGtFlags=5).

[0101] In addition, the template used for Rice parameter derivation is modified. That is, only the neighbors to the left of the current scan position and the neighbors below the current scan position are considered similar to the local template used for sig_coeff_flag context modeling.

[0102] In coeff_sign_flag context modeling, due to the instability within the symbol sequence and the fact that the prediction residuals are usually biased, the context model can be used to decode the symbol even when the global empirical distribution is almost uniformly distributed. A single dedicated context model can be used for decoding the symbol, and the symbol can be parsed after sig_coeff_flag to keep all the context-coded bits together.

[0103] In addition, the total number of context-coded bits per TU may be limited to the TU area size multiplied by 2, e.g., the maximum number of context-coded bits for a 16×8TU is 16×8×2=256. The budget for context-coded bits is consumed at the TU level, i.e., instead of a separate budget for context-coded bits for each CG, all CGs within the current TU may share one budget for context-coded bits.

[0104] Fig. 7Ais a block diagram illustrating an encoder using a Joint Component Secondary Transform (JCST) for two color components according to some embodiments. Figure 7B is a block diagram illustrating a decoder using a Joint Component Secondary Transform (JCST) for two color components according to some embodiments.

[0105] In VVC draft 6, it supports a mode for joint coding of chroma residual. The use (or activation) of the joint chroma coding mode is indicated by the TU level flag tu_joint_cbcr_residual_flag, and the selected mode is implicitly indicated by the chroma CBF. If either or both of the chroma CBFs of the TU are equal to 1, the flag tu_joint_cbcr_residual_flag exists.

[0106] In the Picture Parameter Set (PPS) and slice header, chroma QP offset values ​​are signaled for the joint chroma residual coding mode to distinguish from the chroma QP offset values ​​signaled for the normal chroma residual coding mode. These chroma QP offset values ​​are used to derive the chroma QP values ​​for those blocks coded using the joint chroma residual coding mode.

[0107] For example, Table 1 below shows the reconstruction process of the chroma residuals (resCb and resCr) from the transmitted transform block, where the value of CSign is the sign value (+1 or -1) that can be specified in the slice header, and resJointC[][] is the transmitted residual.

[0108]

[0109] Table 1

[0110] This chroma QP offset is added to the chroma QP derived from the applied luma during quantization and decoding of the TU, when the corresponding joint chroma coding mode (mode 2 in Table 1) is activated in the TU. For other modes (modes 1 and 3 in Table 1), the chroma QP is derived in the same way as for regular Cb or Cr blocks. When this mode is activated, a single joint chroma residual block (resJointC[x][y] in Table 1) is signaled, and the residual block of Cb (resCb) and the residual block of Cr (resCr) are derived considering information such as tu_cbf_cb, tu_cbf_cr and CSign, where CSign is the symbol value specified in the slice header. The above three joint chroma coding modes are only supported in intra-coded CUs. In inter-coded CUs, only mode 2 is supported. Thus, for an inter-coded CU, the syntax element tu_joint_cbcr_residual_flag is only present when both chroma cbfs are 1.

[0111] In some embodiments, a method for performing a joint component secondary transform (JCST) is provided. Here, a secondary transform is performed jointly on transform coefficients of multiple color components (e.g., Cb and Cr color components). Fig. 7A , the encoder scheme uses JCST for two color components, where JCST is performed after the forward transform and before quantization. Here, the residual of component 0 and the residual of component 1 are input for forward transform. For example, residual components 0 and 1 can be Cb and Cr transform coefficients. As another example, residual components 0 and 1 can be Y, Cb, and Cr transform coefficients. A joint secondary transform can be performed element by element on the forward transformed residual component 0 and residual component 1, which means that JCST is performed on each pair of Cb and Cr transform coefficients located at the same coordinates. The secondary transformed residual components are then quantized and entropy encoded separately to produce a bitstream.

[0112] Reference Figure 7B , the decoder scheme uses JCST, where JCST is performed after the dequantization transform and before the inverse (or inverse) transform. Here, when the encoded bitstream is received, the decoder performs parsing on the received bitstream to separate the transform coefficients of the two color components and dequantize the corresponding color components. Then JCST and inverse transform are performed on the dequantized color components for the corresponding color components to produce a residual of component 0 and a residual of component 1.

[0113] However, in the transform domain, especially for chroma components, the energy transform coefficients are usually concentrated in low frequencies. Therefore, when performing JCST applied in the transform domain, it may not be necessary to apply JCST for all frequencies to capture the coding gain of JCST.

[0114] Figure 8 is a flow chart illustrating a method of applying a frequency-dependent joint component secondary transform (FD-JCST) according to some embodiments. In some implementations, Figure 8 One or more processing blocks of may be performed by encoder 303. In some implementations, Figure 8 One or more processing blocks may be performed by another device or another group of devices separate from or including the encoder 303 , for example, by the decoder 310 .

[0115] According to some embodiments, when residual decoding is performed using JCST, a frequency-dependent joint component secondary transform (FD-JCST) may be applied to transform coefficients of multiple color components. That is, the transform kernel used in JCST depends on the relative coordinates of the transform coefficients in the transform coefficient block, i.e., the frequency in the transform domain. In some embodiments, different kernels may be applied to the DC coefficient and the AC coefficient. Here, the DC coefficient may refer to a coefficient located at the upper left (lowest frequency) position of the transform coefficient block, and the AC coefficient may refer to any other coefficient in the transform coefficient block that is not a DC coefficient.

[0116] Reference Figure 8 , the method 800 includes obtaining transform coefficients in a transform coefficient block (S810).

[0117] Based on the transform coefficients obtained in S810, method 800 includes determining whether the transform coefficients in the transform block are low-frequency coefficients. According to some embodiments, the low-frequency coefficients may be determined based on coordinates (x, y) in the transform coefficient block, where both x and y are less than or equal to a given threshold N (e.g., where N is equal to 1, 2, 3, 4, 5, 7, 8, ..., 16, ..., and 32). In another embodiment, the low-frequency coefficients may be determined based on the first N coefficients in a scanning order (e.g., where N is equal to 1, 2, 3, 4, 5, 7, 8, ..., 16, ..., 32, ..., 64, ..., 128, ..., and 256). In another embodiment, the low-frequency coefficients may be determined based on coordinates (x, y) in the transform coefficient block, where x or y is less than or equal to a given threshold N (e.g., where N is equal to 1, 2, 3, 4, 5, 7, 8, ..., 16, ..., 32). In another embodiment, the low-frequency coefficients may be determined based on the coordinates (x, y) in the transform coefficients, where the maximum value (or minimum value) between x and y is less than or equal to a given threshold N (e.g., where N is equal to 1, 2, 3, 4, 5, 7, 8, ..., 16, ..., 32). However, the embodiment is not limited thereto, but may include other N values.

[0118] The method 800 includes: if the transform coefficient is a low-frequency coefficient (S820: yes), determining whether the low-frequency coefficient is a non-zero value. If the low-frequency coefficient is non-zero (S830: yes), performing a joint component secondary transform (JCST) on the transform coefficient and indicating that the JCST is being applied using relevant syntax (S840). Alternatively, if the low-frequency coefficient is zero (S830: no), the JCST is not applied and the relevant syntax is not used.

[0119] According to some embodiments, the relevant syntax may be a high-level syntax (HLS). Thus, the method 800 may include: signaling a transform kernel used in the JCST in the high-level syntax (HLS). For example, if the transform kernel is a 2×2 matrix, only one element of the kernel is signaled, and all other elements are derived based on the orthogonality of the transform kernel and a predefined norm.

[0120] According to another embodiment, a set of fixed transform kernels may be predefined and the index of the transform kernel being used in the set may be signaled in HLS instead of signaling the kernel element.

[0121] Furthermore, the HLS may indicate the cores used for each frequency, the cores used for each prediction mode, and / or the cores used for each associated primary transform type.

[0122] In addition, the method 800 may apply an integer transform kernel to the JCST. Here, the integer kernel may refer to a transform kernel whose elements are all integers.

[0123] Based on applying an integer transform kernel to the JCST, method 800 may also include dividing the output of the JCST by a factor N, where N may be a power of 2. Alternatively, the output of the JCST may be clipped to a given data range such as [a, b]. For example, the values ​​of [a, b] may include [-2 15 ,2 15 -1] or [-2 M ,2 M -1], where M depends on the internal bit depth.

[0124] although Figure 8 Example blocks of method 800 are shown, but in some implementations, Figure 8 Compared to the blocks depicted in , method 800 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks. Additionally or alternatively, two or more blocks of method 800 may be performed simultaneously or in parallel.

[0125] Furthermore, the method 800 may be implemented by a processing circuit system (eg, one or more processors or one or more integrated circuits). In an example, one or more processors may execute a program stored in a non-transitory computer-readable medium to perform one or more of the methods.

[0126] Fig. 9 is a simplified block diagram of an apparatus 900 for applying a frequency-dependent joint component secondary transform (FD-JCST) according to some embodiments.

[0127] The apparatus 900 may include an acquisition code 910 , a first determination code 920 , a second determination code 930 , and a processing code 940 .

[0128] The acquisition code 910 may be configured to cause at least one processor to obtain transform coefficients in a transform coefficient block.

[0129] The first determination code 920 may be configured to cause at least one processor to determine whether a transform coefficient in a transform coefficient block is a low frequency coefficient. As described above with respect to method 800, the first determination code may determine whether a transform coefficient is a low frequency coefficient based on various implementations.

[0130] The second determination code 930 may be configured to cause at least one processor to determine whether a low-frequency coefficient has a non-zero value when determining that a transform coefficient in a transform coefficient block is a low-frequency coefficient.

[0131] The processing code 940 may be configured to cause the at least one processor to perform a joint component secondary transform on the low frequency coefficients and indicate with relevant syntax that a JCST is being applied when it is determined that the low frequency coefficients have non-zero values.

[0132] The relevant syntax may be a high-level syntax (HLS).

[0133] The processing code 940 may also be configured to signal the transform kernel used in the JCST in a high-level syntax (HLS). For example, if the transform kernel is a 2×2 matrix, only one element of the kernel is signaled, and all other elements are derived based on the orthogonality and predefined norm of the transform kernel.

[0134] Alternatively, the processing code 940 may be configured to signal a set of predefined fixed transform kernels in HLS and signal the index of the transform kernel being used in the set instead of signaling the kernel element.

[0135] The HLS may indicate the cores used for each frequency, the cores used for each prediction mode, and / or the cores used for each associated primary transform type.

[0136] In addition, the processing code 940 may be configured to cause at least one processor to apply an integer transform kernel to the JCST. Here, the integer kernel may mean that all elements of the transform kernel are integers.

[0137] Based on applying an integer transform kernel to the JCST, the processing code 940 may be configured to divide the output of the JCST by a factor N, where N may be a power of 2. Alternatively, the output of the JCST may be clipped to a given data range such as [a, b]. For example, the values ​​of [a, b] may include [-2 15 , 2 15 -1] or [-2 M , 2 M -1], where M depends on the internal bit depth.

[0138] The above-described embodiments may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media.

[0139] Fig.10 1 is a block diagram of a computer 1000 suitable for implementing the embodiments.

[0140] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code comprising instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.

[0141] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0142] Fig.10 The components for computer system 1000 shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing the various embodiments. The configuration of components should also not be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of computer system 1000.

[0143] Computer system 1000 may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input by one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, tapping), visual input (such as gestures), olfactory input. Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious input by humans, such as audio (such as voice, music, ambient sounds), images (such as scanned images, photographic images obtained from a still image camera), and videos (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0144] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard 1010 , mouse 1020 , trackpad 1030 , touch screen 1100 , data gloves, joystick 1050 , microphone 1060 , scanner 1070 , camera 1080 .

[0145] The computer system 1000 may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include: tactile output devices (e.g., tactile feedback through touch screen 1100, data glove 1040, or joystick 1050, but there may also be tactile feedback devices that are not used as input devices), audio output devices (such as: speakers 1090 and headphones), visual output devices (e.g., screen 1100, including cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light-emitting diode (OLED) screen, each screen with or without touch screen input capability, each screen with or without tactile feedback capability - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output in a manner such as stereographic output; virtual reality glasses, holographic displays, and cigarette cans, and printers.

[0146] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1200 with CD / DVD or similar media 1210, thumb drives 1220, removable hard drives or solid-state drives 1230, traditional magnetic media such as tapes and floppy disks, dedicated ROM / ASIC / PLD based devices such as security dongles, etc.

[0147] Those skilled in the art should also understand that the term "computer-readable medium" used in conjunction with embodiments of the present disclosure does not include transmission media, carrier waves, or other transient signals.

[0148] The computer system 1000 may also include an interface to one or more communication networks. The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet, wireless local area networks, cellular networks (including global systems for mobile communication (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long-Term Evolution (LTE), etc.); television cable or wireless wide area digital networks, including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks, including CANBus, etc. Certain networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1490) (e.g., the universal serial bus (USB) port of the computer system 1000); other networks are typically integrated into the core of the computer system 1000 by attaching to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smart phone computer system). The computer system 1000 may use any of these networks to communicate with other entities. Such communication may be unidirectional receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CANBus to certain CANBus devices), or bidirectional (e.g., using a local digital network or a wide area digital network to other computer systems). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0149] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core 1400 of the computer system 1000.

[0150] The core 1400 may include one or more central processing units (CPUs) 1410, graphics processing units (GPUs) 1420, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 1430, hardware accelerators 1440 for certain tasks, etc. These devices, along with read-only memory (ROM) 1450, random-access memory (RAM) 1460, internal mass storage devices 1470 such as internal non-user accessible hard disk drives, solid-state drives (SSDs), etc., may be connected via a system bus 1480. In some computer systems, the system bus 1480 may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the system bus 1480 of the core directly or via a peripheral bus 1490. The architecture of the peripheral bus includes peripheral component interconnect (PCI), USB, etc.

[0151] The CPU 1410, GPU 1420, FPGA 1430, and accelerator 1440 may execute certain instructions, which in combination may constitute the above-mentioned computer code. The computer code may be stored in ROM 1450 or RAM 1460. Transient data may also be stored in RAM 1460, while permanent data may be stored in, for example, an internal mass storage device 1470. Fast storage and retrieval of any of the storage devices may be achieved by using a cache memory, which may be closely associated with one or more CPUs 1410, GPUs 1420, mass storage devices 1470, ROM 1450, RAM 1460, etc.

[0152] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the embodiments, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.

[0153] As an example and not limitation, a computer system 1000 having an architecture, and in particular a core 1400, can provide functionality due to a processor (including a CPU, a GPU, an FPGA, an accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device introduced above and certain storage devices of the core 1400 having non-transient properties, such as a core internal mass storage device 1470 or a ROM 1450. Software implementing various embodiments can be stored in such a device and executed by the core 1400. Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core 1400 and in particular the processor therein (including a CPU, a GPU, an FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM 1460 and modifying such a data structure according to a process defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic contained in circuits (e.g., accelerator 1440) that are hardwired or otherwise included, which may run in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and conversely, references to logic may also include software. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits containing logic for execution, or both. Implementations include any suitable combination of hardware and software.

[0154] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that those skilled in the art will be able to conceive of many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. A method for performing a joint component quadratic transform (JCST), in, The method comprises: Obtaining a plurality of transform coefficients of a plurality of color components and a residual component in the plurality of transform coefficients; performing a forward transform on the residual component; performing JCST on the forward transformed residual component; and The residual component is quantized and entropy encoded to obtain a bit stream.

2. The method according to claim 1, in, The residual components include residual component 0 which is a Cb transform coefficient and residual component 1 which is a Cr transform coefficient.

3. The method according to claim 2, in, The execution of JCST includes: Perform JCST element-wise on residual component 0 and residual component 1.

4. The method according to claim 1, in, Also includes: determining whether at least one of the plurality of transform coefficients is a low frequency coefficient; determining whether the low-frequency coefficient has a non-zero value based on determining that at least one of the plurality of transform coefficients is the low-frequency coefficient; and Based on determining that the low-frequency coefficients have non-zero values, a JCST is performed on the low-frequency coefficients.

5. The method according to claim 4, in, The determining whether at least one of the plurality of transform coefficients is a low-frequency coefficient comprises: Based on coordinates (x, y) of a transform coefficient block including a plurality of transform coefficients, it is determined whether at least one of the plurality of transform coefficients is a low-frequency coefficient.

6. The method according to claim 1, in, Also includes: The associated syntax is signaled to indicate execution of the JCST.

7. The method according to claim 6, in, Related grammars include High Level grammar (HLS), and Also includes: Signals the transform kernel used when executing JCST on HLS.

8. The method according to claim 7, in, The HLS indicates at least one of a transform kernel for each frequency of a plurality of transform coefficients, a transform kernel for each prediction mode, or a transform kernel for each transform type.

9. The method according to claim 4, in, The output of the JCST is divided by a factor of N, where N is a power of 2, or The output of the JCST is clipped within a predetermined data range.

10. A computer device for performing a joint component secondary transform (JCST), the computer device include: Processor, communication interface, memory and communication bus; Wherein, the processor, the communication interface and the memory communicate with each other through the communication bus; the communication interface is an interface of a communication module; The memory is used to store a computer program and transmit the computer program to the processor; The processor is used to call the computer program in the memory to execute the method according to any one of claims 1 to 9.

11. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a computer device, the method described in any one of claims 1 to 9 is implemented to generate and store a code stream.

12. A computer program product comprising a computer program, which, when executed on a computer device, causes the computer device to execute the method according to any one of claims 1 to 9.