Data processing method and related device
By introducing transformation mode syntax control information into the video encoding code stream, the problem of inflexible use control of transformation tools in the prior art is solved, and efficient encoding and decoding in different scenarios is achieved.
Patent Information
- Application Number
- PCT/CN2024/131363
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-19
AI Technical Summary
The use control of the transformation tools in the existing video encoding standards is relatively fixed, and it is unable to flexibly adapt to the video needs in different scenarios, resulting in low encoding and decoding efficiency.
By introducing transformation mode syntax control information into the encoded code stream of video, the use of transformation modes during the decoding process can be flexibly controlled, and the encoding and decoding efficiency can be improved.
It realizes adaptive control of the use of transformation mode in different scenarios, improves video encoding and decoding performance, and reduces encoding and decoding time.
Smart Images

Figure CN2024131363_19062025_PF_FP_ABST
Abstract
Description
A data processing method and related equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311709214.1 and application name “A data processing method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a data processing method, a data processing apparatus, a computer device, and a computer-readable storage medium. Background Art
[0003] With the development of computer technology, a wide variety of media content has emerged, including but not limited to video, audio, and images. To facilitate the transmission and storage of this media content, numerous audio and video coding standards have been proposed to implement encoding and decoding processes, thereby achieving compression and decompression of the media content. Currently, many transformation tools have been introduced into video coding standards to improve decoding performance. However, the control over the use of these transformation tools is relatively fixed. Given the varying transformation requirements of videos in different scenarios, fixed transformation modes do not effectively improve performance and may even increase encoding and decoding time, thereby affecting encoding and decoding efficiency.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a data processing method and related devices, which can flexibly control the use of transformation modes and improve encoding and decoding efficiency.
[0006] In one aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0007] Obtaining a video encoding stream; the encoding stream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode;
[0008] Decode the encoded code stream;
[0009] During the decoding process, the use of the transform mode is controlled according to the transform mode syntax control information.
[0010] In an embodiment of the present application, a video encoding stream can be obtained; the encoding stream includes transform mode syntax control information, which is used to control the use of transform modes. It can be seen that by adding transform mode syntax control information to the encoding stream, the use of transform modes during the decoding process can be flexibly controlled through the transform mode syntax control information, thereby effectively guiding the transform processing flow during the decoding process. The encoding stream is decoded, and during the decoding process, the use of transform modes is controlled according to the transform mode syntax control information. In this way, the use of transform modes can be flexibly controlled based on the transform mode syntax control information in the encoding stream during the decoding process, thereby improving encoding and decoding efficiency.
[0011] In one aspect, an embodiment of the present application provides another data processing method, the method comprising:
[0012] Determine the syntax control information of the transformation mode to be adopted by the video; the transformation mode syntax control information is used to control the use of the transformation mode;
[0013] Encode the video;
[0014] During the encoding process, the transform mode syntax control information is encoded into the encoded bitstream of the video.
[0015] In the embodiments of the present application, transform mode syntax control information is determined for a video and encoded into the coded bitstream, thereby accurately instructing the decoder on the use of transform modes. This transform mode syntax control information enables adaptively enabling or disabling corresponding transform modes in different scenarios, thereby enabling flexible and reasonable control over the use of transform modes in the corresponding scenarios, thereby improving video coding performance.
[0016] In one aspect, an embodiment of the present application provides a data processing device, comprising:
[0017] An acquisition unit is used to acquire a video encoding stream; the encoding stream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode;
[0018] A processing unit, configured to decode the encoded code stream;
[0019] The processing unit is further configured to control the use of the transform mode according to the transform mode syntax control information during the decoding process.
[0020] In one aspect, an embodiment of the present application provides another data processing device, the device comprising:
[0021] A determination unit, configured to determine syntax control information of a transform mode to be used for the video; the transform mode syntax control information is used to control the use of the transform mode;
[0022] A processing unit, configured to encode the video;
[0023] The processing unit is further configured to encode the transform mode syntax control information into the encoded bit stream of the video during the encoding process.
[0024] In one aspect, an embodiment of the present application provides a computer device, comprising:
[0025] a processor suitable for executing a computer program;
[0026] Computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned data processing method is implemented.
[0027] On the one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded by a processor and executes the above-mentioned data processing method.
[0028] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the above-mentioned data processing method is implemented.
[0029] On the one hand, an embodiment of the present application provides a method for processing a video stream, where the video stream is generated according to one of the above-mentioned data processing methods, or is decoded based on another of the above-mentioned data processing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1a is a schematic diagram of a video encoding process provided by an exemplary embodiment of the present application;
[0031] FIG1b is a schematic diagram of coefficient flipping provided by an exemplary embodiment of the present application;
[0032] FIG1c is a schematic diagram of division of a coding unit provided by an exemplary embodiment of the present application;
[0033] FIG1d is a schematic diagram of a transformation type adopted by a sub-block transformation provided by an exemplary embodiment of the present application;
[0034] FIG1e is a schematic diagram of the contents of a secondary transformation kernel provided by an exemplary embodiment of the present application;
[0035] FIG2 is an architecture diagram of a data processing system provided by an exemplary embodiment of the present application;
[0036] FIG3 is a flow chart of a data processing method provided by an exemplary embodiment of the present application;
[0037] FIG4a is a schematic structural diagram of a video sequence provided by an exemplary embodiment of the present application;
[0038] FIG4 b is a schematic diagram of a sheet structure provided by an exemplary embodiment of the present application;
[0039] FIG4c is a schematic diagram showing a relationship of a grammatical structure provided by an exemplary embodiment of the present application;
[0040] FIG5 is a flow chart of another data processing method provided by an exemplary embodiment of the present application;
[0041] FIG6a is a schematic structural diagram of a data processing device provided by an exemplary embodiment of the present application;
[0042] FIG6 b is a schematic structural diagram of another data processing device provided by an exemplary embodiment of the present application;
[0043] FIG7 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0044] In this application, the terms "first," "second," and so on are used to distinguish identical or similar items with substantially the same purpose or function. It should be understood that "first," "second," and "nth" do not have a logical or temporal dependency, nor do they limit quantity or order of execution. In this application, the term "at least one" means one or more, and "plurality" means two or more; for example, "at least one grammatical element" means one or more grammatical elements.
[0045] To facilitate understanding of the technical solution provided by this application, the video encoding process involved in the video encoding technology is first introduced below.
[0046] Video signals, depending on how they are acquired, can include video captured by a camera and video generated by a computer. Video captured by a camera is natural content captured by the camera. Computer-generated video is a type of screen content. Screen content simply refers to text, images, animations, videos, etc. generated by a computer. Examples include screen sharing, cloud gaming, and video conferencing. The statistical characteristics of the signal distribution between computer-generated screen content and natural content captured by a camera vary greatly in their distortion sensitivity to the human eye. For example, compared to video captured by a camera, screen content video may have large flat areas, less noise, many repetitive patterns and characters, a limited range of colors, high image contrast, and sharp edges. Therefore, based on the differences in statistical characteristics, the corresponding compression encoding methods may also differ. For example, for videos captured by cameras, HEVC (High Efficiency Video Coding, an international video coding standard, also known as H.265) and VVC (Versatile Video Coding, an international video coding standard, also known as H.266) are often used for compression encoding. For computer-generated videos, video coding standards that support screen content coding (SCC) functions are often used, including but not limited to HEVC-SCC (a video coding standard that supports screen content coding), AV1 (Alliance for Open Media Video1, a first-generation video coding standard developed by the Alliance for Open Media (AOM)), AV2 (Alliance for Open Media Video 2, a second-generation video coding standard developed by the Alliance for Open Media), and AVS3 (Audio Video Coding Standard 3, an audio and video coding standard).
[0047] HEVC, VVC, and AVS3 mentioned above all use a hybrid coding framework, which is a video coding technology that combines multiple coding methods (such as predictive coding, transform coding, quantization, and entropy coding) to improve video compression efficiency and image quality. For the input original video signal, when using the above video coding technologies for compression coding, the following operations and processing are generally included:
[0048] 1) Block partition structure: Based on the size of a unit, the input image can be divided into multiple non-overlapping processing units, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Units, also known as the maximum coding unit) or LCU (the maximum coding unit in HEVC). CTU can be further divided into more refined parts to obtain one or more basic coding units, called CU (Coding Unit). Each CU is the most basic element in the encoding process. The following description is for the various encoding methods that may be used for each CU.
[0049] 2) Predictive Coding: This includes intra-frame prediction (Intra-picture Prediction) and inter-frame prediction (Inter-picture Prediction). In intra-frame prediction, the predicted signal comes from a previously coded and reconstructed area within the same image; in inter-frame prediction, the predicted signal comes from a previously coded image (called a reference image) that is different from the current image. The original video signal is predicted using the selected reconstructed video signal to produce a residual video signal. The encoder selects an appropriate predictive coding mode for the current CU from a variety of predictive coding modes and notifies the decoder.
[0050] 3) Transform Coding and Quantization: The residual video signal undergoes transform operations such as the Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT) to convert the video signal into the transform domain, obtaining transform coefficients (i.e., DC coefficients). In one implementation, the transform coefficients can be subjected to a secondary transform to improve transform efficiency and overall encoder performance. A secondary transform is a second transform performed on the frequency domain signal (primary transform coefficients) after the primary transform, converting the signal from one transform domain to another. The primary transform is the first transform applied during the encoding process, while the secondary transform is applied after the primary transform and is typically used to further optimize the encoding performance. The rationale behind the secondary transform is that the DC coefficients (primary transform coefficients) typically have high statistical redundancy and large amplitude, making entropy coding expensive. A secondary transform can further remove redundancy, thereby achieving relatively better performance. The signal in the transform domain is subjected to a lossy quantization operation, which discards a certain amount of information, making the quantized signal conducive to compression expression. The degree of quantization fineness can be determined by the quantization parameter (QP). When the QP value is large, coefficients representing a larger value range will be quantized to the same output, which usually results in greater distortion and a lower bit rate. Conversely, when the QP value is small, coefficients representing a smaller value range will be quantized to the same output, which usually results in less distortion and a corresponding higher bit rate. In some video coding standards, there may be multiple transform modes (also called transform methods) to choose from. Therefore, the encoder can select one of the transform modes for the current CU and inform the decoder.
[0051] 4) Entropy Coding or Statistical Coding: The quantized transform domain signal can be statistically compressed and encoded based on the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream is output. At the same time, other information generated by the encoding (such as the selected mode, motion vector, etc.) also needs to be entropy encoded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0052] 5) Loop Filtering: The encoded image is subjected to inverse quantization, inverse transformation and prediction compensation (the reverse operations of 2) to 4) above) to obtain a decoded reconstructed image. Compared with the original image, the reconstructed image has some information different from the original image due to the influence of quantization, resulting in distortion. By performing filtering operations on the reconstructed image, such as deblocking filtering, SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filtering) filters, the degree of distortion caused by quantization can be effectively reduced. Since these filtered reconstructed images will be used as a reference for subsequent encoded images and used to predict future signals, the above filtering operation is also called loop filtering, and the filtering operation within the encoding loop.
[0053] The operations and processes described in 1)-5) above can be performed by a video encoder, and a basic flow chart of exemplary video encoding can be provided, as shown in FIG1a. This basic flow may include: inputting video data (i.e., an input video picture) into the video encoder. The video picture may be input in coding order and partitioned into blocks. For the current coding block (e.g., the coding unit being encoded) in the video picture, an intra-frame prediction mode decision and motion estimation are used to determine whether the prediction coding mode is intra-picture prediction or inter-frame prediction (i.e., motion-compensation prediction). Information about the prediction coding mode is then communicated to a coding mode decider. The current coding block can then be predictively coded according to the determined prediction coding mode to generate a prediction signal. Simultaneously, the coding mode decider may entropy code the relevant information (including coding modes, intra-frame prediction modes, and motion data) for encoding into the coded bitstream. The predicted signal and the current coded signal can be combined to produce a residual signal. This residual signal undergoes transformation and quantization, followed by inverse quantization and scaling. This signal is then combined with the predicted signal to produce a reconstructed signal. This reconstructed signal can be stored in a buffer (buffer for current picture) for intra-frame prediction. After in-loop filtering, the reconstructed signal can be stored in a decoder buffer for subsequent inter-frame prediction, where it can be decoded to produce a cached image (decoded picture buffer) as a reference for motion estimation. Furthermore, the reconstructed signal undergoes entropy coding to produce a coded bitstream. The transformation, quantization, and motion estimation processes are controlled by an encoder controller.It should be noted that, unless otherwise specified, the terms "encoded code stream", "decoded code stream", "encoded bit stream", "compressed code stream" and "code stream" in this application are equivalent.
[0054] Based on the encoding process described above, on the decoding end, for each CU, after receiving the encoded bitstream, the decoder first performs entropy decoding to obtain various coding mode information and quantized transform coefficients. Each quantized transform coefficient undergoes inverse quantization and inverse transformation to produce a residual signal. Furthermore, based on the known coding mode information, the prediction signal corresponding to the CU is obtained. These two signals are added together to produce a reconstructed signal. Finally, the reconstructed signal of the decoded image undergoes loop filtering to produce the final output signal, thereby achieving video decoding.
[0055] Currently, mainstream video coding standards such as HEVC, VVC, AVS3, AV1, and AV2 all adopt a block-based hybrid coding framework and follow the aforementioned coding process: first, the original video is divided into a series of coding units. Then, a combination of prediction, transform, and entropy coding methods is used to generate a coded bitstream, thereby achieving video compression. Different video coding standards support certain partitioning structures, transform modes, and other technologies. For example, AVS3 supports partitioning structures such as quadtree (QT), binary tree (BT), and extended quadtree (EQT). The largest coding unit (CTU) can be divided down layer by layer into coding units (CUs) using quadtree, binary tree, or extended quadtree structures. The CU is the basic unit of coding, and within the CU, processing can be performed on a transform unit (TU) basis. For example, compared to previous-generation coding standards, AVS3 supports more flexible transforms, including IST (Implicit Selection of Transforms), ISTS (Implicted Selected of Transform Skip), PBT (Position Based Transform), SBT (Sub-Block Transform), ST (Secondary Transform), EST (Enhanced Secondary Transform), and DEST (Decoding Enhanced Secondary Transform). For large coding units used in ultra-high-definition video encoding, a DCT transform of up to 64×64 can be used.
[0056] The following introduces each transformation mode (or transformation tool, transformation method) in the AVS3 standard mentioned above.
[0057] (1) IST (Implicit Selection of Transforms)
[0058] AVS3 Phase II adopted the IST technology, a video coding technique used to optimize the transform selection process to improve coding efficiency and image quality. The concept of IST is to dynamically select the most appropriate transform method based on the characteristics of the input signal. IST is a new transform tool added in AVS3 for intra blocks (i.e., intra-frame residual blocks). For intra-frame residual blocks, IST introduces the DST-VII transform in addition to the DCT-II transform. DCT-II and DST-VII are two separable transform kernels provided by IST for intra-frame residual blocks. When using DST-VII to transform intra-frame prediction residuals, DST-VII is used as the transform kernel for both horizontal and vertical transforms. In IST, the use of DST-VII is indicated by the parity of the number of even coefficients. Odd indicates the use of DST-VII, while even indicates the use of DCT-II. IST can be applied to coding units of sizes from 4×4 to 32×32, but not to coding units partitioned using a derived tree (DT). In IST technology, the encoder can select the optimal transform kernel based on RDO. However, the index of the selected transform kernel is not transmitted in the encoded bitstream. Instead, the index is hidden in the parity of the non-zero transform coefficients. The decoder can derive the corresponding transform kernel based on the parity.
[0059] (2) ISTS (Implicted Selected of Transform Skip)
[0060] ISTS is a technology used in video coding to improve coding efficiency. It allows the encoder to skip the transform step in certain circumstances, thereby reducing computational effort and bit rate. The second stage of AVS3 supports transform skip mode, and indicates whether the transform needs to be skipped by the parity of the number of non-zero coefficients in the coefficient block. When the implicit change skip is turned on in the sequence header, the decoder determines whether to perform an inverse transform or skip the transform based on the parity of the number of non-zero coefficients in the current reference block. For inter-frame images, the image header also has another switch to control whether the inter-frame CU (i.e., the CU using the inter-frame prediction mode) of this image uses transform skip. ISTS is similar to IST and is also for intra blocks. When ISTS is turned on, if the CU uses the transform skip mode, the flag is not transmitted directly, but is derived based on the parity.
[0061] 1) For intra CU
[0062] An ETS flag is transmitted in the coded bitstream. If the flag is 0, an inverse DCT-II (i.e., DCT2) transform is performed. If the flag is 1, whether to perform residual reverse (RTS) is determined based on the parity of the number of non-zero coefficients. If the number of non-zero coefficients is odd (index 1), the transform is skipped but reversed according to an exemplary coefficient flipping method shown in Figure 1b (e.g., RTS-I (RTS-II)). As shown in Figure 1b, the characters in the boxes (e.g., "a," "b," etc.) are used to indicate the position of the corresponding coefficients. After flipping, the positions of the coefficients at positions g and j remain unchanged. If the number of non-zero coefficients is even (index 0), the transform is skipped (Transform Skip, TS). The following table shows the indications of intra-frame skipping and flipping.
[0063] Table 1a
[0064] 2) Non-direct / skip inter-frame CU, non-SBT inter-frame CU
[0065] If the number of non-zero coefficients is odd, the inverse transform is not skipped directly; if the number of non-zero coefficients is even, an inverse DCT-II transform is performed. Implicit transform skipping is applicable to intra / inter predicted blocks or IBC (Intra Block Copy) coded blocks with sizes ranging from 4x4 to 32x32.
[0066] 3) For inter-frame SBT CU
[0067] Skip the inverse transform directly.
[0068] 4) For inter-frame direct / skip and non-SBT CU
[0069] Directly use the inverse DCT-II transform.
[0070] (3) PBT (Position Based Transform)
[0071] PBT is a technique used in video coding that optimizes transform selection by considering the characteristics of different image regions, thereby improving coding efficiency and image quality. PBT can be applied to inter-frame prediction residual blocks to better fit the characteristics of inter-frame residuals. For example, in Figure 1c, the coding unit is divided into four regions: 0, 1, 2, and 3. The corresponding horizontal and vertical transform combinations are shown in Table 1b below.
[0072] Table 1b
[0073] In the above table, Sub-block part represents the sub-block part, Horizontal transform represents the horizontal transform, and Vertical transform represents the vertical transform. By using different transform combinations in different areas, more efficient residual signal expression is achieved. The transform types allowed by PBT (also called transform kernels in PBT) include DCT8 and DST7. The maximum size of the coding unit using this method is 32x32, the minimum is 8x8, and the aspect ratio of the coding unit is not greater than 2. PBT takes into account the prediction residual characteristics of different positions in the coding unit, divides a transform unit into four units according to the quadtree structure, and uses DCT8 and DST7 transform kernels for row transform and column transform respectively, which further improves the efficiency of transform coding. Moreover, this position-based binding method can effectively adapt to the distribution law of residuals between frames at different positions, improves the performance of the transform and has lower complexity.
[0074] (4) SBT (Sub-Block Transform)
[0075] The second phase of AVS3 adopted SBT technology. SBT is a technology used in video coding and image processing to improve coding efficiency and image quality. SBT divides large blocks in an image or video frame into smaller sub-blocks and transforms these sub-blocks independently, so that it can more flexibly handle the characteristics of different regions. SBT divides the inter residual into two sub-blocks, where the residual of one sub-block defaults to 0 and the residual of the other sub-block defaults to non-zero. There are 8 options for the size and position of the non-zero residual sub-block in AVS3 (this information is transmitted in the bitstream). The transform of the non-zero residual sub-block adaptively selects DCT8 / DST7 transform as the horizontal transform and vertical transform according to the position of the sub-block. SBT is applied to the luminance residual block of the inter mode CU whose width and height are both less than or equal to 64.
[0076] 1) There are four sizes / orientations of non-zero residual sub-blocks: ① SBT-V-1: The width of the sub-block is 1 / 2 of the width of the residual block, and the height of the sub-block is the height of the residual block. ② SBT-V-2: The width of the sub-block is 1 / 4 of the width of the residual block, and the height of the sub-block is the height of the residual block. ③ SBT-H-1: The height of the sub-block is 1 / 2 of the height of the residual block, and the width of the sub-block is the width of the residual block. ④ SBT-H-2: The height of the sub-block is 1 / 4 of the height of the residual block, and the width of the sub-block is the width of the residual block.
[0077] 2) There are two locations for non-zero residual sub-blocks: ① the left side (for SBT-V) or the top side (for SBT-H) of the residual block. ② the right side (for SBT-V) or the bottom side (for SBT-H) of the residual block.
[0078] Combining the size / direction and position types shown in 1) and 2) above, there are a total of eight size / direction and position combinations, as shown in Figure 1d: SBT-V with position 0 / position 1, and SBT-H with position 0 / position 1. Position 0 represents the left / top side of the residual block, and position 1 represents the right / bottom side of the residual block. The size / direction combination is described by transmitting two flag bits in the bitstream. The position of the non-zero residual sub-block is derived from the parity of the number of non-zero coefficients. In addition, a flag bit is required to indicate whether the SBT mode is used. When the width or height of the non-zero residual sub-block is 64, the horizontal and vertical transforms of the non-zero residual sub-block are both DCT-2. In other cases, the horizontal and vertical transforms can be selected as shown in Figure 1d above.
[0079] Note that: ① Deblocking filtering is not performed on sub-block boundaries within the CU. ② If sub-block mode is used, the luminance CBF of the CU is 1, and the luminance CBF of all 4x4 positions in the CU is stored as 1.
[0080] The secondary conversion mode mentioned in the embodiments of the present application refers to a technology for implementing secondary conversion, which may include at least one of the following: ST, EST, and DEST. ST, EST, and DEST are introduced below.
[0081] (5) ST (Secondary Transform)
[0082] The first stage of AVS3 inherits the secondary transform in AVS2. The secondary transform is applied to the intra-predicted blocks, and only the upper left 4x4 block of the transform coefficients after the first transform is subjected to the secondary transform. The secondary transform of the first stage of AVS3 has no code stream indication, and all intra-predicted blocks will be subjected to the transform. The transform of the second stage of AVS3 is extended (called EST) and there is a flag in the code stream to indicate whether it is to be performed, and the secondary transform is extended to chroma. In the use of the secondary transform, for blocks encoded using the intra-prediction method, considering the special statistical characteristics of their residuals, the coefficients of the upper left 4×4 block can be subjected to the secondary transform, which further reduces the coding redundancy and makes the transform coefficients further concentrated.
[0083] (6) EST (Enhanced Secondary Transform)
[0084] In some cases, ST may not be suitable for intra CUs. Therefore, EST adds a flag, st_flag, to each TU (Transform Unit) level on top of ST to control the use of ST. "0" indicates that ST is not used, and "1" indicates that ST is used. Furthermore, ST can be applied to non-DT (Derived Tree) intra-coded CUs with sizes ranging from 4×4 to 64×64. In addition, a new 8x8 ST kernel has been introduced for CUs larger than 4x4.
[0085] The EST processing principle is described as follows: 1. Obtain the transform coefficient matrix. 2. Perform an inverse vertical transform on the transform coefficient matrix to obtain matrix K. 3. Perform an inverse horizontal transform on matrix K to obtain matrix H. 4. Use matrix H directly as the residual sample matrix ResidueMatrix.
[0086] For each step of implementation, please refer to the contents shown in the following process 1-4.
[0087] 1. If the current transform block is an intra prediction residual block, the value of M1 or M2 is greater than 4, and the value of StEnableFlag (ST enable flag) is equal to 1, perform the following operations on the CoeffMatrix (transform coefficient matrix). M1 and M2 are the size parameters of the intra prediction residual block, M1 represents the height of the intra prediction residual block, and M2 represents the width of the intra prediction residual block. The operations on the transform coefficient matrix mainly include the contents shown in the following a)-d).
[0088] a) Obtain a 4×4 matrix C from the transformation coefficient matrix CoeffMatrix. The code logic is as follows:
[0089] for(i=0;i<4;i++){
[0090] for(j=0;j<4;j++){
[0091] C[i][j]=CoeffMatrix[i][j]
[0092] }
[0093] }
[0094] b) If the current transform block meets one of the following four conditions: ① The current transform block is a luma prediction residual block, EnhancedStEnableFlag (EST enable flag) is equal to 0 and IntraLumaPredMode (intra luma prediction mode) is equal to 0-2 or 13-32 or 44-65, and the reference sample with coordinates (x0-1, y0+j-1) (j=1-M2) is "available". ② The current transform block is a luma prediction residual block, DtSplitFlag (DT split flag) is equal to 1, IntraLumaPredMode (intra luma prediction mode) is equal to 0-2 or 13-32 or 44-65, and the reference sample with coordinates (x0-1, y0+j-1) (j=1-M2) is "available". ③ The current transform block is a luma prediction residual block, EnhancedStEnableFlag (EST enable flag) and EstTuFlag (TU level EST flag) are both equal to 1. ④ The current transform block is a chroma prediction residual block, and both EnhancedStEnableFlag (EST enable flag) and ChromaStFlag (chroma level ST flag) are equal to 1. First, we obtain matrix P, then C. See the code logic below. S4 is the 4×4 inverse transform matrix.
[0095] P=C×S4
[0096] for(i=0;i<4;i++){
[0097] for(j=0;j<4;j++){
[0098] C[i][j]=Clip3(-32768,32767,(P[i][j]+2 6 )>>7)
[0099] }
[0100] }
[0101] c) If the current transform block meets one of the following four conditions: ① The current transform block is a luma prediction residual block, EnhancedStEnableFlag (EST enable flag) is equal to 0 and IntraLumaPredMode (intra luma prediction mode) is equal to 0 to 23 or 34 to 57, and the reference sample with coordinates (x0+i-1, y0-1) (i=1 to M1) is "available". ② The current transform block is a luma prediction residual block, DtSplitFlag (DT split flag) is equal to 1 and IntraLumaPredMode (intra luma prediction mode) is equal to 0 to 23 or 34 to 57, and the reference sample with coordinates (x0+i-1, y0-1) (i=1 to M1) is "available". ③ The current transform block is a luma prediction residual block, EnhancedStEnableFlag (EST enable flag) and EstTuFlag (TU level EST flag) are both equal to 1. ④ The current transform block is a chroma prediction residual block, and EnhancedStEnableFlag (EST enable flag) and ChromaStFlag (chroma level ST flag) are both equal to 1. Then first get the matrix Q, then get C, you can refer to the code logic shown below where S4 T It is the transpose of S4. The so-called transpose refers to the operation of swapping the rows and columns of a matrix.
[0102] Q=S4 T ×C
[0103] for(i=0;i<4;i++){
[0104] for(j=0;j<4;j++){
[0105] C[i][j]=Clip3(-32768,32767,(Q[i][j]+2 6 )>>7)
[0106] }
[0107] }
[0108] d) Modify the transformation coefficient matrix according to matrix C. The code logic is as follows.
[0109] for(i=0;i<4;i++){
[0110] for(j=0;j<4;j++){
[0111] CoeffMatrixf[i][j]=C[i][j]
[0112] }
[0113] }
[0114] 2. Perform vertical inverse transformation on the transformation coefficient matrix to obtain matrix K.
[0115] a) Obtain the matrix V from the transformation coefficient matrix and the inverse transformation matrix.
[0116] b) If M1 and M2 are both equal to 4 and StEnableFlag is equal to 1 and one of the following conditions ①-④ is met: ① If the current transform block is a luma intra-frame prediction residual block and EnhancedStEnableFlag (EST enable flag) is equal to 0; ② If the current transform block is a luma intra-frame prediction residual block and DtSplitFlag (DT split flag) is equal to 1; ③ If the current transform block is a luma intra-frame prediction residual block and EnhancedStEnableFlag (EST enable flag) and EstTuFlag (TU level EST flag) are both equal to 1; ④ If the current transform block is a chroma intra-frame prediction residual block and EnhancedStEnableFlag (EST enable flag) and ChromaStFlag (chroma level ST flag) are both equal to 1. Then the calculation logic of the matrix V is as follows, where D4 T It is the transpose of the 4×4 inverse transformation matrix D4.
[0117] V=D4 T ×CoeffMatrix
[0118] c) Obtain matrix K from matrix V. The code logic is as follows.
[0119] for(i=0;i <M1;i++){
[0120] for(j=0;j <M2;j++){
[0121] K[i][j]=Clip3(-32768,32767,(V[i][j]+2 4 )>>5)
[0122] }
[0123] }
[0124] 3. Perform horizontal inverse transform on matrix K to obtain matrix H.
[0125] a) Obtain matrix W from the transform coefficient matrix and the inverse transform matrix, and obtain the maximum value (MaxValue) and the conversion data shift1. Specifically, if the values of M1 and M2 are both equal to 4 and StEnableFlag (ST enable flag) is equal to 1 and one of the following conditions ①-④ is met: ① If the current transform block is a luma intra prediction residual block and EnhancedStEnableFlag is equal to 0; ② If the current transform block is a luma intra prediction residual block and DtSplitFlag (DT split flag) is equal to 1; ③ If the current transform block is a luma intra prediction residual block and EnhancedStEnableFlag (EST enable flag) and EstTuFlag (TU level EST flag) are both equal to 1; ④ If the current transform block is a chroma intra prediction residual block and EnhancedStEnableFlag (EST enable flag) and ChromaStFlag (chroma level ST flag) are both equal to 1. Then the calculation logic of the maximum value MaxValue, the conversion data shift1, and the matrix W is as follows, where D4 is a 4×4 inverse transform matrix.
[0126] W=K×D4
[0127] MaxValue=(1< <BitDepth)–1
[0128] shift1=22-BitDepth
[0129] b) Obtain matrix H from matrix W. The code logic is as follows.
[0130] for(i=0;i <M1;i++){
[0131] for(j=0;j <M2;j++){
[0132] H[i][j]=Clip3(-MaxValue-1,MaxValue,(W[i][j]+2 shift1-1 )>>shift1)
[0133] }
[0134] }
[0135] 4. Use the matrix H directly as the residual sample matrix (ResidueMatrix).
[0136] The secondary transformation inverse transformation matrices S4 and D4 involved in the above process are as follows:
[0137] S4={
[0138] {123,-35,-8,-3}
[0139] {-32,-120,30,10}
[0140] {14,25,123,-22}
[0141] {8,13,19,126}}
[0142] D4={
[0143] {34,58,72,81}
[0144] {77,69,-7,-75}
[0145] {79,-33 -75,58}
[0146] {55,-84,73,-28}}
[0147] (7) DEST (Decoder-side Enhanced Secondary Transform)
[0148] AVS3's next-generation reference software platform (Enhanced Video Model, EVM) incorporates DEST technology, an optimized and improved solution for the secondary transform of the luma component in EVM. This solution introduces a larger secondary transform kernel, increases the number of secondary transform candidates, and adapts and updates the representation of the secondary transform type. The luma secondary transform in AVS3 uses a flag bit to explicitly indicate whether the current block uses a secondary transform. This method cannot further expand the transform types, and encoding the flag bit also incurs a certain amount of additional overhead. To address this, DEST introduces new transforms, adds transform types, and updates the representation method. In its implementation, DEST introduces an 8x8 secondary transform kernel for sizes greater than 8, as shown in Figure 1e. The new secondary transform type introduced in DEST is primarily achieved by flipping the current residual block. Three different flipping methods are used: horizontal flip, vertical flip, and 180° diagonal flip.
[0149] In one implementation, DEST uses a decoder-side prediction approach. Specifically, at the decoder side, pixels within the current block's neighborhood are used to predict the boundary pixel value P of the current block. The transform coefficients are inversely transformed using different secondary transform types to obtain four candidate boundary pixel value results Qi (i = 1, 2, 3, 4). The absolute values Ri of the differences between the four candidate boundary pixel values Qi and the predicted boundary pixel value P are then calculated. Finally, the four absolute values Ri are sorted, and the type with the smallest absolute value Ri is selected as the secondary transform type for actual decoding.
[0150] Based on the DEST secondary transform type in the above method, the secondary transform type with the smallest Ri is selected by default. In another feasible implementation, the encoder can also perform the above-mentioned decoding end prediction. After sorting Ri, the encoder can send the index of the selected secondary transform type in the list. The decoder decodes the index and performs the same rearrangement operation to derive the list and confirm the selected secondary transform type. In this way, the decoder can determine the secondary transform type based solely on the index, thereby improving decoding performance.
[0151] Based on the above related description, an embodiment of the present application provides a data processing solution, which includes a processing flow at the encoding end and a processing flow at the decoding end.
[0152] (1) The processing flow of the encoding end is roughly as follows:
[0153] ① Determine the transformation mode syntax control information to be used for the video; the transformation mode syntax control information is used to control the use of the transformation mode; ② Encode the video; ③ During the encoding process, encode the transformation mode syntax control information into the video encoding stream.
[0154] (2) The processing flow of the decoding end is roughly as follows:
[0155] ① Obtain the encoded bitstream of the video; the encoded bitstream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode; ② Decode the encoded bitstream; ③ During the decoding process, control the use of the transform mode according to the transform mode syntax control information.
[0156] As can be seen from the above scheme, by adding transform mode syntax control information to the coded bitstream, the present application can flexibly control the use of transform modes during the decoding process through this transform mode syntax control information, thereby effectively guiding the transform processing flow during the decoding process. The coded bitstream is decoded, and during the decoding process, the use of transform modes is controlled according to the transform mode syntax control information. In this way, the use of transform modes can be flexibly controlled during the decoding process based on the transform mode syntax control information in the coded bitstream, thereby improving encoding and decoding efficiency. Through this transform mode syntax control information, corresponding types of transform modes can be adaptively turned on or off in different scenarios, making the use of transform modes in the corresponding scenarios flexible and reasonable, which is conducive to improving video encoding and decoding performance.
[0157] The following describes a data processing system suitable for implementing the embodiments of the present application in conjunction with FIG2 . As shown in FIG2 , the data processing system 20 may include a service device 201 and a consumer device 202 . The service device 201 may serve as a video encoding end; the service device 201 may be a terminal device or a server. The consumer device 202 may serve as a video decoding end; the consumer device 202 may be a terminal device or a server. A communication connection may be established between the service device 201 and the consumer device 202 . The terminal device may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, smart TV, etc., but is not limited thereto. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0158] The process of data processing performed by the service device 201 and the consumption device 202 is as follows: for the service device 201, it mainly includes the following data processing processes: ① the video acquisition process; ② determining the transformation mode syntax control information required for the video; ③ the video encoding process; for the consumption device 202, it mainly includes the following data processing processes: ④ the video decoding process; ⑤ controlling the use of the transformation mode according to the transformation mode syntax control information.
[0159] In addition, the video transmission process between the service device 201 and the consumer device 202 can be carried out based on various transmission protocols (or transmission signaling). The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, smart media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.
[0160] The following is a detailed description of the various processes involved in data processing:
[0161] ①Video acquisition process
[0162] Service device 201 can acquire video. Video can be acquired through scene capture or device generation. Scene capture refers to capturing real-world visual scenes using a capture device associated with service device 201. The capture device is used to provide video acquisition services for service device 201. Capture devices may include, but are not limited to, any of the following: cameras, sensors, and scanners. Cameras may include standard cameras, stereo cameras, and light field cameras. Sensors may include lasers and radars. Scanners may include 3D laser scanners. The capture device associated with service device 201 may refer to a hardware component within service device 201, such as a terminal camera or sensor. It may also refer to a hardware device connected to service device 201, such as a camera connected to service device 201. Device generation refers to the service device 201 generating video based on virtual objects, such as recording screen content or generating stereoscopic video based on virtual 3D objects and scenes generated through 3D modeling, thereby providing a more immersive visual experience.
[0163] ② Determine the transformation mode syntax control information required for the video.
[0164] In one implementation, the service device 201 can determine the syntax elements required for the syntax structure at at least one level for the video based on the encoding scenario to which the video belongs, and each syntax element is used to control the use of the secondary transform mode by the object of the corresponding syntax structure (such as a video sequence); thereby, transform mode syntax control information is formed based on at least one syntax element, and the transform mode syntax control information is used to control the use of the secondary transform mode.
[0165] ③Video encoding process
[0166] The service device 201 can encode the video to obtain the encoded code stream of the video. For the video encoding process, the encoding process shown in Figure 1a above can be referred to, the video is divided into multiple coding units, and the encoding process is performed in combination with video encoding methods such as prediction, transformation and entropy coding. In the process of encoding the video, the service device 201 can determine the transformation mode to be used based on the transformation mode syntax control information, and then perform transformation processing according to the transformation mode. For example, a certain coding unit uses the intra-frame prediction mode for prediction processing in the prediction stage, and uses the DEST mode for secondary transformation processing in the transformation processing stage. A certain coding unit uses the inter-frame prediction mode for prediction processing in the prediction stage, and uses the PBT mode for transformation processing in the transformation processing stage.
[0167] In one implementation, after obtaining the transform mode syntax control information, the service device 201 can encode the transform mode syntax control information into the coded bitstream of the video. If the transform mode syntax control information is used to control the use of the secondary transform mode, then during the encoding process, syntax elements can be added to the syntax structure of each level to achieve the encoding of the transform mode syntax control information into the coded bitstream. For example, the acquired video is an ultra-high-definition video, and the corresponding encoding scenario is an ultra-high-definition video encoding scenario. In order to improve the encoding efficiency of the ultra-high-definition video encoding scenario, syntax elements can be added to the sequence header, picture header, and slice header to control the opening or closing of the DEST mode, and these syntax elements are the contents of the transform mode syntax control information. After the encoding process, a coded bitstream containing the transform mode syntax control information can be obtained, so that during the decoding process, the use of the transform mode can be controlled based on the transform mode syntax control information. In this application, the coded bitstream of the video refers to the binary data stream obtained after the video is encoded. The original data can be converted into a series of binary data through encoding, and includes relevant information in the encoding process (such as the above-mentioned transform mode syntax control information) so that the video can be effectively restored during transmission or storage.
[0168] ④Video decoding process
[0169] The consumer device 202 can obtain the encoded bitstream of the video and decode the encoded bitstream. The decoding process can obtain the transform mode syntax control information in the encoded bitstream. During the decoding process, the consumer device 202 can control the use of the transform mode according to the transform mode syntax control information. For example, the transform mode syntax control information is used to control the use of various types of secondary transform modes (such as the DEST mode). Therefore, the transform mode syntax control information can be used to determine whether the DEST mode is permitted.
[0170] ⑤Control the use of transformation mode according to the transformation mode syntax control information
[0171] The transformation modes in the present application may include but are not limited to secondary transformation modes (such as ST, EST and DEST) and other transformation modes (such as PBT, SBT or IST, etc.). Based on the type of transformation mode, the control of the use of the transformation mode by the transformation mode syntax control information may include: allowing the use of certain types of transformation modes and / or prohibiting the use of certain types of transformation modes. As for the control of the use of a certain type of transformation mode, it may include: indicating that a certain type of transformation mode is allowed to be used or indicating that a certain type of transformation mode is prohibited to be used. For example, if the transformation mode includes a secondary transformation mode, then the transformation mode syntax control information can be used to control the use of the secondary transformation mode, specifically to indicate that the use of the DEST mode is allowed.
[0172] In the technical solution provided by this application, the transform mode syntax control information can be high-level syntax control information related to the transform mode. It can serve as a high-level syntax control switch to enable or disable the transform mode within the high-level syntax structure, thereby flexibly controlling the use of the transform mode. Furthermore, this transform mode syntax control information can be provided for different video encoding and decoding scenarios. This high-level syntax control information is used to determine whether to use the corresponding transform mode, allowing the corresponding type of transform mode to be adaptively enabled or disabled based on different scenarios, rather than remaining enabled by default. This allows the transform mode to be adapted to the needs of the encoding and decoding scenario, thereby increasing the flexibility of controlling the use of the transform mode and improving encoding and decoding performance.
[0173] The technical solution provided by the embodiments of the present application can be applied to any video encoding and decoding scenario, including but not limited to: screen content encoding and decoding scenario, non-screen content encoding and decoding scenario, ultra-high-definition video encoding and decoding scenario, high-definition video encoding and decoding scenario, etc. The transformation mode includes a secondary transformation mode, so it can also be applied to video codecs or video compression products that use secondary transformation (such as DEST) technology, and this application does not limit this. For example, the EVM platform adopts DEST technology, and DEST is turned on by default (i.e., DEST is allowed to be used). By testing test video sequences of different scenarios, it is found that in the screen content encoding and decoding scenario, if DEST is kept on, the encoding performance of the test video sequence in the screen content encoding and decoding scenario cannot be improved, and it may also cause an increase in encoding time. However, this solution can control the use of DEST in the screen content encoding and decoding scenario through transformation mode syntax control information, such as turning off DEST in the high-level syntax (i.e., prohibiting the use of DEST), so that based on the flexible control of DEST, encoding time and decoding time can be saved to a certain extent, thereby improving encoding performance and decoding performance.
[0174] Please refer to Figure 3, which is a flow chart of a data processing method provided in an embodiment of the present application. The data processing method can be executed by the consumer device 202 in the data processing system, and the method includes the following steps S301-S303.
[0175] S301, obtaining a coded bitstream of a video; the coded bitstream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode.
[0176] In one specific implementation, the transform mode includes at least one of the following: position-based transform (PBT), sub-block transform (SBT), and secondary transform mode. The secondary transform mode includes multiple types, including but not limited to secondary transform mode (ST) (referred to as ST mode), enhanced secondary transform mode (EST) (referred to as EST mode), and decoder-side enhanced secondary transform mode (DEST) (referred to as DEST mode). Transform mode syntax control information can be used to control the use of one or more of these transform modes. For example, if the transform mode includes a secondary transform mode, the transform mode syntax control information can be used to control the use of the secondary transform mode. Regarding the control of the use of the secondary transform mode, based on the multiple types of secondary transform modes, for a particular type of secondary transform mode, controlling its use can include allowing or prohibiting the use of a certain type of secondary transform mode. More broadly, the transform mode syntax control information can be used to indicate the types of secondary transform modes that are permitted and / or prohibited. For example, the secondary transform mode ST is permitted, while the enhanced secondary transform mode EST and decoder-side enhanced secondary transform mode DEST are prohibited.
[0177] In one implementation, the transform mode syntax control information is set in a syntax structure of a coded bitstream, where the syntax structure includes at least one of the following: a video sequence, a picture, a slice, a coding unit, and a transform unit.
[0178] The syntax structure of a video coded stream refers to the unit structure used to process the video. In one implementation, when a coded stream includes multiple syntax structures, the various syntax structures are hierarchically ordered. Optionally, the syntax structure of the coded stream may include the following: video sequence, picture, slice, maximum coding unit, coding unit, and transform unit. In terms of hierarchy, the following order applies: video sequence > picture > slice > maximum coding unit > coding unit > transform unit. The video sequence is the highest-level syntax structure in the coded stream. Based on this hierarchy, decoding of the coded stream begins at the highest level. Note that since the encoding process includes determining parameters for each level and writing the parameters into the bitstream, parameters can be written into the bitstream during video encoding. The order in which parameters are written into the bitstream is similar to the order in which they are decoded: higher-level information is written first, followed by lower-level information. This ensures that higher-level information is available from the outset of decoding.
[0179] A video sequence is a dataset consisting of a series of consecutive video frames (i.e., images). A video sequence begins with a sequence header and ends with a sequence end code or video editing code. The sequence header stores key parameters and information that describe the properties of the video sequence, making it a crucial component of video coding standards. For example, the sequence header includes information such as image size, frame rate, color space, and encoding parameters. Before decoding a video sequence, the sequence header can be read and parsed to obtain the sequence configuration information. Decoding is then performed based on this information to restore the original video frames. The sequence end code or video editing code indicates the end of a video sequence. The sequence headers between the first sequence header and the first occurrence of the sequence end code or video editing code are repeated sequence headers. Each sequence header is followed by one or more images, and each image must be preceded by an image header. Images are arranged in the coded bitstream in stream order. The stream order is the same as the decoding order, but the decoding order may differ from the display order. For a structural diagram of a video sequence, see FIG4a , which shows a 2-second video sequence, including a sequence header, followed by N frames of images, and a sequence end code after the N frames of images to indicate the end of the video sequence.
[0180] An image is a coded image. An image can be a frame or a field. Its coded data begins with a picture start code and ends with a sequence start code, a sequence end code, or the next picture start code. In the coded bitstream, the coded data of the two fields of an interlaced image can appear sequentially or interleaved. The decoding and display order of the two fields is specified in the image header. For interlaced images, even-numbered lines are called top field lines, odd-numbered lines are called bottom field lines, all top field lines are called top fields, and all bottom field lines are called bottom fields. Therefore, the two fields mentioned above include the bottom field and the top field. Image types can include: I-frame images (Intra Frame, which can be abbreviated as I-frame), P-frame images (Predicted Frame, which can be abbreviated as P-frame), and B-frame images (Bi-directional predicted Frame, which can be abbreviated as B-frame). Among them, I-frames are key frames. Each I-frame is a complete image frame. It does not rely on other frames for decoding and can be displayed independently. P frames are predicted frames, generated by performing motion estimation and differential coding on a forward reference frame. They only store the difference between the frame and the previous one, and by referencing the previous frame for decoding, the complete image can be restored. Compared to I frames, P frames have higher compression rates because they only store the changed portion. B frames are bidirectionally predicted frames, referencing both the previous and next frames for motion estimation and differential coding. B frames store the difference between the frame and the previous and next reference frames, and by referencing these two reference frames for decoding, the complete image can be restored. Compared to P frames, B frames have higher compression rates because they store more difference information.
[0181] A tile (or slice) refers to a rectangular area in an image, which can contain parts of multiple maximum coding units within the image, and the tiles do not overlap. This is because the image can be divided into maximum coding units, the maximum coding units do not overlap, and the samples in the upper left corner of the maximum coding unit do not support exceeding the image boundary, while the samples in the lower right corner of the maximum coding unit support exceeding the image boundary. For example, please refer to the structural diagram of the tile as shown in Figure 4b, where the area identified by each character (including A, B, C, D, E, F) represents a tile structure. It should be noted that each tile in the image has its own header, and each header is a part used to describe and configure the parameters and information of each tile. The header contains some important parameters, such as motion vectors, quantization parameters, etc., which are used by the decoder to correctly decode each tile.
[0182] The maximum coding unit (MCU) is obtained by dividing an image. MCUs do not overlap and can be divided into one or more coding units, determined by a coding tree (such as a binary tree or quadtree). A coding unit can be further divided into multiple transform units. Based on the above description, a schematic diagram of the relationship between the grammatical structures of video sequences, images, slices, and MCUs can be seen in Figure 4c.
[0183] In the present application, the transform mode syntax control information can be set in one or more syntax structures in the video sequence, image, slice, coding unit and transform unit. Based on the existence of syntax structures such as coding unit and transform unit, the video sequence, image and slice are relatively high-level syntax structures. Therefore, the transform mode syntax control information set in the video sequence, image and slice is a high-level syntax control information.
[0184] In one feasible implementation, if the transform mode includes a secondary transform mode, the transform mode syntax control information can be used to control the use of the secondary transform mode. The transform mode syntax control information includes at least one syntax element, and each syntax element is used to control the use of the secondary transform mode by the current object; the current object refers to the syntax structure in the coded bitstream that is being decoded. In other words, the use of the secondary transform mode by various syntax structures can be controlled by the syntax elements in the transform mode syntax control information. For example, the transform mode syntax control information includes two syntax elements, one syntax element is used to control the use of the secondary transform mode by a video sequence that is being decoded in the coded bitstream, and the other syntax element is used to control the use of the secondary transform mode by an image that is being decoded in the coded bitstream. Controlling the use of the secondary transform mode by the current object includes instructing the current object to allow the use of certain types of secondary transform modes and / or prohibiting the current object from using certain types of secondary transform modes.
[0185] Based on the fact that each syntax element is used to control the use of the secondary transform mode by the current object (i.e., a syntax structure in the coded bitstream that is being decoded), the syntax structure corresponding to the current object involved is the syntax structure acted upon by the syntax element. For example, if a certain syntax element is used to control the use of the secondary transform mode by the current video sequence, it can be understood that this syntax element acts on the current video sequence, and the video sequence is the syntax structure acted upon by this syntax element, and this syntax element is set in the sequence header. Based on this, the syntax elements included in the transform mode syntax control information can be set in the corresponding type of video header (e.g., sequence header, picture header, or slice header) according to the syntax structure acted upon by the syntax element. Each type of video header corresponds to a syntax structure and is a part of the corresponding syntax structure, used to store and transmit key parameters and information of the corresponding syntax structure. The syntax elements included in the transform mode syntax control information are introduced in detail below. Please refer to ①-③ below.
[0186] ① If the syntax structure includes a video sequence, the transform mode syntax control information is set in the sequence header of the current video sequence. The transform mode syntax control information is the sequence header syntax element (seq_flag). The current object is the current video sequence. The sequence header syntax element is used to control the use of the secondary transform mode by the current video sequence.
[0187] The current video sequence refers to the video sequence that is being decoded in the coded bitstream. The transform mode syntax control information includes a sequence header syntax element (also called a sequence header flag) in the sequence header of the current video sequence. The sequence header syntax element acts on the current video sequence. Specifically, based on the multiple types of secondary transform modes, the sequence header syntax element can be used to indicate one or more secondary transform modes that are allowed to be used by the current video sequence, and / or one or more secondary transform modes that are prohibited from being used, when controlling the use of the secondary transform mode by the current video sequence. For example, the sequence header syntax element is used to indicate that the current video sequence allows the use of ST mode, EST mode, and DEST mode. For another example, the sequence header syntax element is used to indicate that the current video sequence allows the use of ST mode, but prohibits the use of EST mode and DEST mode.
[0188] ② If the syntax structure includes an image, the transform mode syntax control information is set in the image header of the current image. The transform mode syntax control information is the image header syntax element (pic_flag). The current object is the current image. The image header syntax element is used to control the use of the secondary transform mode by the current image.
[0189] The current image refers to the image that is being decoded in the coded code stream. The transform mode syntax control information includes an image header syntax element (also called an image header flag) set in the image header of the current image, and the image header syntax element acts on the current image. Specifically, based on the multiple types of secondary transform modes, the image header syntax element can be used to indicate one or more secondary transform modes that are allowed to be used for the current image, and / or one or more secondary transform modes that are prohibited from being used, when controlling the use of the secondary transform mode by the current image. For example, the image header syntax element is used to indicate that the current image allows the use of ST mode, EST mode, and DEST mode. For another example, the image header syntax element is used to indicate that the current image allows the use of ST mode, but prohibits the use of EST mode and DEST mode.
[0190] ③ If the syntax structure includes slices, when the transform mode syntax control information is set in the slice header of the current slice, the transform mode syntax control information is the slice header syntax element (slice_flag), the current object is the current slice, and the slice header syntax element is used to control the use of the secondary transform mode by the current slice.
[0191] The current slice refers to the slice in the coded bitstream that is undergoing decoding processing. The transform mode syntax control information includes a slice header syntax element (also called a slice flag or a slice header flag) set in the slice header of the current slice. The slice header syntax element acts on the current slice. Specifically, based on the multiple types of secondary transform modes, the slice header syntax element can be used to indicate one or more secondary transform modes that are allowed to be used by the current slice, and / or one or more secondary transform modes that are prohibited from being used, when controlling the use of the secondary transform mode by the current slice. For example, the slice header syntax element is used to indicate that the current slice allows the use of ST mode, EST mode, and DEST mode. For another example, the slice header syntax element is used to indicate that the current slice allows the use of ST mode, but prohibits the use of EST mode and DEST mode.
[0192] It is understood that the syntax structure based on the coded bitstream can include multiple types of video sequences, pictures, and slices, and the above-mentioned syntax elements can also be combined into multiple types. That is, the transform mode syntax control information can include multiple types of sequence header syntax elements, picture header syntax elements, and slice header syntax elements. For example, the transform mode syntax control information includes: sequence header syntax elements and picture header syntax elements. For another example, the transform mode syntax control information includes: sequence header syntax elements and slice header syntax elements. For another example, the transform mode syntax control information includes: sequence header syntax elements, picture header syntax elements, and slice header syntax elements.
[0193] To facilitate understanding of the decoding of various syntax elements in this application, the following briefly introduces the bitstream description method and descriptor.
[0194] In this application, transform mode syntax control information may be included in a bitstream syntax description, similar to the C language. Bitstream syntax elements are represented in boldface. Each syntax element is described by its name (a group of lowercase letters separated by underscores), syntax, and semantics. Syntax element values in syntax tables and text are represented in regular font.
[0195] In some cases, syntax tables may use other variable values derived from syntax elements. Such variables are named using a mix of lowercase and uppercase letters without underscores in syntax tables and the text. Variables beginning with an uppercase letter are used to decode the current and related syntax structures, as well as subsequent syntax structures. Variables beginning with a lowercase letter are used only within the subsection in which they appear.
[0196] The relationship between syntax element value mnemonics and variable value mnemonics and their values is described in the text. In some cases, the two are used equivalently. Mnemonics are represented by one or more letter groups separated by underscores. Each letter group begins with a capital letter and may contain multiple capital letters.
[0197] When the length of a bit string is an integer multiple of 4, hexadecimal notation may be used. The hexadecimal prefix is "0x." For example, "0x1a" represents the bit string "00011010." In conditional statements, 0 represents FALSE, and non-zero represents TRUE. The syntax table describes a superset of all bitstream syntaxes that conform to this document. Additional syntax restrictions are specified in the relevant clauses. When a syntax element appears, it indicates that a data unit is read from the bitstream.
[0198] Table 2a below provides a pseudo code example of syntax description: When a syntax element appears, it indicates that a data unit is read from the bit stream.
[0199] Table 2a
[0200] The parsing and decoding processes are described using text and C-like pseudocode.
[0201] Returns the next n bits of the bitstream, MSB first, and advances the bitstream pointer by n bits. If n is 0, 0 is returned and the bitstream pointer is not advanced. This function is also used to describe the parsing and decoding processes.
[0202] The descriptor represents the parsing process of different syntax elements, and the descriptor shown in Table 2b below can be seen.
[0203] Table 2b
[0204] For the decoding of various syntax elements in this application, please refer to the above content.
[0205] In one embodiment, a syntax element may include one or more mode flags, and each mode flag may be used to control the use of one or more types of secondary transform modes. The types of secondary transform modes include the following: secondary transform mode ST (which may be referred to as ST mode), enhanced secondary transform mode EST (which may be referred to as EST mode), and decoder-side enhanced secondary transform mode DEST (which may be referred to as DEST mode). For ease of understanding, any syntax element in the transform mode syntax control information is represented as a target syntax element. Similar logic applies to each syntax element. Taking the target syntax element as an example, the specific values of each mode flag in the target syntax element and the role of the target syntax element are introduced below, including the contents shown in (1)-(3) below.
[0206] (1) The target syntax element includes a first mode flag (st_flag). The first mode flag can be used to indicate the use of the secondary transform mode ST, the enhanced secondary transform mode EST, and the decoder-side enhanced secondary transform mode DEST by the current object. It can be understood that, based on the above content, if the target syntax element is a sequence header syntax element, then the current object is the current video sequence; if the target syntax element is a picture header syntax element, then the current object is the current picture; if the target syntax element is a slice header syntax element, then the current object is the current slice. The same applies to the target syntax elements mentioned below.
[0207] In the case where the target syntax element includes only one mode flag, the target syntax element may control the use of the secondary transform mode by including the following methods 1 and 2.
[0208] Method 1: ① If the value of the first mode flag is the first numerical value, the target syntax element is used to indicate that the current object prohibits the use of the secondary transform mode ST, the enhanced secondary transform mode EST, and the decoding-side enhanced secondary transform mode DEST; ② If the value of the first mode flag is the second numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding-side enhanced secondary transform mode DEST; ③ If the value of the first mode flag is the third numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding-side enhanced secondary transform mode DEST; ④ If the value of the first mode flag is the fourth numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST, and the decoding-side enhanced secondary transform mode DEST.
[0209] In the present application, the first numerical value is, for example, "0", the second numerical value is, for example, "1", the third numerical value is, for example, "2", and the fourth numerical value is, for example, "3". Prohibited use means that use is not allowed. As can be seen from the above method, the target syntax element only includes one mode flag (i.e., the first mode flag). Different values of the first mode flag make the role of the target syntax element different, or it can also be understood that the semantics corresponding to the target syntax element are different. Specifically, the target syntax element controls the use of the secondary transform mode, which is actually controlled by the first mode flag to control the current object's use of the secondary transform mode ST, enhanced secondary transform mode EST and decoding-end enhanced secondary transform mode DEST.
[0210] In one implementation, the target syntax element is a sequence header syntax element (seq_flag). The first mode flag included in the sequence header syntax element is denoted as seq_st_flag (sequence header secondary transform flag). The sequence header syntax element is used to control the use of secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST in the current video sequence. For the values of the first mode flag in the sequence header syntax element and the corresponding semantic content under different values, please refer to Table 3a below.
[0211] Table 3a
[0212] The logic for decoding the sequence header flags can be found in Table 3a-1 below:
[0213] Table 3a-1
[0214] Here, n=ceil(log2(number of patterns)), for example, n=ceil(log2(4))=2.
[0215] The sequence header secondary transform flag (seq_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of the ST mode; a value greater than '0' allows the use of the ST mode. The value of SeqStFlag is equal to the value of seq_st_flag. If seq_st_flag does not exist in the bitstream, the value of SeqStFlag is 0.
[0216] In another implementation, the target syntax element is a picture header syntax element (pic_flag, picture header flag). The first mode flag included in the picture header syntax element is denoted as pic_st_flag (picture header secondary transform flag). The picture header syntax element is used to control the use of secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST for the current picture pair. For the values of the first mode flag in the picture header syntax element and the corresponding semantic content under different values, please refer to Table 3b below.
[0217] Table 3b
[0218] For the decoding logic of the above-mentioned image header flag, please refer to the following table 3b-1:
[0219] Table 3b-1
[0220] Here, n=ceil(log2(number of patterns)), for example, n=ceil(log2(4))=2.
[0221] The picture header secondary transform flag (pic_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of ST mode; a value greater than '0' allows the use of ST mode. The value of PicStFlag is equal to the value of pic_st_flag. If pic_st_flag does not exist in the bitstream, the value of PicStFlag is 0.
[0222] In another implementation, the target syntax element is a slice header syntax element (slice_flag). The first mode flag included in the slice header syntax element is slice_st_flag (slice header secondary transform flag). The slice header syntax element is used to control whether to prohibit the use of secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST in the current slice. For the values of the first mode flag in the slice header syntax element and the corresponding semantic content under different values, please refer to Table 3c below.
[0223] Table 3c
[0224] For the decoding logic of the above-mentioned slice header flag, please refer to the content in the following Table 3c-1:
[0225] Table 3c-1
[0226] Here, n=ceil(log2(number of patterns)), for example, n=ceil(log2(4))=2.
[0227] The slice header secondary transform flag (slice_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of the ST mode; a value greater than '0' allows the use of the ST mode. The value of SliceStFlag is equal to the value of slice_st_flag. If slice_st_flag does not exist in the bitstream, the value of SliceStFlag is 0.
[0228] Method 2: ① If the value of the first mode flag is the first numerical value, the target syntax element is used to indicate that the current object prohibits the use of the secondary transform mode ST, the enhanced secondary transform mode EST, and the decoding-end enhanced secondary transform mode DEST; ② If the value of the first mode flag is the second numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding-end enhanced secondary transform mode DEST; ③ If the value of the first mode flag is the third numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST, and the decoding-end enhanced secondary transform mode DEST.
[0229] In this approach, the first mode flag included in the target syntax element has only three possible values. The secondary transform mode (ST) and the enhanced secondary transform mode (EST) can be understood as being controlled together, meaning that either both are permitted or neither is. When both are permitted, the use of the enhanced secondary transform mode (DEST) at the decoder can be flexibly controlled based on the different values of the first mode flag. When the secondary transform mode (ST) is not permitted, both the enhanced secondary transform mode (EST) and the enhanced secondary transform mode (DEST) at the decoder are prohibited.
[0230] In one implementation, the target syntax element is a sequence header syntax element (seq_flag). The first mode flag included in the sequence header syntax element is denoted as seq_st_flag (sequence header secondary transform flag). The sequence header syntax element is used to control the use of the secondary transform mode (ST), enhanced secondary transform mode (EST), and decoder-side enhanced secondary transform mode (DEST) in the current video sequence. For the values of the first mode flag in the sequence header syntax element and the corresponding semantic content under different values, please refer to Table 4a below.
[0231] Table 4a
[0232] The logic for decoding the sequence header flags can be referred to the following table 4a-1:
[0233] Table 4a-1
[0234] Among them, v indicates that variable-length coding is used.
[0235] The sequence header secondary transform flag (seq_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of the ST mode; a value greater than '0' allows the use of the ST mode. The value of SeqStFlag is equal to the value of seq_st_flag. If seq_st_flag does not exist in the bitstream, the value of SeqStFlag is 0.
[0236] In another implementation, the target syntax element is a picture header syntax element (pic_flag, picture header flag). The first mode flag included in the picture header syntax element is denoted as pic_st_flag (picture header secondary transform flag). The picture header syntax element is used to control the use of secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST for the current picture pair. For the values of the first mode flag in the picture header syntax element and the corresponding semantic content under different values, please refer to Table 4b below.
[0237] Table 4b
[0238] For the decoding logic of the above-mentioned image header flag, please refer to the following table 4b-1:
[0239] Table 4b-1
[0240] Among them, v indicates that variable-length coding is used.
[0241] The picture header secondary transform flag (pic_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of ST mode; a value greater than '0' allows the use of ST mode. The value of PicStFlag is equal to the value of pic_st_flag. If pic_st_flag does not exist in the bitstream, the value of PicStFlag is 0.
[0242] In another implementation, the target syntax element is a slice header syntax element (slice_flag), and the first mode flag included in the picture header syntax element is denoted as slice_st_flag (slice header secondary transform flag). The slice header syntax element is used to control whether to prohibit the use of secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST in the current slice. For the values of the first mode flag in the slice header syntax element and the corresponding semantic content under different values, please refer to Table 4c below.
[0243] Table 4c
[0244] For the decoding logic of the above slice flag, please refer to the following table 4c-1:
[0245] Table 4c-1
[0246] Among them, v indicates that variable-length coding is used.
[0247] The slice header secondary transform flag (slice_st_flag) indicates the secondary transform mode. A value of '0' prohibits the use of the ST mode; a value greater than '0' allows the use of the ST mode. The value of SliceStFlag is equal to the value of slice_st_flag. If slice_st_flag does not exist in the bitstream, the value of SliceStFlag is 0.
[0248] As can be seen, when the target syntax element includes only one mode flag, the value of that one mode flag can be used to control the use of multiple types of secondary transform modes, making the corresponding instructions more concise. Furthermore, compared to Method 2, Method 1 offers more diverse control methods, and compared to Method 2, Method 1 offers greater flexibility in controlling the use of various secondary transform modes.
[0249] (2) The target syntax element includes a reference mode flag and a third mode flag (dest_flag), and the reference mode flag is the first mode flag (st_flag) or the second mode flag (est_flag). When the target syntax element includes two mode flags (st_flag and dest_flag, or est_flag and dest_flag), the reference mode flag (st_flag or est_flag) is used to control the current object's use of the secondary transform mode ST and the enhanced secondary transform mode EST, and the third mode flag is used to control the current object's use of the decoder-side enhanced secondary transform mode DEST. The use of multiple types of secondary transform modes can be jointly controlled by the reference mode flag and the third mode flag. Based on this, for different mode flags included in the target syntax element, under different value combinations, the target syntax element has different control functions, that is, corresponds to different semantic contents. Specifically, it may include the following (2.1) and (2.3).
[0250] (2.1) If the value of the reference mode flag is the first value, the target syntax element is used to indicate that the current object prohibits the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST.
[0251] In a specific implementation, when the reference mode flag takes a first value (e.g., "0"), it indicates that the reference mode flag is used to indicate that the current object prohibits the use of the secondary transform mode (ST) and the enhanced secondary transform mode (EST). Based on the relationship between the ST mode, EST mode, and DEST mode, if the ST mode and EST mode are prohibited, the decoder-side enhanced secondary transform mode (DEST) cannot be used either. Therefore, the target syntax element is used not only to indicate that the current object prohibits the use of the ST mode and EST mode, but also to indicate that the current object prohibits the use of the DEST mode.
[0252] In addition, it should be noted that when the reference mode flag takes the first numerical value (for example, the numerical value "0"), since both the ST mode and the EST mode are prohibited from use, no matter how the third mode flag used to control the use of the DEST mode is set, the DEST mode is prohibited from use. Therefore, in order to avoid wasting unnecessary decoding resources, the third mode flag in this application can take the value of the target character (for example, "x") to indicate that the third mode flag does not need to be decoded, that is, when the reference mode flag takes the first numerical value, there is no need to decode the third mode flag. For any mode flag that takes the value of the target character, the target syntax element, in addition to controlling the use of some transformation modes, is also used to indicate that the mode flag does not need to be decoded.
[0253] (2.2) If the value of the reference mode flag is the second numerical value and the value of the third mode flag is the first numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding-side enhanced secondary transform mode DEST.
[0254] In a specific implementation, the reference mode flag takes the second value (e.g., "1"), indicating that the reference mode flag is used to indicate that the current object allows the use of the secondary transform mode (ST) and the enhanced secondary transform mode (EST). The third mode flag takes the first value (e.g., "0"), indicating that the third mode flag is used to indicate that the current object prohibits the use of the decoder-side enhanced secondary transform mode (DEST). Therefore, with the two mode flags acting together, the target syntax element is used to indicate that the current object allows the use of the ST mode and the EST mode, but prohibits the use of the DEST mode.
[0255] (2.3) If the value of the reference mode flag is the second numerical value and the value of the third mode flag is the second numerical value, the target syntax element is used to indicate that the current object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-side enhanced secondary transform mode DEST.
[0256] In a specific implementation, when both the reference mode flag and the third mode flag are set to the second value (e.g., "1"), the reference mode flag indicates that the current object allows the use of the secondary transform mode (ST) and the enhanced secondary transform mode (EST), and the third mode flag indicates that the current object allows the use of the decoder-side enhanced secondary transform mode (DEST). With these two mode flags acting together, the target syntax element indicates that the current object allows the use of the ST mode, the EST mode, and the DEST mode.
[0257] For the several situations shown in (2.1)-(2.3) above, the corresponding syntax elements are substituted in for summary below. In one implementation, the target syntax element is a sequence header syntax element (seq_flag, i.e., sequence header flag), and the reference mode flag included in the sequence header syntax element can be used to control the use of the secondary transform mode ST and the enhanced secondary transform mode EST by the current video sequence. When the reference mode flag is the first mode flag, the reference mode flag can be recorded as seq_st_flag (sequence header secondary transform flag), and when the reference mode flag is the second mode flag, the reference mode flag can be recorded as seq_est_flag (sequence header enhanced secondary transform flag). The third mode flag included in the sequence header syntax element is used to indicate the use of the decoding-end enhanced secondary transform mode DEST by the current video sequence. In this case, the third mode flag can be recorded as seq_dest_flag (sequence header decoding-end enhanced secondary transform flag). Based on this, taking the reference mode flag as the first mode flag as an example, the values and corresponding semantic contents of the two mode flags shown in Table 5a below are provided.
[0258] Table 5a
[0259] The logic for decoding the sequence header flags can be found in Table 5a-1 below:
[0260] Table 5a-1
[0261] The Sequence Head Secondary Transformation Flag (seq_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of SeqStFlag is equal to the value of seq_st_flag. If seq_st_flag does not exist in the bitstream, the value of SeqStFlag is 0.
[0262] The sequence header decoding enhanced secondary transform flag (seq_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SeqDestFlag is equal to the value of seq_dest_flag. If seq_dest_flag is not present in the bitstream, the value of SeqDestFlag is 0.
[0263] In another implementation, the target syntax element is a picture header syntax element (pic_flag, i.e., picture header flag), and the reference mode flag included in the picture header syntax element can be used to control the use of the secondary transform mode ST and the enhanced secondary transform mode EST for the current picture. When the reference mode flag is the first mode flag, the reference mode flag can be recorded as pic_st_flag (picture header secondary transform flag), and when the reference mode flag is the second mode flag, the reference mode flag can be recorded as pic_est_flag (picture header enhanced secondary transform flag). The third mode flag included in the picture header syntax element is used to indicate the use of the decoding-side enhanced secondary transform mode DEST for the current picture. In this case, the third mode flag can be recorded as pic_dest_flag (picture header decoding-side secondary transform flag). Based on this, taking the reference mode flag as the first mode flag (i.e., pic_st_flag) as an example, the values and corresponding semantic contents of the two mode flags shown in Table 5b below are provided.
[0264] Table 5b
[0265] For the decoding logic of the above-mentioned image header flag, please refer to the following table 5b-1:
[0266] Table 5b-1
[0267] The picture header secondary transform flag (pic_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of PicStFlag is equal to the value of pic_st_flag. If pic_st_flag is not present in the bitstream, the value of PicStFlag is 0.
[0268] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0269] It can be understood that if the reference mode flag is the second mode flag (i.e., pic_est_flag), its corresponding value and semantics are the same as those in Table 5b, and the decoding logic for pic_est_flag and pic dest_flag included in the picture header flag is similar to the logic shown in Table 5b-1 above, as shown in Table 5b-2 below, that is:
[0270] Table 5b-2
[0271] The picture header enhanced secondary transform flag (pic_est_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of PicStFlag is equal to the value of pic_est_flag. If pic_est_flag is not present in the bitstream, the value of PicStFlag is 0.
[0272] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0273] In another implementation, the target syntax element is a slice header syntax element (slice_flag, i.e., slice flag), and the reference mode flag included in the slice header syntax element can be used to control the use of the secondary transform mode ST and the enhanced secondary transform mode EST by the current slice. When the reference mode flag is the first mode flag, the reference mode flag can be recorded as slice_st_flag (slice header secondary transform flag), and when the reference mode flag is the second mode flag, the reference mode flag can be recorded as slice_est_flag (slice header enhanced secondary transform flag). The third mode flag included in the slice header syntax element is used to indicate the use of the decoding-side enhanced secondary transform mode DEST by the current slice. In this case, the third mode flag can be recorded as slice_dest_flag (slice header decoding-side enhanced secondary transform flag). Based on this, taking the reference mode flag as the first mode flag as an example, the values and corresponding semantic contents of the two mode flags shown in Table 5c below are provided.
[0274] Table 5c
[0275] The decoding logic of the slice flags can be referred to the following table 5c-1:
[0276] Table 5c-1
[0277] The slice header secondary transform flag (slice_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of SliceStFlag is equal to the value of slice_st_flag. If slice_st_flag does not exist in the bitstream, the value of SliceStFlag is 0.
[0278] The slice header decoder enhanced secondary transform flag (slice_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SliceDestFlag is equal to the value of slice_dest_flag. If slice_dest_flag is not present in the bitstream, the value of SliceDestFlag is 0.
[0279] It should be noted that the third mode flag (dest_flag) included in the various syntax elements shown above takes the value of character x, indicating that the third mode flag does not need to be decoded. This is because when the ST mode or EST mode is not allowed, the DEST mode is also not allowed. Even if the value of the third mode flag is the second numerical value to allow the DEST mode to be used, the DEST mode cannot be truly used. Therefore, in order to avoid unnecessary decoding processing and save decoding resources, the value of the third mode flag can be set to a target character, thereby explicitly indicating that the third mode flag does not need to be decoded.
[0280] (3) The target syntax element includes a first mode flag (st_flag), a second mode flag (est_flag), and a third mode flag (dest_flag). Under the premise that the target syntax element includes three mode flags, each mode flag can be used to control the current object's use of a type of secondary transform mode. That is, the first mode flag is used to control the current object's use of the secondary transform mode ST, the second mode flag is used to control the current object's use of the enhanced secondary transform mode EST, and the third mode flag is used to control the current object's use of the decoding-end enhanced secondary transform mode DEST. Based on this, under the joint action of the three mode flags, based on the different value combinations of each mode flag, the target syntax element has different control functions. For details, please refer to the contents shown in (3.1)-(3.4) below.
[0281] (3.1) If the value of the first mode flag is the first value, the target syntax element is used to indicate that the current object prohibits the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST.
[0282] In a specific implementation, the first mode flag takes a first numerical value (e.g., "0"). The first mode flag has the following control function: the first mode flag is used to indicate that the current object prohibits the use of the secondary transform mode (ST). Based on the relationship between the ST mode, the EST mode, and the DEST mode, when the secondary transform mode (ST) is prohibited, the enhanced secondary transform mode (EST) is also prohibited, and thus the enhanced secondary transform mode (DEST) on the decoding side is also prohibited. Therefore, the target syntax element is used to indicate that the current object prohibits the use of the ST mode, the EST mode, and the DEST mode.
[0283] It should be noted that when the first mode flag takes the first value (e.g., the value "0"), since neither the EST mode nor the DEST mode can be used, in order to avoid wasting resources decoding other mode flags, the second mode flag for controlling the use of the EST mode and the third mode flag for controlling the use of the DEST mode can both take the target character (e.g., "x") to indicate that neither the second mode flag nor the third mode flag needs to be decoded. That is, when the first mode flag takes the first value, the target syntax element, in addition to controlling the use of some transform modes, is also used to indicate that the second mode flag and the third mode flag do not need to be decoded.
[0284] (3.2) If the value of the first mode flag is the second numerical value and the value of the second mode flag is the first numerical value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding-side enhanced secondary transform mode DEST.
[0285] In a specific implementation, when the value of the first mode flag is a second numerical value (for example, a numerical value of "1"), the first mode flag is used to indicate that the current object is allowed to use the secondary transform mode ST. Based on the relationship between the ST mode, the EST mode, and the DEST mode, when the secondary transform mode ST is allowed to be used, whether the enhanced secondary transform mode EST and the decoding-end enhanced secondary transform mode DEST are allowed to be used can be flexibly controlled by the value of the corresponding mode flag. If the value of the second mode flag is a first numerical value (for example, a numerical value of "0"), then the second mode flag is used to indicate that the current object is prohibited from using the enhanced secondary transform mode EST, that is, the enhanced secondary transform mode EST is prohibited from being used, thereby directly determining that the decoding-end enhanced secondary transform mode DEST is prohibited from being used. Therefore, the target syntax element is used to indicate that the current object is allowed to use the ST mode, but is prohibited from using the EST mode and the DEST mode.
[0286] In addition, it should be noted that when the second mode flag takes the first value (e.g., the value "0"), since the DEST mode cannot be used when the EST mode is prohibited, to avoid wasting resources decoding other mode flags, the third mode flag can take the target character (e.g., "x") to indicate that the third mode flag does not need to be decoded. In other words, when the first mode flag takes the second value (e.g., the value "1") and the second mode flag takes the first value (e.g., the value "0"), the target syntax element, in addition to controlling the use of some transform modes, is also used to indicate that the third mode flag does not need to be decoded.
[0287] (3.3) If the first mode flag and the second mode flag both take the second value, and the third mode flag takes the first value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding-side enhanced secondary transform mode DEST.
[0288] In a specific implementation, when the value of the first mode flag is the second numerical value (for example, the numerical value "1"), the first mode flag is used to indicate that the current object is allowed to use the secondary transform mode ST. When the value of the second mode flag is the second numerical value (for example, the numerical value "1"), the second mode flag is used to indicate that the current object is allowed to use the enhanced secondary transform mode EST. Based on the relationship between the ST mode, the EST mode and the DEST mode, when both the ST mode and the EST mode are allowed to be used, whether the DEST mode is allowed to be used can be flexibly controlled by the value of the third mode flag. If the value of the second mode flag is the first numerical value (for example, the numerical value "0"), then the third mode flag is used to indicate that the current object is prohibited from using the decoding-end enhanced secondary transform mode DEST. Combining the control functions of the above three mode flags, the target syntax element is used to indicate that the current object is allowed to use the ST mode and the EST mode, but is prohibited from using the DEST mode.
[0289] (3.4) If the first mode flag and the second mode flag both take the second numerical value, and the third mode flag takes the second numerical value, the target syntax element is used to indicate that the current object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-side enhanced secondary transform mode DEST.
[0290] Similar to the above (3.3), the values of the first mode flag and the second mode flag are both the second numerical value (for example, the numerical value "1"), then these two mode flags work together to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST. Whether the DEST mode is allowed to be used can be flexibly controlled by the value of the third mode flag. When the value of the third mode flag is also the second numerical value (for example, the numerical value "1"), the third mode flag is used to indicate that the current object allows the use of the decoding-side enhanced secondary transform mode DEST. Thus, through the values of the above three mode flags, the target syntax element is used to indicate that the current object allows the use of the ST mode, EST mode and DEST mode.
[0291] For the several situations shown in (3.1)-(3.4) above, different syntax elements are used for explanation below. In one implementation, the target syntax element is a sequence header syntax element (seq_flag). The first mode flag included in the sequence header syntax element is used to control the use of the secondary transform mode ST by the current video sequence. In this case, the first mode flag can be recorded as seq_st_flag (i.e., the sequence header secondary transform flag); the second mode flag included in the sequence header syntax element is used to control the use of the enhanced secondary transform mode EST by the current video sequence. In this case, the second mode flag can be recorded as seq_est_flag (i.e., the sequence header enhanced secondary transform flag). The third mode flag included in the sequence header syntax element is used to control the use of the decoding-end enhanced secondary transform mode DEST by the current video sequence. In this case, the third mode flag can be recorded as seq_dest_flag (i.e., the sequence header decoding-end enhanced secondary transform flag). Based on this, different values and corresponding semantic contents of the three mode flags are provided as shown in Table 6a below.
[0292] Table 6a
[0293] For the decoding logic of the above sequence header flag, please refer to the content in the following Table 6a-1:
[0294] Table 6a-1
[0295] The Sequence Head Secondary Transformation Flag (seq_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of SeqStFlag is equal to the value of seq_st_flag. If seq_st_flag does not exist in the bitstream, the value of SeqStFlag is 0.
[0296] The Sequence Header Enhanced Secondary Transformation Flag (eq_est_flag) is a binary variable. A value of '1' indicates that the EST mode is enabled; a value of '0' indicates that the EST mode is disabled. The value of SeqEstFlag is equal to the value of seq_est_flag. If seq_est_flag is not present in the bitstream, the value of SeqEstFlag is 0.
[0297] The sequence header decoding enhanced secondary transform flag (seq_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SeqDestFlag is equal to the value of seq_dest_flag. If seq_dest_flag is not present in the bitstream, the value of SeqDestFlag is 0.
[0298] In another implementation, the target syntax element is a picture header syntax element (pic_flag, i.e., picture header flag). The first mode flag included in the picture header syntax element is used to control the use of the secondary transform mode ST for the current picture. In this case, the first mode flag can be recorded as pic_st_flag (i.e., picture header secondary transform flag). The second mode flag included in the picture header syntax element is used to control the use of the enhanced secondary transform mode EST for the current picture. In this case, the second mode flag can be recorded as pic_est_flag (i.e., picture header enhanced secondary transform flag). The third mode flag included in the picture header syntax element is used to control the use of the decoder-side enhanced secondary transform mode DEST for the current picture. In this case, the third mode flag can be recorded as pic_dest_flag (i.e., picture header decoder-side enhanced secondary transform flag). Based on this, three mode flags with different values and specific semantic contents are provided as shown in Table 6b below.
[0299] Table 6b
[0300] For the decoding logic of the above picture header flags, please refer to the following table 6b-1:
[0301] Table 6b-1
[0302] The picture header secondary transform flag (pic_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of PicStFlag is equal to the value of pic_st_flag. If pic_st_flag is not present in the bitstream, the value of PicStFlag is 0.
[0303] The picture header enhanced secondary transform flag (eq_est_flag) is a binary variable. A value of '1' indicates that the EST mode is enabled; a value of '0' indicates that the EST mode is disabled. The value of PicEstFlag is equal to the value of pic_est_flag. If pic_est_flag is not present in the bitstream, the value of PicEstFlag is 0.
[0304] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0305] In another implementation, the target syntax element is a slice header syntax element (slice_flag, i.e., slice flag). The first mode flag included in the slice header syntax element is used to control the use of the secondary transform mode ST by the current slice. In this case, the first mode flag can be recorded as slice_st_flag (i.e., slice header secondary transform flag); the second mode flag included in the slice header syntax element is used to control the use of the enhanced secondary transform mode EST by the current video slice. In this case, the second mode flag can be recorded as slice_est_flag (i.e., slice header enhanced secondary transform flag). The third mode flag included in the slice header syntax element is used to control the use of the decoding-side enhanced secondary transform mode DEST by the current slice. In this case, the third mode flag can be recorded as slice_dest_flag (i.e., slice header decoding-side enhanced secondary transform flag). Based on this, three mode flags are provided as shown in Table 6c below with different values and specific semantic contents.
[0306] Table 6c
[0307] For the decoding logic of the above slice flags, please refer to the following table 6c-1:
[0308] Table 6c-1
[0309] The slice header secondary transform flag (slice_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of SliceStFlag is equal to the value of slice_st_flag. If slice_st_flag does not exist in the bitstream, the value of SliceStFlag is 0.
[0310] The slice header enhanced secondary transform flag (eq_est_flag) is a binary variable. A value of '1' indicates that the EST mode is enabled; a value of '0' indicates that the EST mode is disabled. The value of SliceEstFlag is equal to the value of slice_est_flag. If slice_est_flag is not present in the bitstream, the value of SliceEstFlag is 0.
[0311] The slice header decoder enhanced secondary transform flag (slice_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SliceDestFlag is equal to the value of slice_dest_flag. If slice_dest_flag is not present in the bitstream, the value of SliceDestFlag is 0.
[0312] It can be seen that in the above exemplary table, when any syntax element includes multiple mode flags, the use of various types of secondary transformation modes can be more finely controlled based on the values of the multiple mode flags, thereby further improving the flexibility of the use of various types of secondary transformation modes.
[0313] In one embodiment, when the transform mode syntax control information includes multiple syntax elements and the low-level syntax elements only include one mode flag, the decoding of the low-level syntax elements depends on the high-level syntax elements. The high and low levels of the syntax elements here are relative and are determined based on the level of the syntax structure to which the syntax elements act. For example, the transform mode syntax control information includes sequence header syntax elements and picture header syntax elements. The sequence header syntax elements act on the current video sequence and the syntax structure to which they act is the video sequence. The picture header syntax elements act on the current picture and the syntax structure to which they act is the picture. Since the level of the video sequence is higher than the level of the picture, the sequence header syntax elements are high-level syntax elements, while the picture header syntax elements are low-level syntax elements.
[0314] The following takes the example of transform mode syntax control information including two syntax elements to exemplify the dependency between syntax elements in syntax structures at different levels. The transform mode syntax control information includes a first syntax element and a second syntax element; the level of the syntax structure to which the second syntax element acts is higher than the level of the syntax structure to which the first syntax element acts. For example, the second syntax element is a sequence header syntax element (seq_flag), and the first syntax element is a picture header syntax element (pic_flag). For another example, the second syntax element is a picture header syntax element (pic_flag), and the first syntax element is a slice header syntax element (slice_flag). For another example, the second syntax element is a sequence header syntax element (seq_flag), and the first syntax element is a slice header syntax element (slice_flag).
[0315] The first syntax element includes a third mode flag, and the third mode flag in the first syntax element is used to control the first object's use of the decoder-side enhanced secondary transform mode DEST. Specifically, the different values of the third mode flag can be used to indicate whether the first object is prohibited from using or allowed to use the DEST mode. The second syntax element includes one or more of the first mode flag, the second mode flag, and the third mode flag; the first mode flag in the second syntax element is used to control the second object's use of the decoder-side enhanced secondary transform mode ST; the second mode flag in the second syntax element is used to control the second object's use of the enhanced secondary transform mode EST; and the third mode flag in the second syntax element is used to control the second object's use of the secondary transform mode DEST.
[0316] If the first syntax element is a picture header syntax element, the first object is the current picture, the second syntax element is a sequence header syntax element, and the second object is the current video sequence. If the first syntax element is a slice header syntax element, the first object is the current slice, the second syntax element is a picture header syntax element, and the second object is the current picture. If the first syntax element is a slice header syntax element, the first object is the current slice, the second syntax element is a sequence header syntax element, and the second object is the current video sequence. As can be seen from the above, the level of the syntax structure corresponding to the second object is higher than the level of the syntax structure corresponding to the first object.
[0317] In the case where the first syntax element includes only one mode flag (here, the third mode flag), the third mode flag in the first syntax element depends on any mode flag in the second syntax element. In one implementation, if the second syntax element includes only one mode flag (such as any one of the first mode flag, the second mode flag, and the third mode flag), then the third mode flag in the first syntax element may depend on the mode flag included in the second syntax element. In another implementation, if the second syntax element includes two or more mode flags (such as the first mode flag and the second mode flag, or the first mode flag and the third mode flag, or the second mode flag and the third mode flag, or the first mode flag, the second mode flag, and the third mode flag), then the third mode flag in the first syntax element may depend on any one of the multiple mode flags included in the second syntax element (such as the first mode flag or the second mode flag).
[0318] Regarding the dependency between the mode flags of the first syntax element and the second syntax element, there may be several cases as shown in the following I-III.
[0319] I. The second syntax element includes the first mode flag, and the third mode flag in the first syntax element depends on the first mode flag in the second syntax element, including the following:
[0320] (I-1) If the value of the first mode flag in the second syntax element is the first value, the transform mode syntax control information is used to indicate that the first object prohibits the use of the secondary transform mode ST, enhanced secondary transform mode EST and decoding-side enhanced secondary transform mode DEST.
[0321] Specifically, if the value of the first mode flag in the second syntax element is a first value (e.g., a value of "0"), the second syntax element is used to indicate that the second object prohibits the use of the secondary transform mode ST. Because the syntax structure corresponding to the second object is at a higher level than the syntax structure corresponding to the first object, if the second object is prohibited from using the ST mode, the first object is also prohibited from using the secondary transform mode ST, enhanced secondary transform mode EST, and decoder-side enhanced secondary transform mode DEST. Therefore, this transform mode syntax control information can be used to indicate that the first object prohibits the use of the ST mode, EST mode, and DEST mode.
[0322] It should be noted that if the ST mode is prohibited in a higher-level syntax element, the lower-level syntax structure is not allowed to use the ST mode, EST mode, and DEST mode, regardless of how the mode flag in the lower-level syntax element is set. To avoid unnecessary decoding of the third mode flag and save decoding resources, the value of the third mode flag can be set to a target character (e.g., the character "x"), so that the target syntax element can also be used to indicate that the third mode flag does not need to be decoded.
[0323] (I-2) If the value of the first mode flag in the second syntax element is the second numerical value, and the value of the third mode flag in the first syntax element is the first numerical value, the transform mode syntax control information is used to indicate that the first object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding-side enhanced secondary transform mode DEST.
[0324] Specifically, if the value of the first mode flag in the second syntax element is a second numerical value (for example, a numerical value of "1"), then the second syntax element is used to indicate that the second object is allowed to use the ST mode. In the case where the second object at a higher level is allowed to use the ST mode, the first object at a lower level may be allowed to use the ST mode, EST mode, and DEST mode, and whether the DEST mode is ultimately allowed to be used can be determined by the value of the third mode flag in the first syntax element. In this way, the third mode flag in the first syntax element takes the first numerical value (for example, a numerical value of "0"), which indicates that the first object is prohibited from using the DEST mode. Therefore, under the joint action of the first syntax element and the second syntax element, the transformation mode syntax control information can be used to indicate that the first object is allowed to use the ST mode and EST mode, but is prohibited from using the DEST mode.
[0325] (I-3) If the value of the first mode flag in the second syntax element is the second numerical value, and the value of the third mode flag in the first syntax element is the second numerical value, the transform mode syntax control information is used to indicate that the first object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding-end enhanced secondary transform mode DEST.
[0326] Specifically, if the value of the first mode flag in the second syntax element is a second numerical value (for example, a numerical value of "1"), then the second syntax element is used to indicate that the second object is allowed to use the ST mode. In the case where the second object at a higher level is allowed to use the ST mode, the first object at a lower level may be allowed to use the ST mode, EST mode, and DEST mode, and whether the DEST mode is ultimately allowed to be used can be determined by the value of the third mode flag in the first syntax element. In this way, the third mode flag in the first syntax element takes the second numerical value (for example, a numerical value of "1"), which indicates that the first object is allowed to use the DEST mode. Thus, under the joint action of the first syntax element and the second syntax element, the transformation mode syntax control information can be used to indicate that the first object is allowed to use the ST mode, EST mode, and DEST mode.
[0327] For better understanding, the values of the corresponding mode flags in different syntax elements shown in Tables 7a, 7b, and 7c, as well as the corresponding semantic content, are provided below. In one implementation, if the first syntax element is a picture header syntax element (pic_flag), the second syntax element is a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is denoted as pic_dest_flag, and the first mode flag included in the second syntax element is denoted as seq_st_flag. The pic_dest_flag in the picture header syntax element depends on the seq_st_flag in the sequence header syntax element, and may include decoding of the pic_dest_flag that depends on the seq_st_flag. According to the contents described in (I-1) to (I-3) above, the contents shown in Table 7a below correspond to the contents.
[0328] Table 7a
[0329] The decoding logic for the first and second syntax elements may include the following contents in Table 7a-1 and Table 7a-2:
[0330] In the sequence header:
[0331] Table 7a-1
[0332] The sequence header secondary transform flag (seq_st_flag) indicates the secondary transform mode. A value of '1' allows the use of the ST mode; a value of '0' prohibits the use of the ST mode. The value of SeqStFlag is equal to the value of seq_st_flag. If seq_st_flag does not exist in the bitstream, the value of SeqStFlag is 0.
[0333] In the image header:
[0334] Table 7a-2
[0335] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0336] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element is a picture header syntax element (pic_flag). The third mode flag included in the first syntax element is denoted as slice_dest_flag, and the first mode flag included in the second syntax element is denoted as pic_st_flag. The slice_dest_flag in the slice header syntax element depends on the pic_st_flag in the picture header syntax element, and may include decoding of the slice_dest_flag that depends on the pic_st_flag. According to the contents described in (I-1)-(I-3) above, the contents shown in Table 7b below correspond to the contents.
[0337] Table 7b
[0338] For the decoding logic of the first syntax element and the second syntax element, please refer to the following Table 7b-1 and Table 7b-2:
[0339] In the image header:
[0340] Table 7b-1
[0341] The picture header secondary transform flag (pic_st_flag) is a binary variable. A value of '1' indicates that ST mode is enabled; a value of '0' indicates that ST mode is disabled. The value of PicStFlag is equal to the value of pic_st_flag. If pic_st_flag is not present in the bitstream, the value of PicStFlag is 0.
[0342] In the opening credits:
[0343] Table 7b-2
[0344] The image header decoder enhanced secondary transform flag (slice_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SliceDestFlag is equal to the value of slice_dest_flag. If slice_dest_flag is not present in the bitstream, the value of SliceDestFlag is 0.
[0345] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element may also be a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is recorded as slice_dest_flag, and the first mode flag included in the second syntax element is recorded as seq_st_flag. The slice_dest_flag in the slice header syntax element depends on the seq_st_flag in the sequence header syntax element, and may include decoding of the slice_dest_flag that depends on the seq_st_flag. According to the contents described in (I-1) to (I-3) above, the contents shown in Table 7c below correspond to the contents.
[0346] Table 7c
[0347] As can be seen in the example table above, if the first mode flag (st_flag) in the second syntax element of the higher level takes the value "0," then the first object of the higher level prohibits the use of ST mode. The second object of the lower level cannot use ST mode, and thus cannot use EST mode, nor can it use DEST mode. For example, when the first mode flag seq_st_flag in the sequence header syntax element takes the value "0," this first mode flag is used to indicate that ST mode is prohibited for the current video sequence. Therefore, ST mode, EST mode, and DEST mode cannot be used for the current picture and current slice. In other words, if the prohibition of ST mode is controlled through syntax elements at the sequence level, then ST mode, EST mode, and DEST mode cannot be used at levels lower than the sequence level (such as the picture level and slice level). Therefore, the third mode flag pic_dest_flag in the picture header syntax element does not need to be decoded, and the third mode flag slice_dest_flag in the slice header syntax element does not need to be decoded.
[0348] II. The second syntax element includes the second mode flag. The third mode flag in the first syntax element depends on the second mode flag in the second syntax element and may include the following content:
[0349] (II-1) If the value of the second mode flag in the second syntax element is the first value, the transform mode syntax control information is used to indicate that the first object prohibits the use of the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST.
[0350] Specifically, if the value of the second mode flag in the second syntax element is a first numerical value (e.g., a numerical value of "0"), the second syntax element is used to indicate that the second object is prohibited from using the EST mode. Since the hierarchy of the syntax structure corresponding to the second object is higher than the hierarchy of the syntax structure corresponding to the first object, if the second object is prohibited from using the EST mode, the second object is not allowed to use the EST mode, and the first object is also not allowed to use the DEST mode. Therefore, under the combined effect of the first syntax element and the second syntax element, the transform mode syntax control information can be used to indicate that the first object is prohibited from using both the EST mode and the DEST mode.
[0351] It should be noted that if the EST mode is prohibited in a higher-level syntax element, the lower-level syntax structure is not allowed to use the EST mode or the DEST mode, regardless of how the mode flag in the lower-level syntax element is set. To avoid unnecessary decoding of the third mode flag and save decoding resources, the value of the third mode flag can be set to a target character (e.g., the character "x"), so that the target syntax element can also be used to indicate that the third mode flag does not need to be decoded.
[0352] (II-2) If the value of the second mode flag in the second syntax element is the second numerical value and the value of the third mode flag in the first syntax element is the first numerical value, the transform mode syntax control information is used to indicate that the first object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding-side enhanced secondary transform mode DEST.
[0353] Specifically, if the value of the second mode flag in the second syntax element is a second numerical value (for example, a numerical value of "1"), then the second syntax element is used to indicate that the second object is allowed to use the EST mode. And when the EST mode is allowed, it means that the second object is allowed to use the ST mode. Therefore, when the second object at a higher level is allowed to use the ST mode and the EST mode, the first object at a lower level may be allowed to use the ST mode, the EST mode, and the DEST mode, and whether the DEST mode is ultimately allowed to be used can be determined by the value of the third mode flag in the first syntax element. In this way, the third mode flag in the first syntax element takes the first numerical value (for example, a numerical value of "0"), which indicates that the first object is prohibited from using the DEST mode, so that under the joint action of the first syntax element and the second syntax element, the transformation mode syntax control information can be used to indicate that the first object is allowed to use the ST mode and the EST mode, but is prohibited from using the DEST mode.
[0354] (II-3) If the value of the second mode flag in the second syntax element is the second numerical value, and the value of the third mode flag in the first syntax element is the second numerical value, the transform mode syntax control information is used to indicate that the first object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding-end enhanced secondary transform mode DEST.
[0355] Similar to (II-2) above, if the value of the second mode flag in the second syntax element is a second numerical value (e.g., a numerical value of "1"), then the second syntax element is used to indicate that the second object is allowed to use the EST mode, thereby allowing the first object to use the ST mode, EST mode, and DEST mode. In this manner, the third mode flag in the first syntax element is a second numerical value (e.g., a numerical value of "1"), indicating that the first object is allowed to use the DEST mode. Therefore, under the combined effect of the first and second syntax elements, the transform mode syntax control information can be used to indicate that the first object is allowed to use the ST mode, EST mode, and DEST mode.
[0356] Based on the contents described in (II-1)-(II-3) above, the values of corresponding mode flags in different syntax elements and the corresponding semantic contents can be provided as shown in the following Tables 8a, 8b and 8c.
[0357] In one implementation, if the first syntax element is a picture header syntax element (pic_flag), the second syntax element is a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is denoted as pic_dest_flag, and the second mode flag included in the second syntax element is denoted as seq_est_flag. The pic_dest_flag in the picture header syntax element depends on the seq_est_flag in the sequence header syntax element, and may include decoding of the pic_dest_flag that depends on the seq_est_flag. According to the contents described in (II-1)-(II-3) above, the contents shown in Table 8a below correspond to the contents.
[0358] Table 8a
[0359] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element is a picture header syntax element (pic_flag). The third mode flag included in the first syntax element is denoted as slice_dest_flag, and the second mode flag included in the second syntax element is denoted as pic_est_flag. The slice_dest_flag in the slice header syntax element depends on the pic_est_flag in the picture header syntax element, and may include decoding of the slice_dest_flag depending on the pic_est_flag. According to the contents described in (II-1) to (II-3) above, the contents shown in Table 8b below correspond to the contents.
[0360] Table 8b
[0361] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element may also be a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is recorded as slice_dest_flag, and the second mode flag included in the second syntax element is recorded as seq_est_flag. The slice_dest_flag in the slice header syntax element depends on the seq_est_flag in the sequence header syntax element, and may include decoding of the slice_dest_flag depending on the seq_est_flag. According to the contents described in (II-1) to (II-3) above, the contents shown in the following Table 8c correspond.
[0362] Table 8c
[0363] III. The second syntax element includes a third mode flag. The third mode flag in the first syntax element depends on the third mode flag in the second syntax element and may include the following content:
[0364] (III-1) If the value of the third mode flag in the second syntax element is the first value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the decoding-side enhanced secondary transform mode DEST.
[0365] Specifically, if the value of the third mode flag in the second syntax element is a first value (e.g., a value of "0"), the second syntax element is used to indicate that the second object is prohibited from using the DEST mode. Because the hierarchy of the syntax structure corresponding to the second object is higher than the hierarchy of the syntax structure corresponding to the first object, if the second object is prohibited from using the DEST mode, the first object is also prohibited from using the DEST mode. Therefore, this transform mode syntax control information can be used to indicate that the first object is prohibited from using the DEST mode.
[0366] It should be noted that if the syntax element at a higher level prohibits the use of the DEST mode, then regardless of how the mode flag in the syntax element at a lower level is set, the syntax structure at the lower level is not allowed to use the DEST mode. To avoid unnecessary decoding of the third mode flag and save decoding resources, the value of the third mode flag can be set to a target character (e.g., the character "x"), so that the target syntax element can also be used to indicate that the third mode flag does not need to be decoded.
[0367] (III-2) If the value of the third mode flag in the second syntax element is the second numerical value, and the value of the third mode flag in the first syntax element is the first numerical value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the decoding-side enhanced secondary transform mode DEST.
[0368] Specifically, if the value of the second mode flag in the second syntax element is a second numerical value (for example, a numerical value of "1"), then the second syntax element is used to indicate that the second object is allowed to use the DEST mode. Therefore, when the second object at a higher level is allowed to use the DEST mode, the first object at a lower level may be allowed to use the DEST mode, and whether the DEST mode is ultimately allowed to be used can be determined by the value of the third mode flag in the first syntax element. In this way, the third mode flag in the first syntax element takes the first numerical value (for example, a numerical value of "0"), which indicates that the first object is prohibited from using the DEST mode. Therefore, under the joint action of the first syntax element and the second syntax element, the transformation mode syntax control information can be used to indicate that the first object is prohibited from using the DEST mode.
[0369] (III-3) If the value of the third mode flag in the second syntax element is the second numerical value, and the value of the third mode flag in the first syntax element is the second numerical value, then the transform mode syntax control information is used to indicate that the first object allows the use of the decoding-side enhanced secondary transform mode DEST.
[0370] Similar to (III-2) above, if the value of the third mode flag in the second syntax element is the second numerical value (e.g., the numerical value "1"), then the third mode flag in the second syntax element is used to indicate that the second object allows the use of DEST mode. If the third mode flag in the first syntax element is the second numerical value (e.g., the numerical value "1"), then it indicates that the first object allows the use of DEST mode. Thus, under the combined effect of the first and second syntax elements, the transform mode syntax control information can be used to indicate that the first object allows the use of DEST mode.
[0371] Based on the contents described in (III-1)-(III-3) above, the values of corresponding mode flags in different syntax elements and the corresponding semantic contents can be provided as shown in the following Tables 9a, 9b and 9c.
[0372] In one implementation, if the first syntax element is a picture header syntax element (pic_flag), the second syntax element is a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is denoted as pic_dest_flag, and the third mode flag included in the second syntax element is denoted as seq_dest_flag. The pic_dest_flag in the picture header syntax element depends on the seq_dest_flag in the sequence header syntax element. Specifically, the decoding of the pic_dest_flag may depend on the seq_dest_flag. According to the contents described in (III-1)-(III-3) above, the contents shown in Table 9a below correspond to the contents.
[0373] Table 9a
[0374] For the decoding logic of the first and second syntax elements, please refer to the following Table 9a-1 and Table 9a-2:
[0375] In the sequence header:
[0376] Table 9a-1
[0377] The sequence header decoding enhanced secondary transform flag (seq_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SeqDestFlag is equal to the value of seq_dest_flag. If seq_dest_flag is not present in the bitstream, the value of SeqDestFlag is 0.
[0378] In the image header:
[0379] Table 9a-2
[0380] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0381] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element is a picture header syntax element (pic_flag). The third mode flag included in the first syntax element is denoted as slice_dest_flag, and the third mode flag included in the second syntax element is denoted as pic_dest_flag. The slice_dest_flag in the slice header syntax element depends on the pic_dest_flag in the picture header syntax element, and may include decoding of the slice_dest_flag that depends on the pic_dest_flag. According to the contents described in (III-1) to (III-3) above, the contents shown in Table 9b below correspond to the contents.
[0382] Table 9b
[0383] For the decoding logic of the first and second syntax elements, please refer to the following Table 9b-1 and Table 9b-2:
[0384] In the image header:
[0385] Table 9b-1
[0386] The picture header decoder enhanced secondary transform flag (pic_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of PicDestFlag is equal to the value of pic_dest_flag. If pic_dest_flag is not present in the bitstream, the value of PicDestFlag is 0.
[0387] In the opening credits:
[0388] Table 9b-2
[0389] The slice header decoder enhanced secondary transform flag (slice_dest_flag) is a binary variable. A value of '1' indicates that the DEST mode is enabled; a value of '0' indicates that the DEST mode is disabled. The value of SliceDestFlag is equal to the value of slice_dest_flag. If slice_dest_flag is not present in the bitstream, the value of SliceDestFlag is 0.
[0390] In another implementation, if the first syntax element is a slice header syntax element (slice_flag), the second syntax element may also be a sequence header syntax element (seq_flag). The third mode flag included in the first syntax element is recorded as slice_dest_flag, and the third mode flag included in the second syntax element is recorded as seq_dest_flag. The slice_dest_flag in the slice header syntax element depends on the seq_est_flag in the sequence header syntax element, and specifically may include that the decoding of the slice_dest_flag depends on the seq_est_flag. According to the contents described in (III-1)-(III-3) above, the contents shown in the following Table 9c correspond.
[0391] Table 9c
[0392] As can be seen, in the above methods I-III, when the first syntax element at the lower level includes only a mode flag, the control of the use of the DEST mode will also depend on the value of the mode flag used to control the use of the secondary transform mode in the second syntax element at the higher level. This control method can control the use of the DEST mode in the lower-level syntax structure in the higher-level syntax element, thereby determining whether to apply the DEST mode after decoding less content, which is beneficial to improving coding performance.
[0393] S302: Decode the encoded code stream.
[0394] In one specific implementation, the process of decoding a coded bitstream includes: first, entropy decoding the coded bitstream to obtain mode information (such as prediction mode) used in various encoding processes and quantized transform coefficients. The quantized transform coefficients are then sequentially dequantized and inverse transformed to obtain residual data. A prediction mode is determined based on the mode information and then predicted to obtain predicted data. Reconstructed data is obtained based on the residual data and the predicted data. The inverse transform process involves the use of several transform modes, so at this stage, transform mode syntax control information in the coded bitstream can be used to determine which transform modes to use.
[0395] S303: During the decoding process, control the use of the transform mode according to the transform mode syntax control information.
[0396] In a specific implementation, based on the multiple types of transform modes, during the decoding process, the transform mode syntax control information can be used to determine the transform mode that can be used, and the current object can be transformed according to the determined transform mode. For example, if the transform mode includes a secondary transform mode, if the transform mode syntax control information is used to indicate that the DEST mode is allowed, then the DEST mode can be used in the intra CU.
[0397] The following various implementations are described on the premise that the transform mode includes a secondary transform mode, and the transform mode syntax control information is used to control the use of the secondary transform mode.
[0398] In one implementation, secondary transform modes include, but are not limited to, ST mode, EST mode, and DEST mode. If the transform mode syntax control information includes a single syntax element, such as a sequence header syntax element, a picture header syntax element, or a slice header syntax element, then when controlling the use of the secondary transform mode according to the transform mode syntax control information, the syntax structure to which the syntax element applies, as well as syntax structures at a lower level than the syntax structure, can be controlled based on the indication of the syntax element. For example, if the transform mode syntax control information includes only a sequence header syntax element, and the sequence header syntax element indicates that the current video sequence is permitted to use ST mode, EST mode, and DEST mode, then the ST mode, EST mode, and DEST mode are permitted for pictures, slices, coding units, and transform units. In one implementation, while these secondary transform modes are permitted, whether they are actually used during decoding can be further determined based on other information about the current object being decoded.
[0399] In another implementation, if the transform mode syntax control information includes multiple syntax elements, based on the dependency between the syntax elements, taking the transform mode syntax control information including the first syntax element and the second syntax element as an example, when executing the above S303, the consumer device can execute the following steps 1.1 to 1.3.
[0400] Step 1.1: Compare the level of a first syntax structure to which a first syntax element applies and the level of a second syntax structure to which a second syntax element applies.
[0401] Specifically, the syntax structure used by each syntax element refers to the syntax structure corresponding to the current object involved in controlling the secondary transform mode. For example, the first syntax element is a picture header syntax element and is used to indicate the use of the secondary transform mode for the current picture. The syntax structure corresponding to the current picture is picture, so the first syntax structure is picture. The second syntax element is a sequence header syntax element and is used to indicate the use of the secondary transform mode for the current video sequence. The syntax structure corresponding to the current video sequence is video sequence, so the second syntax structure is video sequence.
[0402] The consumer device can compare the hierarchies of the first and second syntax structures to determine the higher-level syntax. For example, if the first syntax structure is an image and the second syntax structure is a video sequence, a hierarchical comparison of the syntax structures can determine that the second syntax structure is the higher-level syntax structure. Based on the control of the use of the secondary transform mode by the syntax elements corresponding to the higher-level syntax structure, the decoding of the syntax elements corresponding to the lower-level syntax structure is determined, or the use of the secondary transform is controlled according to the syntax elements corresponding to the lower-level syntax structure. See steps 1.2 and 1.3 below.
[0403] Step 1.2: If the syntax element corresponding to the higher-level one of the first and second syntax structures indicates prohibition of the secondary transform mode, there is no need to decode the syntax element corresponding to the lower-level one of the first and second syntax structures.
[0404] In a specific implementation, if the level of the first grammatical structure is higher than that of the second grammatical structure, the first grammatical structure is the one with the higher level, the first grammatical element is the grammatical element corresponding to the one with the higher level, and the second grammatical element is the grammatical element corresponding to the one with the lower level. If the level of the second grammatical structure is higher than that of the first grammatical structure, the second grammatical element is the one with the higher level, the second grammatical element is the grammatical element corresponding to the one with the higher level, and the first grammatical element is the grammatical element corresponding to the one with the lower level.
[0405] If the syntax element corresponding to a higher level indicates that the secondary transform mode (e.g., DEST mode) is prohibited, then based on the dependency relationship between syntax elements of different syntax structures, it is not necessary to decode the syntax element corresponding to the lower level. For example, if the first syntax element is a picture header syntax element (pic_flag) and the second syntax element is a sequence header syntax element (seq_flag), if the sequence header syntax element indicates that the DEST mode is prohibited, it is not necessary to decode the picture header syntax element.
[0406] Step 1.3: If the syntax element corresponding to the higher level of the first syntax structure and the second syntax structure indicates that the secondary transform mode is allowed, control the use of the secondary transform mode according to the syntax element corresponding to the lower level of the first syntax structure and the second syntax structure.
[0407] In a specific implementation, when the syntax element corresponding to the higher level indicates that the secondary transform mode (for example, the DEST mode) is allowed to be used, the syntax element corresponding to the lower level can be decoded, so as to control the use of the secondary transform mode according to the syntax element corresponding to the lower level. The syntax element corresponding to the lower level can be used to indicate the type of secondary transform mode that is allowed to be used and / or the type of secondary transform mode that is prohibited to be used. For the indication of the syntax elements, please refer to the various feasible implementation methods under the aforementioned target syntax element, which will not be repeated here. For example, the first syntax element is the picture header syntax element (pic_flag), and the second syntax element is the sequence header syntax element. When the sequence header syntax element indicates that the DEST mode is allowed to be used, the picture header syntax element can be decoded, and when the picture header syntax element indicates that the DEST mode is allowed to be used, the secondary transform processing can be performed according to the DEST mode.
[0408] It can be seen that if the syntax element corresponding to the higher-level layer indicates that the secondary transform mode (such as DEST mode) is prohibited, there is no need to decode the syntax element corresponding to the lower-level layer, and the lower-level layer will not use the secondary transform mode. If the higher-level layer indicates that the secondary transform mode is allowed, then the syntax element of the lower level layer can determine which type of secondary transform mode to use based on the value of the mode flag included in the syntax element. For example, if the sequence header syntax element is used to indicate that the DEST mode is allowed, the picture header syntax element can be decoded, and the use of the DEST mode for the picture can be determined based on the decoded picture header syntax element.
[0409] It will be appreciated that if the transform mode syntax control information includes two or more syntax elements, the syntax structure hierarchies affected by the syntax elements can be sorted from highest to lowest. Thus: ① If a syntax element at a higher syntax level indicates that the secondary transform mode is prohibited, there is no need to decode syntax elements at lower syntax levels. For example, the transform mode syntax control information includes sequence header syntax elements, picture header syntax elements, and slice header syntax elements. If the highest-level syntax element is the sequence header syntax element, and this sequence header syntax element indicates that the DEST mode is not permitted, then the picture header syntax elements, slice header syntax elements, and so on do not need to be decoded, and it can be determined that the DEST mode is not used for both the picture and the slice. If the sequence header syntax element indicates that the DEST mode is permitted, then the picture header syntax element can be decoded. If the picture header syntax element also indicates that the DEST mode is not permitted, then there is no need to decode the slice header syntax element. ② If a syntax element at a higher syntax level indicates that the secondary transform mode is prohibited, the use of the secondary transform can be controlled according to the syntax elements at lower levels. For example, continuing with the above example, when the sequence header syntax element is used to indicate that the DEST mode is allowed, the picture header syntax element can be decoded, and the use of the DEST mode conversion mode can be controlled according to the picture header syntax element. Specifically, if the picture header syntax element indicates that the DEST mode is allowed, the slice header syntax element can be decoded, and the use of the DEST mode conversion mode can be controlled according to the slice header syntax element.
[0410] In summary, the above steps 1.2 and 1.3 are configured based on the decoding order of each syntax structure in the coded bitstream and the priority of syntax elements. Based on the decoding order of each syntax structure in the coded bitstream, the syntax elements within the corresponding syntax structure also have a decoding priority: sequence header > picture header > slice header > coding unit > transform unit. When a syntax element with a higher decoding priority indicates that the secondary transform mode (e.g., DEST) is not allowed, the next level does not need to decode the syntax elements related to the secondary transform mode (e.g., DEST). The priority of syntax elements is transform unit > coding unit > slice header > picture header > sequence header. When a syntax element with a higher priority indicates that the secondary transform mode (e.g., DEST) is not allowed, the secondary transform mode (e.g., DEST) is not used, even if a syntax element with a lower priority indicates that the secondary transform mode (e.g., DEST) is allowed. This configuration allows the secondary transform mode to be controlled at a higher level, reducing unnecessary decoding processing and improving decoding efficiency.
[0411] In one embodiment, the transformation mode includes a secondary transformation mode, and the transformation mode syntax control information is used to control the use of the secondary transformation mode. Based on this, the consumer device may further perform the following steps 2.1 to 2.3.
[0412] Step 2.1: During the decoding process, determine the current image in the coded bitstream; the current image refers to the image being decoded.
[0413] Step 2.2: Get the image type of the current image.
[0414] Step 2.3: If the image type of the current image is an I-frame image, trigger the execution of the step of controlling the use of the transformation mode according to the transformation mode syntax control information.
[0415] In a specific implementation, based on the decoding order of the encoded code stream, a video sequence can be decoded first. The video sequence includes multiple frames of images, and the frame of image being decoded is the current image. Each frame of image belongs to an image type, and the image type can include I-frame images, B-frame images, and P-frame images. The image type of the current image can be one of I-frame images, B-frame images, and P-frame images. For example, the image type of the current image is an I-frame image. Since the I-frame image uses intra-frame prediction, and the secondary transform mode is applied to the intra-frame prediction frame, better encoding performance can be achieved. Therefore, when the image type of the current image is an I-frame image or other types of images, it can be decided whether the secondary transform mode is used.
[0416] ① If the image type of the current image is an I-frame image, indicating that the current image uses the intra-frame prediction mode, then the use of the secondary transform mode can be controlled according to the transform mode syntax control information. The transform mode syntax control information can be decoded to obtain the syntax elements included therein, and the current image can be subjected to secondary transform processing according to the secondary transform mode allowed for use indicated by the syntax elements, such as using the DEST mode to perform secondary transform processing on the current image. ② If the image type of the current image is a P-frame image or a B-frame image. Since P-frame images and B-frame images mostly use the inter-frame prediction mode for prediction processing and are not suitable for the secondary transform mode, no secondary transform is required, and there is no need to trigger the execution of the steps of controlling the use of the transform mode according to the transform mode syntax control information, and there is no need to decode the transform mode syntax control information to control the use of the secondary transform mode.
[0417] Steps 2.1 through 2.3 above determine whether the transform mode syntax control information is decoded and applied based on the picture type. Specifically, the transform mode syntax control information for controlling the use of the secondary transform mode is decoded only in I-frame pictures, thereby performing the corresponding secondary transform process. For example, the syntax element dest_flag indicating the use of the DEST mode only needs to be decoded in I-frame pictures.
[0418] In one embodiment, the transform mode includes a secondary transform mode, and the transform mode syntax control information is used to control the use of the secondary transform mode. Based on this, when controlling the use of the transform mode according to the transform mode syntax control information, the consumer device may perform the following: if the transform mode syntax control information indicates that the decoder-side enhanced secondary transform mode DEST is allowed to be used for the current object, then: ① determine the decoder-side enhanced secondary transform mode DEST as the target secondary transform mode to be used for the current object; or ② determine the target secondary transform mode to be used for the current object based on reference information of the current object; where the current object refers to a transform unit or coding unit in the coded bitstream undergoing decoding processing.
[0419] In other words, whether the current object uses the DEST mode can be determined based on the transform mode syntax control information. Since the transform mode syntax control information is high-level syntax information used to control the use of the secondary transform mode, it can be used to control the use of the secondary transform mode by a video sequence, picture, or slice. Control over the use of the secondary transform mode is transitive between different syntax structures. This transitivity means that if a higher-level syntax structure allows the use of the secondary transform mode, the lower-level syntax structure can flexibly control the use of the secondary transform mode using its own syntax elements. If a higher-level syntax structure does not allow the use of the secondary transform mode, the lower-level syntax structure also does not allow the use of the secondary transform mode. Based on this transitivity, whether a higher-level syntax structure allows the use of the ST mode, EST mode, or DEST mode can directly affect the use of the DEST transform mode by a coding unit or a transform unit. For example, if the transform mode syntax control information indicates that the current picture or slice allows the use of the DEST mode, it can be determined that the current object (i.e., the coding unit or the current transform unit) allows the use of the DEST mode. If the current picture does not allow the use of the ST mode or the DEST mode, it can be determined that the current object (i.e., the coding unit or the current transform unit) does not allow the use of the DEST mode.
[0420] If the DEST mode is determined to be permitted for the current object, in one implementation, the DEST mode can be directly determined as the target secondary transformation mode to be used by the current object, thereby performing DEST-related processing on the current object according to the DEST mode. For DEST-related processing, please refer to the aforementioned description of the DEST mode and will not be further elaborated here.
[0421] It's worth noting that if the transform mode syntax control information indicates that DEST mode is prohibited, then the current object does not need to be processed using DEST. Furthermore, if the current object contains control information indicating the use of DEST, this information does not need to be decoded. For example, if the current object contains a flag explicitly indicating that DEST mode is permitted, this flag will not be decoded during actual processing, eliminating the need for a secondary transform on the current object.
[0422] In another implementation, a target secondary transformation mode to be used for the current object may be determined based on the reference information of the current object, thereby determining whether to perform DEST processing. If the DEST mode is determined to be the target secondary transformation mode to be used based on the reference information of the current object, then DEST-related processing may be performed on the current object. Conversely, if the DEST mode is determined not to be the target secondary transformation mode to be used based on the reference information of the current object, then DEST-related processing is not performed on the current object.
[0423] The reference information of the current object may be some attribute information of the current object. Optionally, the reference information includes at least one of the following: the size of the current object, the prediction mode used by the current object, and the partitioning method used by the current object. When determining the target secondary transform mode to be used for the current object based on the reference information, any of the following ①-③ may be included:
[0424] ① If the size of the current object is within the preset size range, the decoding end enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used for the current object.
[0425] Specifically, the preset size range is the size range supported by the DEST mode. The size includes width and height, and being in the preset size range means that the width is in the preset width range and the height is in the preset height range. For example, the preset size range is: 4×4 to 64×64. If the size of the current object is 8×8, since the width 8 is in the preset width range [4, 64] and the height 8 is in the preset height range [4, 64], then it can be said that the current object can use the DEST mode, so that the DEST mode can be determined as the required target secondary transformation mode, so that the current object can be subjected to secondary transformation processing according to the DEST mode. Conversely, if the size of the current object is not in the preset size range, that is, the width is not in the preset width range or the height is not in the preset height range, then it can be determined that the current object does not use the DEST mode, and the current object does not perform DEST related processing.
[0426] ② If the prediction mode used by the current object is an intra-frame prediction mode of a specified type, the decoding end enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object.
[0427] There are multiple types of intra-frame prediction modes, such as DC mode (suitable for large flat areas), angular prediction mode, and Plane mode (including horizontal and vertical prediction). A specified type of intra-frame prediction mode can include one or more, and one or more intra-frame prediction modes can be selected from the multiple types of intra-frame prediction modes as the specified type of intra-frame prediction mode. For example, DC mode and angular prediction mode are selected as the specified type of intra-frame prediction mode.
[0428] If the intra-frame prediction mode used by the current object is the specified type of intra-frame prediction mode, it means that the current object performs intra-frame prediction processing, so the DEST mode can be determined as the target secondary transform mode to be used by the current object. Conversely, if the prediction mode used by the current object is another type of intra-frame prediction mode (i.e., an intra-frame prediction mode other than the specified type of intra-frame prediction mode among multiple intra-frame prediction modes) or an inter-frame prediction mode, then it can be determined that the current object does not use the DEST mode.
[0429] ③ If the partitioning method used by the current object is a non-derivative tree partitioning method, the decoding end enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object.
[0430] Since the ST mode and other modes can be applied to non-DT (derivative tree) intra-frame coded coding units, if the partitioning method used by the current object is non-DT, the DEST mode can be applied to the current object, that is, the target second transform mode to be used by the current object is the DEST mode. Conversely, if the partitioning method used by the current object is DT, it can be determined that the DEST mode is not used for the current object.
[0431] In one embodiment, the transform mode syntax control information is also used to control the use of transform information required for the transform mode. The transform mode here includes not only the secondary transform mode, but also other transform modes, including but not limited to: position-based transform mode PBT and sub-block transform mode SBT. That is, the transform mode syntax control information is not only used to indicate the use of transform information required for the secondary transform mode, but also to indicate the use of transform information required for other transform modes. The transform information includes one or more of a transform kernel and a transform type. The transform kernel here includes a separable transform kernel and an inseparable transform kernel. For example, DCT-II and DST-VII are separable transform kernels.
[0432] For example, the transform mode syntax control information can indicate the transform kernel used in the DEST mode, or the transform kernel used in the ST mode, such as an 8x8 ST transform kernel. It can also be used to indicate the DEST transform type to be used (such as horizontal flip, vertical flip, or diagonal flip). For another example, based on the transform types allowed in the PBT mode, which include DCT8 and DST7, the transform mode syntax control information can also be used to indicate that one or more of DCT8 and DST7 are used in the PBT mode. The transform types allowed in the SBT mode include DCT8, DST7, and DCT-2. The transform mode syntax control information can also be used to indicate that one or more of DCT8, DST7, and DCT-2 are used in the PBT mode. Based on the transform mode information controlled by the transform mode syntax control information, one or more of the following can be determined based on the transform mode syntax control information: the ST transform kernel to be used, the DEST transform kernel, the transform type of the DEST transform to be used, the transform type of the SBT transform to be used, the transform type of the PBT transform, and the transform type of the PBT transform.
[0433] The data processing method provided in the embodiment of the present application can flexibly control the use of transform modes, including secondary transform modes, position-based transform modes, sub-block transform modes, etc., during the decoding process based on transform mode syntax control information by decoding the encoded code stream. Based on the flexible control of the transform mode, the decoding process can be made more efficient, thereby improving the decoding efficiency. In the specific implementation method, the use of various types of secondary transform modes is mainly controlled by adding control information such as syntax elements to the high-level syntax structure, and based on the different numbers of mode flags included in the syntax elements, a rich control method is designed to further enhance the flexibility of the control of the use of the secondary transform mode.
[0434] Please refer to Figure 5, which is a flow chart of a data processing method provided in an embodiment of the present application. The data processing method can be executed by the service device 201 in the data processing system. The method includes the following steps S501-S503.
[0435] S501, determining syntax control information of a transformation mode to be adopted by a video; the transformation mode syntax control information is used to indicate the use of a transformation mode.
[0436] In one implementation, a service device may acquire a video and determine transform mode syntax control information for the video. Transform modes include, but are not limited to, position-based transform (PBT), sub-block transform (SBT), and secondary transform mode. For example, when the video type is screen content video, transform mode syntax control information may be set. This transform mode syntax control information can be used to flexibly control the enabling and disabling of the secondary transform mode in higher-level syntax structures such as the sequence level and the image level. For example, the DEST mode may be enabled for intra-predicted frames but disabled for inter-predicted frames.
[0437] S502: Encode the video.
[0438] The service device can encode the video. For a specific encoding process, see Figure 1a. During the encoding process, based on the setting of the transform mode syntax control information, if the transform mode syntax control information is used to control the secondary transform mode, then during the transform phase, the corresponding type of secondary transform mode (e.g., DEST mode) to be used by the coding unit can be determined based on the indication of the transform mode syntax control information, and the coding unit can then be secondary transformed according to the determined secondary transform mode.
[0439] S503 , during the encoding process, encoding the transform mode syntax control information into the encoded bitstream of the video.
[0440] In one implementation, the transform mode includes a secondary transform mode, and the transform mode syntax control information includes a plurality of syntax elements. When encoding the transform mode syntax control information into the coded bitstream of the video, the service device may specifically: determine the syntax structure (e.g., sequence, image, slice) of the current coding object (i.e., the object being coded); and add syntax elements to the corresponding syntax structure, thereby placing the syntax elements in the transform mode syntax control information in the correct syntax structure. For example, the service device may add a sequence header syntax element to the sequence header, an image header syntax element to the image header, and a slice header syntax element to the slice header. By encoding the transform mode syntax control information into the coded bitstream, a coded bitstream containing the transform mode syntax control information may be obtained, so that the decoding end can control the use of the transform mode based on the transform mode syntax control information.
[0441] The data processing method provided by the embodiments of the present application determines transform mode syntax control information for a video and encodes this transform mode syntax control information into the coded bitstream, thereby accurately instructing the decoder on the use of transform modes. Furthermore, the setting of this transform mode syntax control information takes into account the corresponding scene of the video, enabling adaptive control of the use of transform modes based on different scenes, thereby improving encoding performance.
[0442] Please refer to Figure 6a, which is a schematic diagram of the structure of another data processing device provided in an embodiment of the present application. The data processing device can be set in the consumer device provided in an embodiment of the present application. The consumer device can be the service device mentioned in the above method embodiment. The data processing device shown in Figure 6a can be a computer program (including program code) running on the consumer device. The data processing device can be used to perform some or all of the steps in the method embodiment shown in Figure 3. Please refer to Figure 6a. The data processing device may include the following units:
[0443] The acquisition unit 601 is configured to acquire a video encoding stream; the encoding stream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode;
[0444] The processing unit 602 is configured to decode the encoded code stream;
[0445] The processing unit 602 is further configured to control the use of the transform mode according to the transform mode syntax control information during the decoding process.
[0446] Please refer to Figure 6b, which is a schematic diagram of the structure of another data processing device provided in an embodiment of the present application. The data processing device can be set in the service device provided in an embodiment of the present application, and the service device can be the service device mentioned in the above method embodiment. The data processing device shown in Figure 6b can be a computer program (including program code) running in the service device. The data processing device can be used to execute some or all of the steps in the method embodiment shown in Figure 5. Please refer to Figure 6b, the data processing device can include the following units:
[0447] The determining unit 611 is configured to determine the syntax control information of the transform mode to be used for the video; the syntax control information of the transform mode is used to control the use of the transform mode;
[0448] A processing unit 612 is configured to perform encoding processing on the video;
[0449] The processing unit 612 is configured to encode the transform mode syntax control information into the encoded bitstream of the video during the encoding process.
[0450] It is understood that the specific functions of the various units of the data processing device described in the embodiments of the present application can be specifically implemented according to the methods in the above-mentioned method embodiments. The specific implementation process can refer to the relevant description of the above-mentioned method embodiments and will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated.
[0451] Next, the consumer device and service device provided in the embodiments of the present application are described.
[0452] An exemplary embodiment of the present application further provides a schematic diagram of the structure of a computer device, as shown in Figure 7 . The computer device may include a processor 701, an input device 702, an output device 703, and a memory 704. The processor 701, input device 702, output device 703, and memory 704 are connected via a bus. The memory 704 is used to store a computer program, which includes program instructions. The processor 701 is used to execute the program instructions stored in the memory 704.
[0453] In one embodiment, the computer device may be the aforementioned consumer device; in this embodiment, the processor 701 executes the following operations by running the executable program code in the memory 704:
[0454] Obtaining a video encoding stream; the encoding stream includes transform mode syntax control information, the transform mode syntax control information being used to control the use of the transform mode;
[0455] Decoding the encoded code stream;
[0456] During the decoding process, use of the transform mode is controlled according to the transform mode syntax control information.
[0457] In another embodiment, the computer device may be the aforementioned service device; in this embodiment, the processor 701 executes the following operations by running the executable program code in the memory 704:
[0458] Determining syntax control information of a transform mode to be used for the video; the transform mode syntax control information is used to control the use of the transform mode;
[0459] performing encoding processing on the video;
[0460] During the encoding process, the transform mode syntax control information is encoded into the encoded bitstream of the video.
[0461] It should be understood that the computer device described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments above, and can also execute the description of the data processing device in the corresponding embodiments above, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here.
[0462] In addition, it should be pointed out here that: an embodiment of the present application also provides a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the method in the embodiments corresponding to Figures 3 and 5 above. Therefore, it will not be repeated here.
[0463] According to one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, so that the computer device can perform the methods described in the embodiments corresponding to FIG. 3 and FIG. 5 . Therefore, detailed description thereof will not be given here.
[0464] An embodiment of the present application provides a method for processing a video stream. The video stream is generated according to the data processing method of the embodiment corresponding to FIG. 3 , or is decoded based on the data processing method of the embodiment corresponding to FIG. 5 .
[0465] An embodiment of the present application also provides a computer-readable storage medium, which stores a video code stream formed by a computer program. When the computer program is executed by a processor, it can execute the method in the embodiments corresponding to Figures 3 and 5 above. Therefore, it will not be repeated here.
[0466] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0467] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire a coded bitstream of a video; the coded bitstream includes transform mode syntax control information, and the transform mode syntax control information is used to control the use of the transform mode; Decoding the encoded code stream; During the decoding process, use of the transform mode is controlled according to the transform mode syntax control information.
2. The method according to claim 1, characterized in that The transform mode syntax control information is set in the syntax structure of the coded bitstream; the syntax structure includes at least one of the following: a video sequence, a picture, a slice, a coding unit and a transform unit; If the transform mode includes a secondary transform mode, the transform mode syntax control information includes at least one syntax element, each syntax element is used to control the use of the secondary transform mode by the current object; the current object refers to the syntax structure that is being decoded in the encoded bitstream.
3. The method according to claim 1 or 2, characterized in that If the syntax structure includes a video sequence, the transform mode syntax control information is set in a sequence header of the current video sequence; the transform mode syntax control information is a sequence header syntax element, the current object is the current video sequence, and the sequence header syntax element is used to control the use of the secondary transform mode by the current video sequence; If the syntax structure includes an image, the transform mode syntax control information is set in an image header of the current image; the transform mode syntax control information is an image header syntax element, the current object is the current image, and the image header syntax element is used to control the use of the secondary transform mode by the current image; If the syntax structure includes a slice, the transform mode syntax control information is set in the slice header of the current slice; the transform mode syntax control information is a slice header syntax element, the current object is the current slice, and the slice header syntax element is used to control the use of the secondary transform mode by the current slice; The current video sequence refers to the video sequence that is being decoded in the coded bitstream; the current image refers to the image that is being decoded in the coded bitstream; and the current slice refers to the slice that is being decoded in the coded bitstream.
4. The method according to any one of claims 1 to 3, characterized in that Any one of the syntax elements is represented as a target syntax element; the target syntax element includes a first mode flag; If the value of the first mode flag is a first value, the target syntax element is used to indicate that the current object is prohibited from using the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the second value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the third value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the fourth value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding-side enhanced secondary transform mode DEST.
5. The method according to any one of claims 1 to 4, characterized in that Any one of the syntax elements is represented as a target syntax element; the target syntax element includes a first mode flag; If the value of the first mode flag is a first value, the target syntax element is used to indicate that the current object is prohibited from using the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the second value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the third value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding-side enhanced secondary transform mode DEST.
6. The method according to any one of claims 1 to 5, characterized in that Any one of the syntax elements is represented as a target syntax element; the target syntax element includes a reference mode flag and a third mode flag; The reference mode flag is used to control the current object to use the secondary transform mode ST and the enhanced secondary transform mode EST, and the third mode flag is used to control the current object to use the decoding end enhanced secondary transform mode DEST; the reference mode flag is the first mode flag or the second mode flag; If the value of the reference mode flag is the first value, the target syntax element is used to indicate that the current object is prohibited from using the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the reference mode flag is the second value, and the value of the third mode flag is the first value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding end enhanced secondary transform mode DEST; If the value of the reference mode flag is the second numerical value, and the value of the third mode flag is the second numerical value, the target syntax element is used to indicate that the current object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-side enhanced secondary transform mode DEST.
7. The method according to any one of claims 1 to 6, characterized in that Any one of the syntax elements is represented as a target syntax element; the target syntax element includes a first mode flag, a second mode flag and a third mode flag; the first mode flag is used to control the use of the secondary transform mode ST by the current object, the second mode flag is used to control the use of the enhanced secondary transform mode EST by the current object, and the third mode flag is used to control the use of the enhanced secondary transform mode DEST by the current object at the decoding end; If the value of the first mode flag is a first value, the target syntax element is used to indicate that the current object is prohibited from using the secondary transform mode ST, the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the first mode flag is the second value, and the value of the second mode flag is the first value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST, but prohibits the use of the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the first mode flag and the second mode flag both take the second value, and the third mode flag takes the first value, the target syntax element is used to indicate that the current object allows the use of the secondary transform mode ST and the enhanced secondary transform mode EST, but prohibits the use of the decoding end enhanced secondary transform mode DEST; If the first mode flag and the second mode flag both take the second numerical value, and the third mode flag takes the second numerical value, then the target syntax element is used to indicate that the current object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-end enhanced secondary transform mode DEST.
8. The method according to any one of claims 1 to 7, characterized in that The transformation mode syntax control information includes a first syntax element and a second syntax element; the level of the syntax structure to which the second syntax element acts is higher than the level of the syntax structure to which the first syntax element acts; The first syntax element includes a third mode flag, and the third mode flag in the first syntax element is used to control the first object to use the decoding end enhanced secondary transform mode DEST; The second syntax element includes one or more of a first mode flag, a second mode flag, and a third mode flag; the first mode flag in the second syntax element is used to control the use of the second object for the decoding end enhanced secondary transform mode ST; the second mode flag in the second syntax element is used to control the use of the second object for the enhanced secondary transform mode EST; the third mode flag in the second syntax element is used to control the use of the second object for the secondary transform mode DEST; A third mode flag in the first syntax element is dependent on any mode flag in the second syntax element; Among them, if the first syntax element is a picture header syntax element, the first object is the current picture, the second syntax element is a sequence header syntax element, and the second object is the current video sequence; if the first syntax element is a slice header syntax element, the first object is the current slice, the second syntax element is a picture header syntax element, and the second object is the current picture.
9. The method according to any one of claims 1 to 8, characterized in that The second syntax element includes a first mode flag, and a third mode flag in the first syntax element depends on the first mode flag in the second syntax element, including: If the value of the first mode flag in the second syntax element is a first value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the secondary transform mode ST, the enhanced secondary transform mode EST and the decoder enhanced secondary transform mode DEST; If the value of the first mode flag in the second syntax element is the second value, and the value of the third mode flag in the first syntax element is the first value, the transform mode syntax control information is used to indicate that the first object is allowed to use the secondary transform mode ST and the enhanced secondary transform mode EST, but is prohibited from using the decoder enhanced secondary transform mode DEST; If the value of the first mode flag in the second syntax element is a second numerical value and the value of the third mode flag in the first syntax element is a second numerical value, then the transform mode syntax control information is used to indicate that the first object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-end enhanced secondary transform mode DEST.
10. The method according to any one of claims 1 to 9, characterized in that The second syntax element includes a second mode flag, and the third mode flag in the first syntax element depends on the second mode flag in the second syntax element, including: If the value of the second mode flag in the second syntax element is a first value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the enhanced secondary transform mode EST and the decoding end enhanced secondary transform mode DEST; If the value of the second mode flag in the second syntax element is a second value, and the value of the third mode flag in the first syntax element is a first value, the transform mode syntax control information is used to indicate that the first object is allowed to use the secondary transform mode ST and the enhanced secondary transform mode EST, but is prohibited from using the decoding end enhanced secondary transform mode DEST; If the value of the second mode flag in the second syntax element is a second numerical value and the value of the third mode flag in the first syntax element is a second numerical value, then the transform mode syntax control information is used to indicate that the first object allows the use of secondary transform mode ST, enhanced secondary transform mode EST and decoding-end enhanced secondary transform mode DEST.
11. The method according to any one of claims 1 to 10, characterized in that The second syntax element includes a third mode flag, and the third mode flag in the first syntax element depends on the third mode flag in the second syntax element, including: If the value of the third mode flag in the second syntax element is the first value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the decoder-side enhanced secondary transform mode DEST; If the value of the third mode flag in the second syntax element is the second value, and the value of the third mode flag in the first syntax element is the first value, the transform mode syntax control information is used to indicate that the first object is prohibited from using the decoder enhanced secondary transform mode DEST; If the value of the third mode flag in the second syntax element is a second numerical value and the value of the third mode flag in the first syntax element is a second numerical value, the transform mode syntax control information is used to indicate that the first object allows the use of the decoding-side enhanced secondary transform mode DEST.
12. The method according to any one of claims 1 to 11, characterized in that The transform mode includes a secondary transform mode, and the transform mode syntax control information is used to control the use of the secondary transform mode; the method further includes: During the decoding process, the current image in the coded bitstream is determined; the current image is the image being decoded. image; Obtaining the image type of the current image; If the image type of the current image is an I-frame image, the step of controlling the use of the transform mode according to the transform mode syntax control information is triggered.
13. The method according to any one of claims 1 to 12, characterized in that The transform mode includes a secondary transform mode, and the transform mode syntax control information is used to control the use of the secondary transform mode; and controlling the use of the transform mode according to the transform mode syntax control information includes: If the transform mode syntax control information indicates that the current object allows the use of the decoder-side enhanced secondary transform mode DEST, the decoder-side enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object; or the target secondary transform mode to be used by the current object is determined according to the reference information of the current object; The current object refers to a transform unit or a coding unit in the coded bitstream that is being decoded.
14. The method according to any one of claims 1 to 13, characterized in that The reference information includes at least one of the following: the size of the current object, the prediction mode used by the current object, and the division method used by the current object; The determining, according to the reference information, a target secondary transformation mode to be used by the current object comprises: If the size of the current object is within a preset size range, the decoding end enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object; or, If the prediction mode used by the current object is an intra-frame prediction mode of a specified type, the decoding end enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object; or, If the partitioning mode used by the current object is a non-derivative tree partitioning mode, the decoding-side enhanced secondary transform mode DEST is determined as the target secondary transform mode to be used by the current object.
15. The method according to any one of claims 1 to 14, characterized in that The transform mode includes a secondary transform mode, and the transform mode syntax control information is used to control the use of the secondary transform mode; the transform mode syntax control information includes a first syntax element and a second syntax element; and controlling the use of the transform mode according to the transform mode syntax control information includes: comparing a level of a first grammatical structure to which the first grammatical element applies and a level of a second grammatical structure to which the second grammatical element applies; If the syntax element corresponding to the higher level one of the first syntax structure and the second syntax structure indicates prohibition of using the secondary transform mode, there is no need to perform decoding processing on the syntax element corresponding to the lower level one of the first syntax structure and the second syntax structure; If the syntax element corresponding to the higher-level one of the first syntax structure and the second syntax structure indicates that the secondary transform mode is allowed, the use of the secondary transform mode is controlled according to the syntax element corresponding to the lower-level one of the first syntax structure and the second syntax structure.
16. The method according to any one of claims 1 to 15, characterized in that The transform mode syntax control information is also used to control the use of transform information required for the transform mode; Among them, the transformation mode includes at least one of the following: a secondary transformation mode, a position-based transformation mode and a sub-block transformation mode; the transformation information includes at least one of the following: a transformation kernel and a transformation type; the transformation kernel includes: a separable transformation kernel or an inseparable transformation kernel.
17. A data processing method, characterized in that: The method comprises: Determine the syntax control information of the transformation mode to be adopted by the video; the syntax control information of the transformation mode is used to control the use of the transformation mode; Performing encoding processing on the video; During the encoding process, the transform mode syntax control information is encoded into the encoded bitstream of the video.
18. A data processing device, characterized in that: include: An acquisition unit, used for acquiring a coded bitstream of a video; the coded bitstream includes transform mode syntax control information, and the transform mode syntax control information is used for controlling the use of the transform mode; A processing unit, used for decoding the encoded code stream; The processing unit is further configured to control the use of the transform mode according to the transform mode syntax control information during the decoding process.
19. A computer device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the data processing method according to any one of claims 1 to 17 is executed.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 17 is executed.
21. A method for processing a video code stream, characterized in that: The video code stream is generated according to the data processing method according to claim 17, or is decoded based on the data processing method according to any one of claims 1-16.
Citation Information
Patent Citations
Video decoding method and device, computer readable medium and electronic equipment
CN112533000A
Conditional use of reduced secondary transform for video processing
CN113841409A
Apparatus, method and computer program for video coding and decoding
CN114424575A
Codec mode dependent selection of transform skip mode
CN116250241A
Methods and Apparatuses for Transform Skip Mode Information Signaling
US20210160479A1