Computer-implemented method for video encoding and decoding and computer-readable storage medium

By adjusting the parameters of the transform skip mode in the sequence parameters of the video encoding standard, and introducing the coded_RU_flag syntax, the lossless compression and hierarchical mapping throughput problems in the prior art are solved, and efficient video encoding and decoding are achieved.

CN119996692AActive Publication Date: 2025-05-13ALIBABA (CHINA) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510386440.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-08-13
Publication Date
2025-05-13
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

When existing video encoding standards deal with large transform blocks, they cannot realize mathematical lossless compression of blocks, and the hierarchical mapping process brings throughput problems to the decoder hardware implementation.

Method used

By moving the log2_transform_skip_max_size_minus2 parameter in the Sequence Parameter Set (SPS), the transform skip mode is allowed to be used within the maximum TB size range and an additional syntax coded_RU_flag is introduced to reduce dependencies between residual cells.

Benefits of technology

Lossless compression of large transform blocks is realized, complexity and resource consumption of decoder implementation, and coding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996692A_ABST
    Figure CN119996692A_ABST
Patent Text Reader

Abstract

A method and apparatus for video processing, comprising: determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; and signaling the maximum transform size in a sequence parameter set (SPS).
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 62 / 899,738, filed on September 12, 2019, and U.S. Provisional Application No. 62 / 904,880, filed on September 24, 2019, both of which are incorporated by reference into this application. Background Art

[0002] A video is a set of static pictures (or "frames") that capture visual information. In order to reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission, and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are currently a variety of video coding formats that use standardized video coding techniques. The most common ones are video coding formats based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, the AVS standard, etc., are developed by standardization organizations. As more and more advanced video coding technologies are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the invention

[0003] Embodiments of the present application provide a video processing method and apparatus. In an exemplary embodiment, a method includes: determining a transform process for skipping a prediction residual based on a maximum transform size of a prediction block; and signaling the maximum transform size in a sequence parameter set (SPS).

[0004] In another embodiment, an apparatus includes a memory configured to store instructions and a processor configured to cause the apparatus to execute the following instructions: determine a transform process to skip a prediction residual based on a maximum transform size of a prediction block; and signal the maximum transform size in a sequence parameter set (SPS).

[0005] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: determining a transform process to skip a prediction residual based on a maximum transform size of a prediction block; and signaling the maximum transform size in a sequence parameter set (SPS).

[0006] In another example embodiment, a method includes: receiving a code stream of a video sequence; determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; and determining to skip a transform process of a prediction residual of the prediction block based on the maximum transform size.

[0007] In another embodiment, a device includes a memory configured to store instructions and a processor configured to cause the device to execute the following instructions: receive a code stream of a video sequence; determine a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; and determine to skip a transform process for a prediction residual of the prediction block based on the maximum transform size.

[0008] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions that can be executed by at least one processor of a device to cause the device to perform a method. The method includes: receiving a code stream of a video sequence; determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; and determining to skip a transform process of a prediction residual of the prediction block based on the maximum transform size. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Embodiments and various aspects of the present application are illustrated in the following detailed description and accompanying drawings.The various features shown in the drawings are not drawn to scale.

[0010] Figure 1 is a schematic diagram of an example video sequence structure according to some embodiments of the present application.

[0011] Figure 2A A schematic diagram showing an example encoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.

[0012] Figure 2B A schematic diagram showing another example encoding process of a hybrid video encoding system consistent with an embodiment of the present application is shown.

[0013] Figure 3A A schematic diagram showing an example decoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.

[0014] Figure 3B A schematic diagram showing another example decoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.

[0015] Figure 4 A block diagram of an example apparatus for encoding or decoding a video according to some embodiments of the present application is shown.

[0016] Figure 5 Table 1 shows an example syntax structure of a sequence parameter set (SPS) according to some embodiments of the present application.

[0017] Figure 6 Table 2 shows an example syntax structure of a sequence parameter set (SPS) according to some embodiments of the present application.

[0018] Figure 7 Table 3 shows an example syntax structure of a transform unit according to some embodiments of the present application.

[0019] Figure 8 An example syntax structure diagram related to transmit packet differential pulse code modulation (BDPCM) mode according to some embodiments of the present application is shown.

[0020] Fig. 9 Table 5 shows another example syntax structure of an SPS according to some embodiments of the present application.

[0021] Fig.10 Table 6 showing another example syntax structure of a transform unit according to some embodiments of the present application is shown.

[0022] Fig.11 is a schematic diagram of an example diagonal scan of a 64×64 transform block (TB) according to some embodiments of the present application.

[0023] Figures 12A-12D An example residual unit (RU) according to some embodiments of the present application is shown.

[0024] Fig.13 is a schematic diagram showing an example of diagonal scanning of a 64×64 TB according to some embodiments of the present application, where the TB is divided into four 32×32 RUs.

[0025] Figures 14A-14D Table 7 shows an example syntax structure diagram for residual coding when a TB is divided into RUs according to some embodiments.

[0026] Figures 15A-15D Table 8 is shown, which shows another example syntax structure for residual decoding according to some embodiments of the present application.

[0027] Fig.16 Table 9 is shown showing example parameter values ​​derived from a chroma format according to some embodiments of the present application.

[0028] Fig.17 Table 10 is shown, which shows an example syntax structure General Video Coding Draft 6 for residual coding performing inverse level mapping according to some embodiments of the present application.

[0029] Fig.18 is a flowchart of an example decoding method according to some embodiments of the present application.

[0030] Fig.19 Table 11 according to some embodiments of the present application is shown, which shows an example syntax structure diagram for residual decoding without performing inverse level mapping.

[0031] Fig. 20 Table 12 is shown, illustrating an example lookup table for selecting Rice parameters, according to some embodiments of the present application.

[0032] Fig.21 A flowchart of an example process for video processing according to some embodiments of the present application is shown.

[0033] Fig. 22 A flowchart of another example process for video processing according to some embodiments of the present application is illustrated. DETAILED DESCRIPTION

[0034] Reference can now be made in detail to example embodiments, examples of which are shown in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different drawings represent the same or similar elements, unless otherwise specified. The embodiments set forth in the following description of example embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with the aspects related to the present application described in the attached claims. Specific aspects of the present application are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0035] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0036] In order to achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding technologies are incorporated into JEM, JEM achieves higher coding performance than HEVC.

[0037] The VVC standard was developed recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0038] A video is a set of static pictures (or "frames") arranged in a time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, computer, smartphone, tablet, video player, or any end-user terminal with display capability) can be used to display such pictures in time sequence. In addition, in some applications, the video capture device can transmit the captured video to a video playback device (e.g., a computer with a monitor) in real time, such as for monitoring, conferencing, or live broadcasting.

[0039] In order to reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The module for compression is generally referred to as an "encoder", and the module for decompression is generally referred to as a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of encoders and decoders can include circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of encoders and decoders can include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress a video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder".

[0040] The video encoding process can identify and retain useful information that can be used to reconstruct the picture, and ignore unimportant information in the reconstruction process. If the ignored, unimportant information cannot be fully reconstructed, the encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off made to reduce the required storage space and transmission bandwidth.

[0041] Useful information about the picture being encoded (referred to as the "current picture") includes changes relative to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, with position changes being of greatest interest. Changes in the position of a group of pixels representing an object can reflect the motion of the object between the reference picture and the current picture.

[0042] A picture that is encoded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is encoded using a previous picture as a reference picture is called a "P-picture." A picture that is encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B-picture."

[0043] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present application is illustrated. The video sequence 100 may be a real-time video or a video that has been captured and archived. The video 100 may be a real-life video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real-life video with an augmented reality effect). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored in a storage device), or a video input interface (e.g., a video broadcast transceiver) to receive video from a video content provider.

[0044] like Figure 1 As shown, video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, and there are more pictures between pictures 106 and 108. Figure 1 , picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, as indicated by the arrow, and its reference picture is picture 102. Picture 106 is a B picture, as indicated by the arrow, and its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be directly before or after the picture. For example, the reference picture of picture 104 may be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are only examples, and the present application does not limit the embodiments of the reference pictures to Figure 1 The example shown in .

[0045] Due to the computational complexity of such tasks, video codecs typically do not encode or decode an entire picture at once. Instead, they may divide the picture into basic segments and encode or decode the picture segment by segment. In this application, these basic segments are referred to as basic processing units ("BPUs"). For example, Figure 1Structure 110 in shows an example structure of a picture (e.g., any one of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in a picture can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit for a picture can be selected based on a balance between coding efficiency and the level of detail to be retained in the basic processing unit.

[0046] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture may include a luma component (Y) representing achromatic luma information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components may have basic processing units of the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0047] Video encoding has several stages of operation, examples of which are given in Figure 2A-2B and Figure 3A-3BAs shown in . For each stage, the size of the basic processing unit may still be too large to be processed, so it can be further divided into segments referred to as "basic processing subunits" in this application. In some embodiments, the basic processing subunit may be referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in certain other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same or smaller size as the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include storage in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to processing needs. It should also be noted that different schemes can be used to divide the basic processing units at different stages.

[0048] For example, in the mode decision phase (an example of which is in Figure 2B As shown in FIG. 1 , the encoder can decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing sub-unit.

[0049] For another example, in the prediction phase (an example of which is Figure 2A-2B ), the encoder can perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further split the basic processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which prediction operations can be performed.

[0050] For another example, in the transformation phase (an example of which is in Figure 2A-2B), the encoder can perform transform operations on the residual basic processing sub-units (e.g., CUs). However, in some cases, these basic processing sub-units may still be too large to process. The encoder can further divide the basic processing sub-units into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and transform operations can be performed at the level of the segments. It should be noted that the division scheme of the same basic processing sub-unit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0051] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are shown by dotted lines. In different schemes, different basic processing units of the same picture can be divided into different basic processing sub-units.

[0052] In some embodiments, in order to provide parallel processing and fault tolerance for video encoding and decoding, the picture can be divided into multiple processing areas, so that for a certain area of ​​the picture, the encoding or decoding process can be independent of information from any other area of ​​the picture. In other words, each area of ​​the picture can be processed independently. In this way, the codec can process different areas of the picture in parallel, thereby improving coding efficiency. In addition, when the data of one area is damaged during processing or lost in network transmission, the codec can correctly encode or decode other areas of the same picture without relying on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, pictures can be divided into different types of areas. For example, H.265 / HEVC and H.266 / VVC provide two types of areas; "slices" and "tilings". It should also be noted that different pictures of the video sequence 100 can have different partitioning schemes for dividing pictures into areas.

[0053] For example, in Figure 1 , structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and regions of the structure 110 are merely examples, and the present application does not limit their implementation.

[0054] Figure 2A 2 shows a schematic diagram of an example encoding process 200A consistent with an embodiment of the present invention. For example, the encoding process 200A can be performed by an encoder. Figure 2A As shown, the encoder can encode the video sequence 202 into a video code stream 228 according to the encoding process 200A. Figure 1 The video sequence 100 in FIG. 200 may include a set of pictures (referred to as “original pictures”) arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , the encoder may divide each original picture of the video sequence 202 into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform the encoding process 200A at the basic processing unit level for each original picture of the video sequence 202. For example, the encoder may perform the encoding process 200A in an iterative manner, wherein the encoder may encode the basic processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform the process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0055] exist Figure 2A 2, an encoder may input a basic processing unit (referred to as an "original BPU") of an original picture of a video sequence 202 into a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may input the residual BPU 210 into a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may input the prediction data 206 and the quantized transform coefficients 216 into a binary encoding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." During process 200A, after the quantization stage 214, the encoder may input the quantized transform coefficients 216 into an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0056] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0057] Referring to process 200A, an encoder may receive a video sequence generated by a video capture device (e.g., a camera) 202. The term "receiving" as used herein may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or in any way inputting data.

[0058] In the prediction phase 204, in the current iteration, the encoder may receive the original BPU and the prediction reference 224, and perform a prediction operation to generate the prediction data 206 and the prediction BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206 that can be used to reconstruct the original BPU into the prediction BPU 208 from the prediction data 206 and the prediction reference 224.

[0059] Ideally, the predicted BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is usually slightly different from the original BPU. In order to record this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the value of the corresponding pixel of the original BPU from the value (e.g., grayscale value or RGB value) of the pixel corresponding to the predicted BPU 208. Each pixel of the residual BPU 210 can have a residual value as a result of this subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared with the original BPU, the prediction data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly reducing the quality, thereby compressing the original BPU.

[0060] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basic patterns", each of which is associated with a "transform coefficient". The basic patterns can have the same size (e.g., the size of the residual BPU 210). Each basic pattern can represent a frequency-varying (e.g., frequency-varying) component of the residual BPU 210. Any basic pattern cannot be reproduced from any combination (e.g., linear combination) of any other basic patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basic patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0061] Different transform algorithms may use different basic modes. In the transform stage 212, various transform algorithms may be used, such as discrete cosine transform, discrete sine transform, etc. The transform in the transform stage 212 is reversible. That is, the encoder may restore the residual BPU 210 by an inverse operation of the transform (referred to as an "inverse transform"). For example, in order to restore the pixels of the residual BPU 210, the inverse transform may be to multiply the values ​​of the corresponding pixels of the basic mode by the corresponding correlation coefficients, and to add the products to produce a weighted sum. For video coding standards, both the encoder and the decoder may use the same transform algorithm (and therefore the same basic mode). Therefore, the encoder may record only the transform coefficients, and the decoder may reconstruct the residual BPU 210 from the transform coefficients without receiving the basic mode from the encoder. The transform coefficients may have fewer bits than the residual BPU 210, but they may be used to reconstruct the residual BPU 210 without significantly reducing the quality. Therefore, the residual BPU 210 is further compressed.

[0062] The encoder may further compress the transform coefficients in the quantization stage 214. During the transform process, different basic modes may represent different frequencies of change (e.g., frequency of brightness changes). Since the human eye is generally better at recognizing low-frequency changes, the encoder may ignore information about high-frequency changes without causing a significant decrease in decoding quality. For example, in the quantization stage 214, the encoder may generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to its nearest integer. Through this operation, some transform coefficients of high-frequency basic modes may be converted to zero, while transform coefficients of low-frequency basic modes may be converted to smaller integers. The encoder may ignore zero-valued quantized transform coefficients 216, which further compress the transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients in an inverse operation of quantization (called "inverse quantization").

[0063] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. Generally, the quantization stage 214 may cause the greatest information loss in the process 200A. The greater the information loss, the fewer bits may be required to quantize the transform coefficients 216. To achieve different degrees of information loss, the encoder may use different quantization parameter values ​​or any other parameters in the quantization process.

[0064] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type of the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0065] Referring to the reconstruction path of process 200A, at the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0066] It should be noted that other variations of process 200A may also be used to encode video sequence 202. In some embodiments, the encoder may perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit Figure 2A one or more stages in a process.

[0067] Figure 2B A schematic diagram of another example encoding process 200B consistent with an embodiment of the present application is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that conforms to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0068] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-frame prediction") can use pixels from one or more encoded adjacent BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction may include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in a picture. Temporal prediction (e.g., inter-picture prediction or "inter-frame prediction") can use regions from one or more encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction may include an encoded picture. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For the original BPU of the picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate a predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be positioned relative to the original BPU from various directions, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the location (eg, coordinates) of the used neighboring BPU, the size of the used neighboring BPU, extrapolation parameters, the direction of the used neighboring BPU relative to the original BPU, and the like.

[0070] For another example, in the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference picture may be encoded and reconstructed by the BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching area in a range of the reference picture (referred to as a "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a position in the reference picture having the same coordinates as the original BPU in the current picture, and may extend outward by a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) an area similar to the original BPU in the search window, the encoder may determine such an area as a matching area. The matching region may have a different size (e.g., smaller, equal, larger, or different in shape) than the original BPU. Figure 1 ), so it can be considered that over time, the matching area "moves" to the location of the original BPU. The encoder can record the direction and distance of this movement as a "motion vector" when using multiple reference pictures (for example, Figure 1 When the encoder is used to search for matching areas and determine the motion vector associated with each reference picture, the encoder may assign weights to the pixel values ​​of the matching areas of each matching reference picture.

[0071] Motion estimation may be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, etc.

[0072] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of ​​a reference picture according to the motion vector, where the encoder may predict the original BPU of the current picture. Figure 1The encoder may move the matching region of the reference picture according to the corresponding motion vector and the average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values ​​of the matching regions of the respective matching reference pictures, the encoder may add the weighted sum of the pixel values ​​of the moving matching region.

[0073] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same temporal direction as the current picture. For example, Figure 1 The picture 104 in is a unidirectional inter-frame prediction picture, where the reference picture (i.e., picture 102) precedes the picture 104. Bidirectional inter-frame prediction can use one or more reference pictures in two temporal directions relative to the current picture. For example, Figure 1 The picture 106 in is a bidirectional inter-prediction picture, where the reference pictures (ie, pictures 104 and 108 ) are in both temporal directions relative to the picture 104 .

[0074] Still referring to the forward path of process 200B, after spatial prediction stage 2042 and temporal prediction stage 2044, at mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode based on the bit rate of a candidate prediction mode and the distortion of a reference picture reconstructed under the candidate prediction mode to minimize the value of a cost function. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0075] In the reconstruction path of process 200B, if the intra-frame prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can directly input the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If the inter-frame prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs have been encoded and reconstructed), the encoder can input the prediction reference 224 to the loop filter stage 232, at which time the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking effect) introduced by the inter-frame prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., used as an inter-frame prediction reference picture for future pictures of the video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in a temporal prediction stage 2044. In some embodiments, the encoder may encode parameters of the loop filter (e.g., loop filter strength) in a binary encoding stage 226, as well as quantized transform coefficients 216, prediction data 206, and other information.

[0076] Figure 3A A schematic diagram of an example decoding process 300A consistent with an embodiment of the present application is shown. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2 . In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode video code stream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figure 2A-2B 214), typically, the video stream 304 is different from the video sequence 202. Figure 2A-2B In the process 200A and 200B in the decoder, the decoder can perform the process 300A at the basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder can perform the process 300A in an iterative manner, where the decoder can decode the basic processing unit in one iteration of the decoding process 300A. In some embodiments, the decoder can perform the process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0077] In FIG. A, the decoder may input a portion of a video code stream 228 associated with a basic processing unit (referred to as an "encoded BPU") of an encoded picture into a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may input the quantized transform coefficients 216 into an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may input the prediction data 206 into the prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may input the predicted reference 224 into the prediction stage 204 for performing a prediction operation in the next iteration of the process 300A.

[0078] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video code stream 228.

[0079] In the binary decoding stage 302, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may also decode other information in the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the quantization parameter process (e.g., quantization parameter), the encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder may unpack the video bitstream 228 before inputting it into the binary decoding stage 302.

[0080] Figure 3B A schematic diagram of another example decoding process 300B consistent with an embodiment of the present application is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder that conforms to a hybrid video coding standard (e.g., H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0081] In process 300B, for an encoded basic processing unit (referred to as a "current BPU") of an encoded picture being decoded (referred to as a "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder uses intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., an identification value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the position (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, extrapolation parameters, the direction of the neighboring BPU relative to the original BPU, etc. For example, if the encoder uses inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator indicating inter-frame prediction (e.g., an identification value), parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with the reference pictures respectively, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference pictures, one or more motion vectors associated with the matching regions respectively, etc.

[0082] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B After performing such spatial prediction or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, such as Figure 3A Described in .

[0083] In process 300B, the decoder may input the prediction reference 224 into the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may input the prediction reference 224 directly into the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference picture in which all BPUs have been decoded), the encoder may input the prediction reference 224 into the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2B The loop filter is applied to the prediction reference 224 in the manner described in . The loop filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-frame prediction reference picture for a future encoded picture of the video code stream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-frame prediction is used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength).

[0084] Figure 4 4 is a block diagram of an example apparatus 400 for encoding or decoding a video consistent with an embodiment of the present application. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for video encoding or decoding. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logic controller. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0085] The device 400 may also include a memory 404 configured to store data (eg, a set of instructions, computer code, intermediate data, etc.). Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video code stream 228, or video stream 304). Processor 402 can access program instructions and data for processing (e.g., via bus 410) and execute program instructions to operate or manipulate the data for processing. Memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).

[0086] The bus 410 may be a communication device that transmits data between components inside the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a universal serial bus port, a peripheral component interconnect express port), and the like.

[0087] For ease of explanation and without ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this application. The data processing circuits can be fully implemented as hardware, or a combination of software, hardware, or firmware. In addition, the data processing circuit can be a single independent module, or can be fully or partially combined into any other component of the device 400.

[0088] The device 400 may also include a network interface 406 to provide wired or wireless communications with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.) In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0089] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touch pad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or a communication input interface coupled to a video archive), and the like.

[0090] It should be noted that the video codec (e.g., a codec that performs the process 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in the device 400. For example, some or all stages of the process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into the memory 404. For another example, some or all stages of the process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).

[0091] In the quantization and inverse quantization blocks (e.g. Figure 2A or Figure 2B quantization 214 and inverse quantization 218, Figure 3A or Figure 3B In the inverse quantization 218 of the prediction residual, the quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. For example, the initial QP value for picture coding or slice coding can be sent at a high level using the init_QP_minus26 syntax element in the picture parameter set (PPS) and the slice_QP_delta syntax element in the slice header. In addition, the QP value can be adjusted for each CU at a local level using a delta QP value sent at the granularity of a quantization group.

[0092] In Versatile Video Coding draft 6 ("VVC 6"), the residual of a transform block (TB) of video data may be encoded using a transform skip (TS) mode that skips the transform stage. For example, a decoder may decode the video data using the TS mode by decoding the video data to obtain the residual, and then inverse quantize and reconstruct the residual without inverse transform. VVC 6 restricts the applicability of the TS mode by the size of the maximum block, where the TS mode is only applicable to TBs with a width and height of at most 32 pixels. The size of the maximum block to which the TS mode applies may be specified as a picture parameter set (PPS) level syntax, log2_transform_skip_max_size_minus2 (transform skip parameter), which may take a value in the range of 0 to 3. When not present, the value of log2_transform_skip_max_size_minus2 (transform skip parameter) is inferred to be 0. The maximum value MaxTsSize (maximum transform size) of the width or height of the maximum block that limits the TS mode may be determined according to formula (1): MaxTsSize = 1 << ( log2_transform_skip_max_size_minus2 + 2 ) Formula 1

[0093] In other words, when log2_transform_skip_max_size_minus2 (transform skip parameter) is 0, TS mode can be allowed if the TB width and height are at most 4. In the current design of VVC 6, the maximum allowed value of MaxTsSize is 32, because the maximum allowed value of log2_transform_skip_max_size_minus2 (transform skip parameter) is 3. If the width and height of the TB are at most MaxTsSize, a parameter transform_skip_flag (transform skip flag) can be sent to specify whether TS mode is selected. If the width or height of the TB is greater than 32, TS mode is not allowed for this TB.

[0094] In VVC 6, the residual of TS mode is coded using non-overlapping coefficient groups (CGs) of size 4×4. The transform skip coefficient levels of the CGs are coded in three passes over the scan positions.

[0095] The first traversal can be represented by the following pseudocode:

[0096] The second traversal can be represented by the following pseudo code:

[0097] The third traversal can be represented by the following pseudo code: for(n=0;n<=numSbCoeff-1;n++) rice=cctx.templateAbsSumTS(n,coeff); decode abs_remainder_using_RG_Coding

[0098] In the above description, syntax elements in TS mode for residual coding (referred to as "TS residual coding") may use context coding (labeled as "context") or bypass coding (labeled as "bypass").

[0099] In some embodiments, a coding tool called "level mapping" can be used for TS residual coding. The absolute coefficient level parameter absCoeffLevel can be mapped to a modified level that is encoded based on the values ​​of the quantized residual samples that are located to the left and above the current residual samples. Let X0 represent the absolute coefficient level to the left of the current coefficient and let X1 represent the absolute coefficient level above the current coefficient. In order to represent coefficients with absolute coefficient levels ("absCoeff"), the mapping parameter absCoeffMod can be encoded. absCoeffMod can be derived in the following pseudo-code representation:

[0100] There are several challenges in the current design of TS mode. In VVC 6, TS mode is a coding tool that enables mathematical lossless compression of blocks by choosing appropriate quantization parameter values ​​and turning off the loop filter stage. Since VVC 6 does not allow TS mode for TBs with a width or height greater than 32, the current design in VVC 6 cannot achieve mathematical lossless compression of blocks when the TB width or height is greater than 32.

[0101] In addition, the newly adopted level mapping process significantly affects the throughput of context adaptive binary arithmetic coding (CABAC) because for each coefficient level, the decoder needs to calculate the prediction value from the top and the left. Since the derivation process of Rice parameters depends on the actual level, the calculation of the actual level involving inverse mapping needs to be implemented in the CABAC parsing loop. This interleaving of parsing and level decoding is undesirable because it reduces the throughput of the decoder hardware implementation.

[0102] In VVC 6, in addition to the log2_transform_skip_max_size_minus2 (transform skip parameter) as described above, another sequence parameter set (SPS) level flag sps_max_luma_transform_size_64_flag can specify the maximum TB size in luma samples. When sps_max_luma_transform_size_64_flag is equal to 1, the maximum TB size in luma samples is equal to 64. When sps_max_luma_transform_size_64_flag is equal to 0, the maximum TB size in luma samples is equal to 32. When the luma coding tree block size (CTU) of a coding tree unit is less than 64, the value of sps_max_luma_transform_size_64_flag is equal to 0. According to sps_max_luma_transform_size_64_flag, the parameter MaxTbLog2SizeY and the maximum TB size MaxTbSizeY can be derived based on formulas (2) and (3): MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6∶5 Formula 2 MaxTbSizeY = 1 << MaxTbLog2SizeY Formula 3

[0103] Based on formulas (2) to (3), the maximum value log2_transform_skip_max_size_minus2 (transform skip parameter) of the PPS level syntax can depend on the SPS level flag sps_max_luma_transform_size_64_flag. log2_transform_skip_max_size_minus2 specifies the size of the maximum block used by the TS mode, and its value can be in the range of 0 to (3+sps_max_luma_transform_size_64_flag). The encoder can be configured to ensure that the value of log2_transform_skip_max_size_minus2 is within the allowed range. When not present, the value of log2_transform_skip_max_size_minus2 can be inferred to be 0. The maximum allowed MaxTsSize can be determined using formula (1). If the width and height of the TB are less than MaxTsSize, the TS mode can be allowed to encode the TB.

[0104] From the above description, it can be seen that in VVC 6, the log2_transform_skip_max_size_Mins2 signal is issued only when sps_transform_skip_enabled_flag is 1. sps_transform_skip_enabled_flag equals 0, indicating that transform_skip_flag does not exist in the transform unit syntax. Therefore, when sps_transform_skip_enabled_flag is 0, there is no need to send the log2_transform_skip_max_size_minus2 signal. At present, this signaling in VVC 6 has the problem of parsing the dependency between SPS and PPS. The above embodiment also has the problem of parsing the dependency between the PPS syntax log2_transform_skip_max_size_minus2 and the SPS syntax sps_max_luma_transform_size_64_flag. This kind of parsing dependency is usually undesirable.

[0105] The embodiments of the present application provide technical solutions for the above technical problems. In order to achieve lossless compression using TS mode for large TB, the present application provides an embodiment in which the TS mode can be extended to apply to TB size up to the maximum TB size allowed by the encoded video sequence. For TS residual coding, different coefficient scanning methods are also provided.

[0106] In accordance with some embodiments of the present application, in order to remove the analytical dependency between SPS and PPS, log2_transform_skip_max_size_minus2 can be moved from PPS to SPS. For example, Figure 5 Table 1 is shown, which shows an example syntax structure of a sequence parameter set (SPS) according to some embodiments of the present disclosure. Figure 6 Table 2 is shown, which shows an example syntax structure of a picture parameter set (SPS) according to some embodiments of the present application. Tables 1 and 2 show that log2_transform_skip_max_size_minus2 is moved from PPS to SPS, as shown in row 502 of Table 1 and rows 602-604 of Table 2.

[0107] Consistent with some embodiments of the present application, the size of the maximum block of blocks to which TS mode is applied may be set to the maximum TB size (MaxTbSizeY), in which case log2_transform_skip_max_size_minus2 is not signaled. In doing so, TS mode may be allowed if the width and height of the TB are less than or equal to MaxTbSizeY. In some embodiments, MaxTbSizeY may be determined based on formulas (2) to (3).

[0108] As an example, Figure 7 Table 3 is shown, which shows an example syntax structure of a transform unit according to some embodiments of the present disclosure. Table 3 shows that according to the example syntax structure of the transform unit, the width and height of the TB can be less than or equal to the maximum value MaxTbSizeY (i.e., 32), as shown in line 706. By doing so, because the size of the largest block to which the TS mode is applied is the same as MaxTbSizeY, all TBs can allow the TS mode, and no additional check is required to determine whether the width and height of a TB is less than or equal to MaxTbSizeY, as shown in lines 702-704. It should be noted that VVC 6 also uses a multiple transform selection (MTS) scheme for residual encoding of inter-frame and intra-frame coded blocks. MTS uses multiple selected transforms from DCT8 / DST7. However, during MTS encoding, additional checks are required because MTS is allowed when both tbWidth and tbHeight are less than or equal to 32.

[0109] VVC 6 provides another coding tool called Block Differential Pulse Code Modulation (BDPCM). In BDPCM mode, horizontal and vertical differential pulse code modulation (DPCM) is applied in the residual domain and the transform stage is skipped. The maximum allowed block width or height for applying BDPCM mode is the same as TS mode.

[0110] According to some embodiments of the present application, the size of the maximum block to which the BDPCM mode is applied can also be extended to the size of the maximum block to which the TS mode is applied. By doing so, if the width and height of the coding unit (CU) are less than or equal to MaxTbSizeY, the BDPCM mode can be allowed. For example, Figure 8 Table 4 is shown according to some embodiments of the present application, which shows an example syntax structure related to signaling a block differential pulse code modulation (BDPCM) mode. Table 4 shows that the size of the maximum block applying the BDPCM mode can be extended to the size of the maximum block applying the TS mode, as shown in line 802.

[0111] In some cases, the allowed values ​​of log2_transform_skip_max_size_minus2 (transform skip parameter) may depend on the profile of the codec. For example, the main profile may specify that the value of log2_transform_skip_max_size_minus2 may be the same as the maximum TB size. Any bitstream that indicates a value of log2_transform_skip_max_size_minus2 that is different from the maximum TB size may be considered a non-conforming bitstream by the codec. If an extended profile exceeds the range of the main profile, the value of log2_transform_skip_max_size_minus2 may be different from the maximum TB size.

[0112] In accordance with some embodiments of the present application, some methods and syntax structures are provided herein to ensure that the value of log2_transform_skip_max_size_minus2 is always the same as the maximum TB size, for example, by not signaling log2_transform_skip_max_size_minus2 and inferring it to be the same as the maximum TB size, or by configuring it through configuration file constraints. By doing so, the burden on the decoder implementation can be reduced because there are fewer combinations of syntax element values ​​to test.

[0113] In some embodiments, the SPS flag may be signaled to indicate that the maximum block size for applying the TS mode is 32 or 64. For example, the SPS flag may be signaled in the same manner as the maximum TB size signal. For example, sps_max_transform_skip_size_64_flag may be set to 0 to specify that the maximum block size for applying the TS mode is 32. For another example, sps_max_transform_skip_size_64_flag may be set to 1 to specify that the maximum block size for applying the TS mode is 64. In some embodiments, when the sps_max_transform_skip_size_64_flag signal is not sent, its value may be inferred to be 0.

[0114] In some embodiments, the size of the maximum block to which the TS mode is applied may be determined based on Formula 4: MaxTsSize=sps_max_transform_skip_size_64_flag? 64:32 formula 4

[0115] In some embodiments, if sps_max_luma_transform_size_64_flag and sps_transform_skip_enabled_flag are both equal to 1, then sps_max_transform_skip_size_64_flag may be signaled.

[0116] As an example, Fig. 9 Table 5 is shown, which shows an example syntax structure of an SPS for sending a sps_max_transform_skip_size_64_flag signal according to some embodiments of the present application. Fig.10 Table 6 is shown, which shows an example syntax structure of a transform unit for signaling sps_max_transform_skip_size_64_flag according to some embodiments of the present application. Tables 5 and 6 show the implementation of signaling sps_max_transform_skip_size_64_flag, as shown in row 902 of Table 5 and rows 1002-1006 of Table 6.

[0117] Consistent with some embodiments of the present application, because the size of the maximum block to which the TS mode or BDPCM mode is applied can be extended to the maximum TB size, the residual coding in the TS mode or BDPCM mode can also be extended to allow the maximum TB size to be encoded therein. According to some embodiments of the application, the residual coding can be directly extended to allow up to the maximum TB size without modifying any scanning mode.

[0118] In some embodiments, similar to VVC draft 6, the transform block can be divided into coefficient groups (CGs) and diagonal scanning can be performed. For example, Fig.11 is a schematic diagram of an example diagonal scan of a 64×64 transform block (TB) according to some embodiments of the present application. Fig.11 A diagonal scan pattern (indicated by the zigzag arrow lines) of 64×64 TB (eg, MaxTbSizeY=64) is shown. Fig.11 Each cell in can represent a 4×4 CG. It should be noted that although Fig.11 A 64×64 TB is shown to illustrate the diagonal scanning process, but the TB can be any size or any shape and is not limited to the examples described herein. For example, when the TB is rectangular rather than square, it has only one dimension equal to 64.

[0119] Scanning the entire TB in residual coding (e.g. Fig.11One challenge with M×N TB (64×64TB) is that the current VVC residual coding needs to be modified to support the above extensions, because the current residual coding in VVC only supports a maximum block size of 32×32. In the current VVC design, even if the transform is applied to 64×64TB (for example, in non-skipped mode), the decoder may still need to apply residual coding only to the 32×32 coefficient block representing the upper left 32×32 block of the 64×64TB. In this case, all remaining high-frequency coefficients are forced to zero (so the remaining coefficients do not need to be encoded). For example, for M×N TB (M is the block width and N is the block height), when M is equal to 64, only the left 32 columns of transform coefficients can be encoded. Similarly, when N is equal to 64, only the first 32 rows of transform coefficients can be encoded.

[0120] Consistent with some embodiments of the present application, in order to reuse the existing VVC 6 residual coding technology, a large TB can be divided into small residual units (RUs). For example, if the width of the TB is greater than 32, the TB can be split horizontally into 2 partitions. For another example, if the height of the TB is greater than 32, the TB can be split vertically into 2 partitions. In another example, if both dimensions of the TB are greater than 32, the TB can be divided into four RUs horizontally and vertically. After splitting, the 32×32RU can be encoded.

[0121] As an example, Figures 12A-12D An example residual unit (RU) according to some embodiments of the present application is shown. Fig. 12A In Figure 1, a 64×64TB is divided into four 32×32RUs (indicated by the dotted lines). Fig. 12B In Figure 1, a 64×16TB is split horizontally into two 32×16Rus (indicated by the dotted lines). Fig. 12C In Figure 1, a 32×64TB is divided vertically into two 32×32Rus (indicated by the dotted lines). Fig.12D In the example, since both the height and width are not greater than 32, no segmentation is performed and the RU size is the same as the TB size. In some embodiments, the maximum allowed RU size is 32×32.

[0122] As an example, Fig.13 is a schematic diagram of an example of diagonal scanning of a 64×64 TB according to some embodiments of the present application, where the TB is divided into four 32×32 RUs. Fig.13 In FIG. 1 , the 64×64 TB is divided into four RUs (indicated by the thick solid lines within the TB), and the coefficients of each RU are scanned individually (e.g., independently) within the RU in the same order as the scanning pattern of the 32×32 TB. Fig.13As shown, the context model and Rice parameter derivation of one RU can be independent of another RU. In some embodiments, the maximum number of context coding containers can also be independently allocated to each RU. This scheme is different from VVC 6, in which the maximum number of context coding containers is defined at the TB level.

[0123] As an example, Figures 14A-14D Table 7 is shown according to some embodiments of the present application, which shows an example syntax structure for residual coding when a TB is divided into RUs.

[0124] In VVC 6, for each coefficient group (CG) of a TS mode block, a coded_sub_block_flag is signaled. coded_sub_block_flag=0 means that all coefficients of the CG are zero. coded_sub_block_flag=1 means that at least one coefficient within the CG is not zero. However, the coded_sub_block_flag of the last CG is not signaled, and is inferred to be 1 if the coded_sub_block_flag of all previously encoded CGs (i.e., before the last CG) is zero. This means that the parsing of the last CG of a TB depends on all previously decoded CGs. In order to eliminate dependencies between RUs, coded_sub_block_flag can be sent for all CGs of the RU (including the last CG).

[0125] Consistent with some embodiments of the present application, an additional syntax coded_RU_flag may be introduced. In some embodiments, a coded_RU_flag signal may be issued when the number of RUs within a TB is greater than 1. In some embodiments, if coded_RU_flag does not exist, it may be inferred to be 1. coded_RU_flag=0 may specify that all coefficients of the RU are zero. coded_RU_flag=1 may specify that at least one of the coefficients of the RU is non-zero. In some embodiments, if all coded_RU_flags except the last RU are zero, the coded_RU_flag signal of the last RU does not need to be issued and may be inferred to be 1. As an example, the following pseudo code shows example signaling of coded_RU_flag:

[0126] As an example, Figures 15A-15DTable 8 according to some embodiments of the present application is shown, which shows another example syntax structure for residual coding when the coded_RU_flag signal is sent. In some embodiments, if the coded_RU_flag signal is sent, the last CG identification can be maintained in the same way as in VVC 6. That is, if coded_sub_block_flag is 0 in the same RU, then coded_sub_block_flag can be inferred to be 1.

[0127] The Joint Video Experts Group (JVET) AHG Lossless and Near Lossless Coding Tools AHG18 released lossless software based on VTM-6.0. The lossless software introduced a CU-level flag called cu_transquant_bypass_flag. cu_transquant_bypass_flag=1 means that the transform and quantization of the CU are skipped, and the CU is encoded in lossless mode. In the current version of the lossless software, sps_max_luma_transform_size_64_flag is set to 0, which means that the maximum TB size in the luma sample is limited to 32×32. For chroma samples, the maximum TB size is adjusted according to the YUV color format (for example, for YUV 420, a maximum of 16×16). In some embodiments, when cu_transquant_bypass_flag=1, the luma transform block size can be increased to 64×64, and when cu_transquant_bypass_flag=1, the above-mentioned residual coding technique can be used.

[0128] In some embodiments, the maximum TB size of the chroma component can be determined using formulas (2) and (3). Based on formulas (2) and (3), the maximum TB width maxTbWidth and height maxTbHeight can be determined according to formulas (5) and (6): maxTbWidth=(cIdx= formula = 0)? MaxTbSizeY : MaxTbSizeY / SubWidthC (5) maxTbHeight=(cIdx= formula = 0)? MaxTbSizeY : MaxTbSizeY / SubHeightC (6)

[0129] In formulas (5) and (6), cIdx=0 represents the luminance component. cIdx=1 and cIdx=2 represent two chrominance components. For example, the values ​​of SubWidthC and SubHeightC can be derived from the chrominance format. Consistent with some embodiments of the present application, Fig.16 Table 9 shows example parameter values ​​derived from the chroma format according to some embodiments of the present application.

[0130] In VVC 6, the inverse hierarchy mapping is embedded into the CABAC module. Fig.17 Table 10 is shown, which shows an example syntax structure in VVC 6 for performing inverse level mapped residual coding according to some embodiments of the present application.

[0131] Consistent with some embodiments of the present application, in order to improve the CABAC throughput of the transform skipping residual parsing, the Rice parameters can be derived based on the mapped level values ​​rather than the actual level values. In some embodiments, both the context model and the Rice parameters can depend on the mapped values, and the inverse mapping operation may not be performed during the residual parsing process. By doing so, the inverse mapping can be decoupled from the residual parsing process. After the residual parsing of the entire TB is completed, the inverse mapping can be performed. In some embodiments, inverse mapping and residual parsing can be performed simultaneously in one traversal, which allows a decision based on actual conditions whether to interleave parsing and mapping or to separate them into two processes.

[0132] As an example, Fig.18 is a flowchart of an exemplary decoding method 1400 according to some embodiments of the present application. The method 1800 may be performed in a case where parsing and inverse mapping are separated. Fig.18 It can be seen that the inverse mapping is performed after the residual analysis of the entire TB is completed and before the inverse quantization, which is decoupled from the residual analysis.

[0133] Consistent with some embodiments of the present application, Fig.19 Table 10 is shown, which shows an example syntax structure for residual coding without performing inverse hierarchy mapping according to some embodiments of the present application. In some embodiments, the inverse hierarchy mapping can be moved to the decoding process, which will be described below.

[0134] Consistent with some embodiments of the present application, the following pseudo code shows an inverse hierarchical mapping process that can be performed after residual parsing and before inverse quantization (eg Fig.18 In the following pseudo code, TransCoeffLevel[xC][yC] represents the coefficient value at the position (xC, yC) after residual analysis, and TransCoeffLevelInvMapped[xC][yC] represents the coefficient value at the position (xC, yC) after inverse mapping:

[0135] According to some embodiments of the present application, Rice parameters can be derived based on mapping values, which is different from the Rice parameters derived based on actual level values ​​in VVC 6. Assuming that the array TransCoeffLevel[xC][yC] is the mapping level value of the TB of a given color component at position (xC, yC), the variable locSumAbs can be derived according to the following pseudo code: locSumAbs=0 AbsLevel[xC][yC]=abs(TransCoeffLevel[xC][yC]) if(xC>0) locSumAbs+=AbsLevel[xC-1][yC] if(yC>0) locSumAbs+=AbsLevel[xC][yC-1] locSumAbs=Clip3(0,31,locSumAbs)

[0136] Consistent with some embodiments of the present application, Fig. 20 Table 12 according to some embodiments of the present application is shown, which shows an example lookup table for selecting Rice parameters. In some embodiments of the application, the value of locSumAbs can be adjusted based on a predefined offset value. In some embodiments, the offset value is calculated based on offline training. The following example pseudo code shows that the offset value is 2. locSumAbs=0 offset = 2; AbsLevel[xC][yC]=abs(TransCoeffLevel[xC][yC]) if(xC>0) locSumAbs+=AbsLevel[xC-1][yC] if(yC>0) locSumAbs+=AbsLevel[xC][yC-1] locSumAbs-=offsetlocSumAbs=Clip3(0,31,locSumAbs)

[0137] Consistent with some embodiments of the present application, Figure 21-22Flowcharts showing example processes 2100-2200 for video processing according to some embodiments of the present application are shown. In some embodiments, processes 2100-2200 may be performed by a codec (e.g., Figure 2A-2B The encoder in Figure 3A-3B For example, the codec may be implemented as one or more software or hardware components of an apparatus for video processing (eg, apparatus 400).

[0138] For example, Fig.21 FIG. 2 is a flowchart of an example process 2100 for video processing according to some embodiments of the present application. In step 2102, a codec (e.g., Figure 2A-2B The encoder in the example may determine the transformation process of skipping the prediction residual based on the maximum value of the size of the luminance sample of the prediction block or one of the maximum values ​​of the size of the prediction block. The transformation process may be Figure 2A-2B The prediction residual may be the residual BPU 210 in 2A-2B. The prediction block may be included in Figure 2A-2B The block in the prediction data 206 in FIG. Figure 11-13 ). The size of the prediction block may include height or width.

[0139] In some embodiments, the codec may decide to skip the transform process based on the determination result that the size of the prediction block is not greater than the threshold, thereby determining to skip the transform process for the prediction residual. In some embodiments, the threshold may be MaxTbSizeY, as shown and described in conjunction with formulas (2) to (3), and the maximum value of the threshold may be equal to one of the maximum value of the luma sample size (e.g., 32, 64, or any number) or the maximum value of the prediction block size (e.g., 32, 64, or any number). In some embodiments, the maximum value of the luma sample size or the maximum value of the prediction block size may be a dynamic value (e.g., not a constant).

[0140] In some embodiments, the threshold is equal to the maximum value of the size of the luma sample indicating the luma information of the prediction block. In some embodiments, the maximum value of the threshold is 64. In some embodiments, the maximum value of the threshold is 32. In some embodiments, the minimum value of the threshold is 4. In some embodiments, the threshold may be equal to the maximum value of the size of the prediction block that allows the transform process to be performed (e.g., MaxTsSize as shown and described in formula (1)).

[0141] Still reference Fig.21In step 2104, the codec may generate residual coefficients. In some embodiments, the maximum value of the threshold is determined based on at least a first parameter in a first parameter set. For example, the first parameter set may be a sequence parameter set (SPS). In some embodiments, the value of the first parameter is 0 or 1. For example, the first parameter may be sps_max_luma_transform_size_64_flag, such as Fig. 9 In some embodiments, the threshold value may be determined based on the value of the first parameter. For example, if the first parameter may be sps_max_luma_transform_size_64_flag, and if the threshold value is MaxTbSizeY, then when sps_max_luma_transform_size_64_flag is equal to 1, MaxTbSizeY may be equal to 64. When sps_max_luma_transform_size_64_flag is equal to 0, MaxTbSizeY is equal to 32.

[0142] In some embodiments, the maximum value of the threshold value may be determined based on at least a first parameter in the first parameter set. In some embodiments, the threshold value may be determined based on a value of a second parameter in the second parameter set. In some embodiments, the second parameter set is a sequence parameter set (SPS). In some embodiments, the second parameter set is a picture parameter set (PPS). The second parameter may be log2_transform_skip_max_size_minus2 (e.g., as combined with Figure 5 1 and described in Table 1 in ). The value of the second parameter can be determined based on the value of the first parameter. In some embodiments, the value of the second parameter (e.g., log2_transform_skip_max_size_minus2) has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter (e.g., sps_max_luma_transform_size_64_flag). For example, log2_transform_skip_max_size_minus2 can be in the range of 0 to (3+sps_max_luma_transform_size_64_flag). In some embodiments, the second parameter can have a first value in a first profile (e.g., a main profile) of the encoder and a second value in a second profile (e.g., an extended profile) of the encoder, and the first value and the second value are different.

[0143] Still reference Fig.21At step 2104, the codec may generate residual coefficients for the prediction residual by performing at least one lossless compression process or a quantization process on the prediction residual. As described herein, the residual coefficients may be coefficients associated with a residual encoding process. The quantization process may be Figure 2A-2B 214 in the quantization stage. The lossless compression process may include generating residual coefficients using coefficient groups (CGs). For example, the coefficient groups may be non-overlapping. In some embodiments, the coefficient groups have a size of 4×4.

[0144] In some embodiments, the codec may generate residual coefficients using a multiple transform selection (MTS) scheme. For example, the codec may determine whether the size of the prediction block is greater than 32. If the size of the prediction block is not greater than 32, the codec may generate residual coefficients using the MTS scheme.

[0145] In some embodiments, the codec may further determine the transform skip coefficient level of the coefficient group using one of a context coding technique or a bypass coding technique. The codec may also determine a Rice parameter based on the transform skip coefficient level. The codec may also generate a bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter.

[0146] In some embodiments, the codec may also map the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block to the left of the prediction block and a second value of a second residual coefficient of a second prediction block at the top of the prediction block.

[0147] In some embodiments, the codec may determine a transform skip coefficient level of a coefficient group using one of a context coding technique or a bypass coding technique, map the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block on the left side of the prediction block and a second value of a second residual coefficient of a second prediction block at the top of the prediction block, generate a context model of the context coding technique based on the modified transform skip coefficient level, determine Rice parameters based on the modified transform skip coefficient level, generate remaining coefficients using the coefficient group, and generate a bitstream, a transform skip coefficient level, or Rice parameters by entropy encoding at least one of the coefficient groups.

[0148] Still reference Fig.21 In step 2106, the codec may generate a bitstream by entropy encoding at least the residual coefficients. The bitstream may be Figure 2A-2B The video code stream 228 in.

[0149] Fig. 22 FIG. 2 is a flowchart of another example process 2200 for video processing according to some embodiments of the present application. For example, the process 2200 may be Figure 3A-3B The decoder in is used to execute.

[0150] like Fig. 22 As shown, at step 2202, the decoder receives a bitstream including coding information of a video sequence. The bitstream includes a sequence parameter set (SPS) of the video sequence.

[0151] At step 2204, the decoder determines the maximum transform size of the prediction block based on the parameters in the sequence parameter set (SPS) of the video sequence. The prediction block may be included in Figure 2A-2B , such as a transform block (e.g., Figure 11-13 In some embodiments, the maximum transform size may correspond to the maximum value of the size of the luminance samples of the prediction block, or the maximum value of the size of the prediction block. The size of the prediction block may include height or width. Figure 5-10 A detailed method for determining the maximum transform size based on the parameters in the SPS is described.

[0152] At step 2206, the decoder determines to skip the transform process for the prediction residual of the prediction block based on the maximum transform size. The transform process may be Figure 2A-2B The transformation stage 212 in FIG.

[0153] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and these instructions can be executed by a device for performing the above method (e.g., an encoder and a decoder disclosed in the present application). Common forms of non-transitory media include, for example, floppy disks, floppy disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge memory, and the same network version. The device may include one or more processors (CPU), input / output interfaces, network interfaces, and / or memories.

[0154] The above embodiments can be further described using the following terms: 1. A video processing method, comprising: Determining to skip a transform process on a prediction residual based on a maximum transform size of the prediction block; and The maximum transform size is signaled in the sequence parameter set (SPS). 2. The method of clause 1, wherein determining to skip the transform process on the prediction residual comprises: Based on the size of the prediction block being not greater than a threshold, determining to skip the transform process, the threshold being equal to the maximum of one of the following: the maximum size of luma samples of the prediction block, or The maximum size of the prediction block. 3. A method according to clause 2, wherein one of the maximum value of the size of the luma samples or the maximum value of the size of the prediction block is a dynamic value. 4. A method according to any of the preceding clauses, further comprising: The skipping of the transform process is further determined based on a parameter indicating a transform skip mode. 5. A method according to clause 2, wherein the size of the prediction block comprises a height or a width. 6. The method according to clause 2, wherein the maximum value of the threshold is determined based on at least one first parameter in the first parameter set. 7. A method as described in clause 6, wherein the first parameter set is a sequence parameter set (SPS). 8. A method according to any of clauses 6-7, wherein the value of the first parameter is 0 or 1. 9. A method according to any of clauses 2-8, wherein the maximum value of the threshold is 64. 10. A method according to any of clauses 2-8, wherein the maximum value of the threshold is 32. 11. A method according to any of clauses 2-10, wherein a maximum value of the threshold is determined based on at least a first parameter in the first parameter set and a third parameter in the first parameter set. 12. A method according to any of clauses 2-11, wherein the minimum value of the threshold is 4. 13. A method according to any of clauses 2-12, wherein the threshold is equal to a maximum value of the size of the luma samples representing luma information of the prediction block. 14. A method according to any of clauses 6-13, wherein the maximum value of the threshold is determined based on the value of a second parameter in a second parameter set, and the value of the second parameter is determined based on the value of the first parameter. 15. The method according to clause 14, wherein the value of the second parameter is at least 0 and at most equal to the sum of 3 and the value of the first parameter. 16. A method according to clause 14, wherein the second parameter has a first value in a first profile of an encoder and a second value in a second profile of the encoder, the first value and the second value being different. 17. A method according to any of clauses 14-16, wherein the second parameter set is an SPS. 18. A method according to any of clauses 14-16, wherein the second parameter set is a picture parameter set (PPS). 19. A method according to any of clauses 2-12, wherein the threshold is equal to a maximum value of the size of the prediction block for which the transform process is allowed to be performed. 20. The method of clause 19, wherein the threshold is determined based on a value of the first parameter. 21. A method according to any of the preceding clauses, further comprising: Residual coefficients are generated for the prediction block using a Multiple Transform Selection (MTS) scheme. 22. The method according to clause 21, further comprising: Determine whether the size of the prediction block is not greater than 32; and Based on determining that the size of the prediction block is not greater than 32, residual coefficients are generated using the MTS scheme. 23. The method according to any one of clauses 2 to 22, further comprising: Determining whether the size of the prediction block is not greater than the threshold; and Based on the judgment result that the size of the prediction block is not greater than the threshold, before generating residual coefficients for the prediction block, block differential pulse code modulation (BDPCM) is performed on the prediction residual. 24. A method according to any of the preceding clauses, further comprising: Residual coefficients for the prediction residual are generated by performing a lossless compression process or on the prediction residual, wherein the lossless compression process includes generating the residual coefficients using coefficient groups, the coefficient groups being non-overlapping. 25. The method of clause 24, wherein the coefficient groups are of size 4x4. 26. The method according to any of clauses 24-25, further comprising: Determining a transform skip coefficient level for the coefficient group using one of a context coding technique or a bypass coding technique; determining Rice parameters based on the transform skip coefficient level; and A code stream is generated by entropy encoding at least one of a coefficient group, a transform skip coefficient level, or a Rice parameter. 27. The method according to any one of clauses 24-26, further comprising: The transform skip coefficient level is mapped to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block at the left side of the prediction block and a second value of a second residual coefficient of a second prediction block at the top of the prediction block. 28. The method according to clauses 24-26, further comprising: Determining a transform skip coefficient level for a coefficient group using one of a context coding technique or a bypass coding technique; mapping a transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block on the left side of the prediction block and a second value of a second residual coefficient of a second prediction block on the top of the prediction block; generating a context model for a context coding technique based on the modified transform skip coefficient hierarchy; determining a Rice parameter based on the modified transform jump coefficient level; generating the residual coefficients using the coefficient group; and The code stream is generated by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter. 29. The method according to clause 28, further comprising: After performing a quantization process and during generating the residual coefficients, the transform skip coefficient level is mapped to the modified transform skip coefficient level. 30. The method according to clause 28, further comprising: After performing a quantization process and before generating residual coefficients, the transform skip coefficient level is mapped to the modified transform skip coefficient level. 31. The method of any of clauses 28-30, wherein determining the Rice parameter comprises: The Rice parameters are determined based on a modified transform skip coefficient level of a color component of the prediction block. 32. A method according to clause 31, wherein the modified transform skip coefficient level of the color component is offset by a predetermined offset value. 33. A method according to clause 32, wherein the predetermined offset value is determined using a machine learning model during an offline training process. 34. A method according to any of clauses 23-33, wherein generating the residual coefficients comprises: At least one of a lossless compression process or BDPCM is performed on the prediction residual using diagonal scanning, wherein a maximum size of a prediction block for diagonal scanning is 64. 35. A method according to any of clauses 23-33, wherein generating the residual coefficients comprises: Based on the size of the prediction block being greater than 32, dividing the prediction block into a plurality of sub-blocks in size; and For each specific sub-block among the multiple sub-blocks, at least one of a lossless compression process or BDPCM is performed on the prediction residual associated with the specific sub-block using diagonal scanning, wherein corresponding parameters and output results of the lossless compression process or BDPCM associated with the multiple sub-blocks are independent. 36. The method according to clause 35, further comprising: On the basis of determining that the two sizes of the prediction block are greater than 32, the prediction block is divided into a plurality of sub-blocks of the two sizes. 37. A method according to any one of clauses 35-36, wherein the corresponding parameters and output results of the lossless compression process or BDPCM associated with the multiple sub-blocks include at least one of the following: a context model associated with the context coding technology, Rice parameters, or a maximum number of context coding boxes associated with the context coding technology. 38. A method according to any of clauses 34-37, wherein the unit of diagonal scanning is the coefficient group. 39. The method according to clause 38, further comprising: A first indication parameter indicating coefficient values ​​in the coefficient group is set for each coefficient group of a specific sub-block. 40. The method according to clause 38, further comprising: A second indication parameter is set for each specific sub-block among the plurality of sub-blocks, where the second indication parameter is used to indicate values ​​of all coefficient groups in the specific sub-block. 41. The method according to clause 40, further comprising: Setting, for each coefficient group of the specific sub-block, a first indication parameter for indicating coefficient values ​​in the coefficient group; and Based on determining that the first indication parameters of all coefficient groups before the last coefficient group of the specific sub-block are zero, the first indication parameter of the last coefficient group is set to 1. 42. A method according to any of clauses 24-41, wherein generating the residual coefficients comprises: The residual coefficients are generated by performing lossless compression processing on the prediction residual based on the parameter indicating the lossless encoding mode, wherein a maximum value of the size of the luma sample is 64. 43. A method according to any of the preceding clauses, further comprising: Receive video images; Split the video picture into multiple blocks; generating a prediction block by performing one of intra prediction or inter prediction on the block; and A prediction residual is generated by subtracting the prediction block from the block. 44. An apparatus comprising: a memory configured to store instructions; and A processor configured to execute the following instructions: Determining to skip a transform process on a prediction residual based on a maximum transform size of the prediction block; and The maximum transform size is signaled in the sequence parameter set (SPS). 45. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: Determining to skip a transform process on a prediction residual based on a maximum transform size of the prediction block; and The maximum transform size is signaled in the sequence parameter set (SPS). 46. ​​A video processing method, comprising: Receive a code stream of a video sequence; Determining a maximum transform size for a prediction block based on a sequence parameter set (SPS) of a video sequence; and Based on the maximum transform size, it is determined to skip a transform process of a prediction residual of the prediction block. 47. A method according to clause 46, wherein determining to skip a transform process of the prediction residual comprises: In response to determining that the size of the prediction block is not greater than a threshold, determining to skip the transform process, the threshold having a maximum value equal to one of: the maximum size of luma samples of the prediction block, or The maximum size of the prediction block. 48. A method according to clause 47, wherein the size of the prediction block comprises a height or a width. 49. A method according to clause 47, wherein a maximum value of the threshold is determined based on at least a first parameter in the SPS. 50. The method of clause 49, wherein the value of the first parameter is 0 or 1. 51. A method according to any of clauses 47-50, wherein the maximum value of the threshold is 64. 52. A method according to any of clauses 47-50, wherein the maximum value of the threshold is 32. 53. A method according to any of clauses 47-52, wherein a maximum value of the threshold is determined based on at least a first parameter in the SPS and a third parameter in the SPS. 54. A method according to any of clauses 47-53, wherein the minimum value of the threshold is 4. 55. A method according to any of clauses 47-54, wherein the threshold is equal to a maximum value of the size of the luma samples used to indicate luma information of the prediction block. 56. A method according to any of clauses 49-54, wherein the maximum value of the threshold is determined based on the value of a second parameter in a second parameter set, and the value of the second parameter is determined based on the value of the first parameter. 57. A method according to clause 56, wherein the value of the second parameter is at least 0 and at most equal to the sum of 3 and the value of the first parameter. 58. A method according to clause 56, wherein the second parameter has a first value in a first profile of an encoder and a second value in a second profile of the encoder, the first value and the second value being different. 59. A method according to any of clauses 56-58, wherein the second parameter set is an SPS. 60. A method according to any of clauses 56-58, wherein the second parameter set is a picture parameter set (PPS). 61. An apparatus comprising: a memory for storing instructions; and The processor is configured to execute the following instructions: Receive a code stream of a video sequence; Determining a maximum transform size for a prediction block based on a sequence parameter set (SPS) of a video sequence; and Based on the maximum transform size, it is determined to skip a transform process of a prediction residual of the prediction block. 62. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: Receive a code stream of a video sequence; determining a maximum transform size for a prediction block based on a sequence parameter set (SPS) of a video sequence; and Based on the maximum transform size, it is determined to skip a transform process of a prediction residual of the prediction block.

[0155] It should be noted that the relational terms such as "first" and "second" in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "includes" and other similar forms have the same meaning and are open-ended, because the one or more items following any of these words are not intended to be an exhaustive list of such items or limited to the listed items.

[0156] As used herein, unless expressly stated otherwise, the term "or" encompasses all possible combinations unless not feasible. For example, if it is stated that a component may include A or B, then unless expressly stated otherwise or not feasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless expressly stated otherwise or not feasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0157] It can be understood that the above embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. When executed by a processor, the software can execute the disclosed method. The computing unit and other functional units described in the present invention can be implemented by hardware, or by software, or by a combination of hardware and software. It can also be understood by those of ordinary skill in the art that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.

[0158] In the foregoing description, embodiments have been described with reference to many specific details, which may vary with different implementations. Certain adjustments and modifications may be made to the described embodiments. Other embodiments considering the specifications and practices of the invention disclosed herein are apparent to those skilled in the art. The foregoing description and embodiments are considered to be examples only, and the true scope and spirit of the invention are indicated by the claims. The order of steps shown in the figures is also intended to be used for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, it will be appreciated by those skilled in the art that these steps may be performed in different orders while implementing the same method.

[0159] In the drawings and the specification, exemplary embodiments have been disclosed. However, many changes and modifications may be made to these embodiments. Therefore, although specific terms are used, they are used only in a general and descriptive sense and not for limiting purposes.

Claims

1. A computer-implemented method for video encoding, comprising: Send the log2_transform_skip_max_size_minus2 parameter in the sequence parameter set, where The log2_transform_skip_max_size_minus2 parameter specifies the maximum block size for transform skipping; Determine the MaxTsSize parameter according to the log2_transform_skip_max_size_minus2 parameter; determining a width and height of a transform block for a chroma component Cr; and Based on determining that one of the width and the height of the transform block for the chroma component Cr is larger than the size of the MaxTsSize parameter, skipping the process of transmitting the transform_skip_flag parameter, wherein the transform_skip_flag parameter specifies whether the transform skip mode is selected.

2. The computer-implemented method for video encoding according to claim 1, characterized in that: The log2_transform_skip_max_size_minus2 parameter takes a value in the range of 0 to 3+sps_max_luma_transform_size_64_flag.

3. The computer-implemented method for video encoding according to claim 1, characterized in that: The value of the MaxTsSize parameter is 32 or 64.

4. A computer-implemented method for video decoding, comprising: Receives the log2_transform_skip_max_size_minus2 parameter in the sequence parameter set, where The log2_transform_skip_max_size_minus2 parameter specifies the maximum block size for transform skipping; Determine the MaxTsSize parameter according to the log2_transform_skip_max_size_minus2 parameter; determining a width and height of a transform block for a chroma component Cr; and Based on determining that one of the width and the height of the transform block for the chroma component Cr is larger than the size of the MaxTsSize parameter, skipping the process of transmitting the transform_skip_flag parameter, wherein the transform_skip_flag parameter specifies whether the transform skip mode is selected.

5. The computer-implemented method for video decoding according to claim 4, characterized in that: The log2_transform_skip_max_size_minus2 parameter takes a value in the range of 0 to 3+sps_max_luma_transform_size_64_flag.

6. The computer-implemented method for video decoding according to claim 4, characterized in that: The value of the MaxTsSize parameter is 32 or 64.

7. A non-transitory computer-readable storage medium having a video bitstream stored thereon, the bitstream being generated by a method executed by a video processing device, the method comprising: sending a log2_transform_skip_max_size_minus2 parameter in a sequence parameter set, wherein the log2_transform_skip_max_size_minus2 parameter specifies a maximum block size for transform skipping; Determine the MaxTsSize parameter according to the log2_transform_skip_max_size_minus2 parameter; determining a width and height of a transform block for a chroma component Cr; and Based on determining that one of the width and the height of the transform block for the chroma component Cr is larger than the size of the MaxTsSize parameter, skipping the process of transmitting the transform_skip_flag parameter, wherein the transform_skip_flag parameter specifies whether the transform skip mode is selected.

8. The non-transitory computer-readable storage medium according to claim 7, wherein: The log2_transform_skip_max_size_minus2 parameter takes a value in the range of 0 to 3+sps_max_luma_transform_size_64_flag.

9. The non-transitory computer-readable storage medium according to claim 7, wherein: The value of the MaxTsSize parameter is 32 or 64.

Citation Information

Patent Citations

  • Method and device for encoding / decoding images

    CN104488270A

  • Method and apparatus for encoding / decoding image

    CN105684442A

  • Method and device for encoding / decoding images, and computer readable medium

    CN108712651A

  • Repositioning of prediction residual blocks in video coding

    US20140226721A1

  • Method and apparatus for intra transform skip mode

    US20150110180A1