Sign data hiding for video recordings
By selectively disabling sign data hiding and adjusting encoding modes, the method optimizes video coding efficiency in residual coding, addressing inefficiencies in existing standards and enhancing compression performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ALIBABA INNOVATION PRIVATE LIMITED
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video coding standards face inefficiencies in residual coding, particularly in modes like conversion skip, block-difference pulse code modulation, and transform-skip residual coding, leading to unnecessary data hiding that affects compression efficiency.
Implementing methods to selectively turn off sign data hiding for residual encoding based on specific encoding modes, such as conversion skip, block-difference pulse code modulation, and transform-skip residual coding, as well as adjusting loop filters and quantization states to optimize compression.
Enhances video coding efficiency by reducing unnecessary data hiding and improving compression performance, aligning with advanced standards like VVC/H.266, while maintaining quality.
Smart Images

Figure 2026074007000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications
[0001] This disclosure claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 994,239, filed on March 24, 2020. The above provisional application is hereby incorporated by reference in its entirety.
[0002] Technical Field
[0002] This disclosure generally relates to video data processing, and more particularly, to residual coding of video data.
Background Art
[0003] Background
[0003] Video is a set of still pictures (or "frames") that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly techniques based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards such as the HEVC (High Efficiency Video Coding) (e.g., HEVC / H.265) standard, the VVC (Versatile Video Coding) (e.g., VVC / H.266) standard, and the AVS standard, which specify a particular video coding format, have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is becoming higher.
Summary of the Invention
Means for Solving the Problems
[0004] Summary of the Disclosure
[0004] Embodiments of the present disclosure provide a method for encoding video data, comprising: receiving a video frame for residual encoding; determining whether the video frame is encoded according to a conversion skip mode at the conversion block level; and, in response to the determination that the video frame is encoded according to a conversion skip mode, turning off sign data hiding for residual encoding.
[0005]
[0005] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether the video frame is encoded according to a block-difference pulse code modulation mode; and, in response to the determination that the video frame is encoded according to a block-difference pulse code modulation mode, turning off sign data hiding for residual coding.
[0006]
[0006] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether the video frame is encoded at the slice level according to a transform-skip residual coding mode; and, in response to the determination that the video frame is not encoded at the slice level according to a transform-skip residual coding mode, turning off sign data hiding for residual coding.
[0007]
[0007] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at the picture level of the video frame and whether conversion-skip residual coding is disabled at the slice level of the video frame; and turning on sign data hiding at the slice level of the video frame in response to the determination that sign data hiding is enabled at the picture level of the video frame and conversion-skip residual coding is enabled at the slice level of the video frame.
[0008]
[0008] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at the picture level of the video frame; turning on sign data hiding at the slice level of the video frame in response to the determination that sign data hiding is enabled at the picture level of the video frame; determining whether sign data hiding is turned off at the slice level of the video frame; and turning off transform skip residual coding at the slice level of the video frame in response to the determination that sign data hiding is turned off at the slice level of the video frame.
[0009]
[0009] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether the video frame is encoded in lossless mode at the slice level; and turning off one or more loop filters at the slice level in response to the determination that the video frame is encoded in lossless mode at the slice level.
[0010]
[0010] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether sign data hiding is turned off at the picture level of the video frame; and, in response to the determination that sign data hiding is turned off at the picture level of the video frame, turning off transform skip residual coding at the slice level of the video frame.
[0011]
[0011] Embodiments of the present disclosure further provide a method for encoding video data, comprising: receiving a video frame for residual coding; determining whether state-dependent quantization is enabled for the video frame; and, in response to the determination that state-dependent quantization is enabled for the video frame, turning off transform-skip residual coding at the slice level of the video frame.
[0012]
[0012] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive video frames for residual coding, determine whether the video frames are coded according to a conversion skip mode at the conversion block level, and, in response to the determination that the video frames are coded according to a conversion skip mode, turn off sign data hiding for residual coding.
[0013]
[0013] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive video frames for residual coding, determine whether the video frames are coded according to a block-difference pulse code modulation mode, and, in response to the determination that the video frames are coded according to a block-difference pulse code modulation mode, turn off sign data hiding for residual coding.
[0014]
[0014] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive video frames for residual coding, determine whether the video frames are coded at the slice level according to a transform-skip residual coding mode, and, in response to the determination that the video frames are not coded at the slice level according to a transform-skip residual coding mode, turn off sign data hiding for residual coding.
[0015]
[0015] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual coding, determine whether sign data hiding is enabled at the picture level of the video frame and whether conversion-skip residual coding is disabled at the slice level of the video frame, and, in response to the determination that sign data hiding is enabled at the picture level of the video frame and conversion-skip residual coding is enabled at the slice level of the video frame, to turn on sign data hiding at the slice level of the video frame.
[0016]
[0016] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to perform: receive a video frame for residual coding; determine whether sign data hiding is enabled at the picture level of the video frame; turn on sign data hiding at the slice level of the video frame in response to the determination that sign data hiding is enabled at the picture level of the video frame; determine whether sign data hiding is turned off at the slice level of the video frame; and turn off transform skip residual coding at the slice level of the video frame in response to the determination that sign data hiding is turned off at the slice level of the video frame.
[0017]
[0017] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive video frames for residual coding, determine whether the video frames are coded in lossless mode at the slice level, and, in response to the determination that the video frames are coded in lossless mode at the slice level, turn off one or more loop filters at the slice level.
[0018]
[0018] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive video frames for residual coding, determine whether sign data hiding is turned off at the picture level of the video frame, and, in response to the determination that sign data hiding is turned off at the picture level of the video frame, turn off transform skip residual coding at the slice level of the video frame.
[0019]
[0019] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising a memory for storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual coding, determine whether state-dependent quantization is enabled for the video frame, and, in response to the determination that state-dependent quantization is enabled for the video frame, to turn off transform-skip residual coding at the slice level of the video frame.
[0020]
[0020] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a conversion skip mode at the conversion block level; and, in response to the determination that the video frame is coded according to a conversion skip mode, turning off sign data hiding for residual coding.
[0021]
[0021] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a block-difference pulse code modulation mode; and, in response to the determination that the video frame is coded according to a block-difference pulse code modulation mode, turning off sign data hiding for residual coding.
[0022]
[0022] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a transform-skip residual coding mode at the slice level; and, in response to the determination that the video frame is not coded according to a transform-skip residual coding mode at the slice level, turning off sign data hiding for residual coding.
[0023]
[0023] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual coding; determining whether sign data hiding is enabled at the picture level of the video frame and whether conversion-skip residual coding is disabled at the slice level of the video frame; and turning on sign data hiding at the slice level of the video frame in response to the determination that sign data hiding is enabled at the picture level of the video frame and conversion-skip residual coding is enabled at the slice level of the video frame.
[0024]
[0024] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including: receiving a video frame for residual coding; determining whether sign data hiding is enabled at the picture level of the video frame; turning on sign data hiding at the slice level of the video frame in response to the determination that sign data hiding is enabled at the picture level of the video frame; determining whether sign data hiding is turned off at the slice level of the video frame; and turning off transform-skip residual coding at the slice level of the video frame in response to the determination that sign data hiding is turned off at the slice level of the video frame.
[0025]
[0025] Embodiments of the present disclosure further provide a non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded in lossless mode at the slice level; and turning off one or more loop filters at the slice level in response to the determination that the video frame is coded in lossless mode at the slice level.
[0026]
[0026] Embodiments of the present disclosure further provide a non - transient computer - readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether sign data hiding is turned off at the picture level of the video frame, and in response to the determination that sign data hiding is turned off at the picture level of the video frame, turning off conversion skip residual encoding at the slice level of the video frame.
[0027]
[0027] Embodiments of the present disclosure further provide a non - transient computer - readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether state - dependent quantization is enabled for the video frame, and in response to the determination that state - dependent quantization is enabled for the video frame, turning off conversion skip residual encoding at the slice level of the video frame.
[0028] Brief Description of the Drawings
[0028] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the figures are not drawn to scale.
Brief Description of the Drawings
[0029] [Figure 1]
[0029] An example structure of a video sequence according to some embodiments of the present disclosure is shown. [Figure 2A]
[0030] A schematic diagram of an example of an encoding process according to some embodiments of the present disclosure is shown. [Figure 2B]
[0031] A schematic diagram of another example of an encoding process according to some embodiments of the present disclosure is shown. [Figure 3A]
[0032] A schematic diagram of an example of a decryption process according to some embodiments of this disclosure is shown. [Figure 3B]
[0033] A schematic diagram of another example of a decryption process according to some embodiments of this disclosure is shown. [Figure 4]
[0034] A block diagram of an example of a device for encoding or decoding video, according to some embodiments of the present disclosure, is shown. [Figure 5]
[0035] The following is an exemplary table showing the support conditions that allow or prohibit SDH of TS and BDPCM blocks according to some embodiments of this disclosure. [Figure 6A]
[0036] The following are exemplary encoder adjustments of a BDPCM block before adjustment, according to some embodiments of the present disclosure. [Figure 6B]
[0037] The following are exemplary encoder adjustments of the adjusted BDPCM block according to some embodiments of the present disclosure. [Figure 7]
[0038] An exemplary table is provided showing the conditions for disabling sign data hiding according to some embodiments of this disclosure. [Figure 8]
[0039] The following are exemplary syntaxes, including some of the residual coding syntax, according to several embodiments of this disclosure. [Figure 9]
[0040] The following is an exemplary table showing the conditions that allow sign data hiding for conversion skip mode and block difference pulse code modulation mode according to some embodiments of the present disclosure. [Figure 10]
[0041] The following are exemplary syntaxes, including a portion of the residual coding syntax for the conditions shown in Figure 9, according to some embodiments of this disclosure. [Figure 11]
[0042] The following are exemplary syntaxes, including portions of residual coding syntax for disabling sign data hiding, according to some embodiments of this disclosure. [Figure 12]
[0043] The following are exemplary syntaxes, including a portion of the slice header syntax for controlling the slice-level sign data hiding flag, according to some embodiments of the present disclosure. [Figure 13]
[0044] The following are exemplary syntaxes, including portions of residual coding syntax for controlling slice-level sign data hiding flags, according to some embodiments of the present disclosure. [Figure 14]
[0045] The following are exemplary syntaxes, including a portion of the syntax for a slice header for a slice-level sign data hiding flag, according to some embodiments of this disclosure. [Figure 15A]
[0046] The following are exemplary syntaxes, including some of the syntax for a slice header for a slice-level reversible flag, according to several embodiments of this disclosure. [Figure 15B]
[0046] Exemplary syntax, including a portion of the syntax for a slice header for a slice-level reversible flag, according to some embodiments of the present disclosure is shown. [Figure 16]
[0047] The following are exemplary syntaxes, including portions of the residual coding syntax for slice-level reversible flags, according to some embodiments of the present disclosure. [Figure 17]
[0048] The following are exemplary syntaxes, including some of the syntax for slice headings with reduced syntax redundancy, according to some embodiments of the present disclosure. [Figure 18]
[0049] This disclosure provides exemplary syntax, including some of the syntax for slice headers for sign data hiding and state-dependent quantization conditions, according to several embodiments of this disclosure. [Figure 19]
[0050] A flowchart shows an example of a video encoding method using a conversion skip mode and sign data hiding according to some embodiments of this disclosure. [Figure 20]
[0051] A flowchart shows an example of a video encoding method using block difference pulse code modulation mode and sign data hiding according to some embodiments of the present disclosure. [Figure 21]
[0052] A flowchart shows an example of a video coding method using conversion skip residual coding and sign data hiding according to some embodiments of the present disclosure. [Figure 22]
[0053] A flowchart shows an example of a video encoding method using conversion-skip residual coding and sign data hiding at the picture level, according to some embodiments of the present disclosure. [Figure 23]
[0054] A flowchart shows an example of a video encoding method using sign data hiding at the picture level, sign data hiding at the slice level, and transformation skip residual coding at the slice level, according to some embodiments of the present disclosure. [Figure 24]
[0055] A flowchart shows an example of a video encoding method using a lossless encoding mode and sign data hiding according to some embodiments of this disclosure. [Figure 25]
[0056] A flowchart shows an example of a video encoding method using sign data hiding at the picture level and transform skip residual coding at the slice level, according to some embodiments of this disclosure. [Figure 26]
[0057] A flowchart shows an example of a video coding method using state-dependent quantization and slice-level transformation-skip residual coding, according to some embodiments of this disclosure. [Modes for carrying out the invention]
[0030] Detailed explanation
[0058] Herein, exemplary embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, where, unless otherwise noted, the same number in different drawings represents the same or similar elements. The implementations described below in the description of exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with the embodiments relating to the present invention enumerated in the accompanying claims. Specific embodiments of this disclosure are described in more detail below. In the event of any conflict between terms and definitions used herein and terms and / or definitions incorporated by reference, the description herein shall prevail.
[0031]
[0059] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0032]
[0060] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, the Joint Video Experts Team ("JVET") has developed a technology that surpasses HEVC using the Joint Exploration Model ("JEM") reference software. Because the encoding technology was incorporated into JEM, JEM achieved significantly improved encoding performance compared to HEVC. VCEG and MPEG have also officially begun development of next-generation video compression standards that surpass HEVC.
[0033]
[0061] The VVC standard is a recent development and continues to incorporate more encoding techniques to provide better compression performance. VVC is based on the same hybrid video encoding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, and so on.
[0034]
[0062] An image is a set of still pictures (or frames) arranged in chronological order to store visual information. An image capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and an image playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display such pictures in chronological order. Furthermore, in some applications, such as surveillance, conferencing, or live broadcasting, the image capture device can transmit the captured image in real time to an image playback device (e.g., a computer with a monitor).
[0035]
[0063] To reduce the memory space and transmission bandwidth required for such applications, video can be compressed. For example, video can be compressed before storage and transmission, and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. Modules or circuits for compression are generally called "encoders," and modules or circuits for decompression are generally called "decoders." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as various appropriate hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders include program code, computer executable instructions, firmware, or any appropriate computer implementation algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. In some applications, a codec can decompress video using a first encoding standard and then recompress the decompressed video using a second encoding standard; in this case, the codec can be called a "transcoder."
[0036]
[0064] A video encoding process can identify and retain useful information that can be used to reconstruct a picture. If the video encoding process cannot completely reconstruct information that has been ignored, it can be called "lossy." Otherwise, the video encoding process can be called "lossy." Most encoding processes are lossy, which is a trade-off to reduce the required memory space and transmission bandwidth.
[0037]
[0065] In many cases, useful information in an encoded picture (referred to as the "current picture") may include changes relative to a reference picture (e.g., a previously encoded or reconstructed picture). Such changes may include changes in pixel position, brightness, or color. Changes in the position of a group of pixels representing an object may reflect the movement of the object between the reference picture and the current picture.
[0038]
[0066] A picture that is encoded without referencing another picture (i.e., a picture is its own reference picture) is called an "I-picture". A picture is called a "P-picture" if some or all of the blocks within it (for example, a block that generally refers to a portion of a video picture) are predicted using intra-prediction or inter-prediction with one reference picture (e.g., one-way prediction). A picture is called a "B-picture" if at least one block within it is predicted using two reference pictures (e.g., two-way prediction).
[0039]
[0067] Figure 1 shows the structure of an example video sequence according to some embodiments of the present disclosure. As shown in Figure 1, the video sequence 100 can be live video or captured and archived video. The video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).
[0040]
[0068] As shown in Figure 1, the video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are sequential, with more pictures between picture 106 and picture 108. In Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and as indicated by the arrow, its reference picture is picture 102. Picture 106 is a B picture, and as indicated by the arrow, its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. Please note that the reference pictures 102-106 are merely examples, and this disclosure is not limited to the examples shown in Figure 1 of the reference pictures.
[0041]
[0069] Typically, a video codec does not encode or decode the entire picture at once because the computation of such a task is complex. More precisely, a video codec can divide the picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to in this disclosure as Basic Processing Units ("BPUs"). For example, structure 110 in Figure 1 shows an example of the structure of a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 Basic Processing Units, the boundaries of which are indicated by dashed lines. In some embodiments, Basic Processing Units may be called "macroblocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or "Coding Tree Units" ("CTUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have pixels of variable size, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size within a picture. The size and shape of the basic processing unit can be selected for each picture based on the balance between encoding efficiency and level of detail maintained within the basic processing unit.
[0042]
[0070] A basic processing unit can be a logical unit that may contain various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture may include a luminance component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements in which the lumina and chroma components may have basic processing units of the same size. In some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC), the lumina and chroma components may be called a "coding tree block" (CTB). Any operation performed on a basic processing unit can be repeated on its lumina and chroma components, respectively.
[0043]
[0071] Video encoding involves multiple operational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the basic processing unit may still be too large to process, and therefore can be further divided into segments referred to in this disclosure as “basic processing subunits.” In some embodiments, a basic processing subunit may be referred to as a “block” in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as a “coding unit” ("CU") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Similar to a basic processing unit, a basic processing subunit is also a logical unit and may contain a collection of various types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing subunit can be repeated on its luma and chroma components, respectively. Note that such divisions can be performed at further levels as processing needs require. Also note that various stages can divide the basic processing unit using various methods.
[0044]
[0072] For example, in the mode determination stage (an example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a determination. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs as in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing subunit.
[0045]
[0073] In another example, during the prediction phase (an example of which is shown in Figures 2A and 2B), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at those levels.
[0046]
[0074] In another example, during the transformation phase (an example of which is shown in Figures 2A and 2B), the encoder can perform transformation operations on residual subunits (e.g., CUs). However, in some cases, the subunit may still be too large to process. The encoder can further divide the subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform transformation operations at that level. Note that the division method for the same subunit may differ between the prediction phase and the transformation phase. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU may have different sizes and numbers.
[0047]
[0075] In the structure 110 of Figure 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, with the boundaries indicated by dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.
[0048]
[0076] In some implementations, to provide parallel processing and error tolerance for video encoding and decoding, a picture can be divided into processing regions, so that the encoding or decoding process for one region of the picture does not depend on the information of any other region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can improve encoding efficiency by processing different regions of the picture in parallel. Furthermore, if the data in a region is corrupted during processing or lost during network transmission, the codec can provide error tolerance by correctly encoding or decoding other regions of the same picture without depending on the corrupted or lost data. Some video encoding standards allow a picture to be divided into various types of regions. For example, H.265 / HEVC and H.266 / VVC offer two types of regions: "slices" and "tiles". It should also be noted that the various pictures in video sequence 100 may have various partitioning schemes for dividing the picture into regions.
[0049]
[0077] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, with their boundaries indicated by solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. Note that the basic processing units, basic sub-units, and regions of structure 110 in Figure 1 are merely examples, and this disclosure does not limit its embodiments.
[0050]
[0078] Figure 2A shows a schematic diagram of an example of an encoding process according to some embodiments of the present disclosure. For example, the encoding process 200A shown in Figure 2A can be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 in Figure 1, the video sequence 202 may contain a set of pictures arranged in chronological order (referred to as “original pictures”). Similar to the structure 110 in Figure 1, each original picture in the video sequence 202 can be divided by the encoder into a basic processing unit, a basic processing subunit, or a region for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original picture in the video sequence 202. For example, the encoder can perform process 200A iteratively, in which case the encoder can encode a basic processing unit in a single iteration of process 200A. In some embodiments, the encoder can execute process 200A in parallel for each region of the original picture in the video sequence 202 (e.g., regions 114-118).
[0051]
[0079] In Figure 2A, the encoder can supply the basic processing unit of the original picture of the video sequence 202 (called the "original BPU") to the prediction stage 204 to generate the predicted data 206 and the predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate the residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214 to generate the quantized conversion coefficients 216. The encoder can supply the predicted data 206 and the quantized conversion coefficients 216 to the binary coding stage 226 to generate the video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be called the "forward path". During process 200A, after the quantization stage 214, the encoder can supply the quantized conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction criterion 224, which is used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be called a “reconstruction path”. The reconstruction path can be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0052]
[0080] The encoder can iterate through process 200A to encode each original BPU of the original picture (in the forward path) and generate a predicted criterion 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0053]
[0081] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term “receive” can mean any act of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data.
[0054]
[0082] In prediction stage 204, in the current iteration, the encoder can receive the original BPU and prediction criterion 224 and perform a prediction operation to generate prediction data 206 and the predicted BPU 208. The prediction criterion 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU as the predicted BPU 208 from the prediction data 206 and prediction criterion 224.
[0055]
[0083] Ideally, the predicted BPU208 should be identical to the original BPU. However, due to imperfections in the prediction and reconstruction operations, the predicted BPU208 is generally slightly different from the original BPU. To record such differences, after generating the predicted BPU208, the encoder can subtract it from the original BPU to generate the residual BPU210. For example, the encoder can subtract the pixel values (e.g., grayscale or RGB values) of the predicted BPU208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU210 may have a residual value resulting from such a subtraction between the original BPU and the corresponding pixel of the predicted BPU208. Compared to the original BPU, the predicted data 206 and residual BPU210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0056]
[0084] To further compress the residual BPU 210, in transformation step 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each basis pattern being associated with "transformation coefficients." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency component (e.g., luminance variation frequency) of the residual BPU 210. No basis pattern can be reproduced from any combination of any other basis patterns (e.g., a linear combination). In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the basis functions.
[0057]
[0085] Different transformation algorithms may use different base patterns. Various transformation algorithms can be used in transformation stage 212, such as discrete cosine transform, discrete sine transform, or similar. The transformation in transformation stage 212 can be reversed; that is, the encoder can reconstruct the residual BPU 210 by performing the reverse operation of the transformation (called the "inverse transform"). For example, to reconstruct the pixels of the residual BPU 210, the inverse transform can be performed by multiplying the values of the corresponding pixels in the base pattern by the respective coefficients and adding the products to produce a weighted sum. In video coding standards, both the encoder and decoder can use the same transformation algorithm (and therefore the same base pattern). Thus, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the base pattern from the encoder. While the transformation coefficients may have fewer bits compared to the residual BPU 210, they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU210 is further compressed.
[0058]
[0086] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye generally perceives lower frequency fluctuations better, the encoder can ignore information from higher frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized conversion coefficients 216 by dividing each conversion coefficient by an integer value (called the "quantization scale parameter") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients of the higher frequency basis patterns can be converted to zero, and the conversion coefficients of the lower frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-value quantized conversion coefficients 216, thereby further compressing the conversion coefficients. The quantization process can also be reversed, in which case the quantized conversion coefficients 216 can be reconstructed into conversion coefficients by the reverse operation of quantization (called "inverse quantization").
[0059]
[0087] Because the encoder ignores the remainder of such division in rounding operations, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can be the source of the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized conversion coefficients 216 may require. To obtain varying levels of information loss, the encoder can use quantization scale factors of varying values or any other parameter of the quantization process.
[0060]
[0088] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transformation coefficients 216 using a binary coding technique such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transformation coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transformation in the transformation stage 212, the parameters of the quantization process (e.g., quantization scale factor), and the encoder control parameters (e.g., bitrate control parameter). The encoder may use the output data from the binary coding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0061]
[0089] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transformation coefficients 216 to generate reconstructed transformation coefficients. In the inverse transformation stage 220, the encoder can generate reconstructed residual BPU 222 based on the reconstructed transformation coefficients. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction criterion 224 to be used in the next iteration of process 200A.
[0062]
[0090] It should be noted that the video sequence 202 can be encoded using other variations of process 200A. In some embodiments, the steps of process 200A can be performed in a different order by the encoder. In some embodiments, one or more steps of process 200A can be combined into a single step. In some embodiments, a single step of process 200A can be divided into multiple steps. For example, the conversion step 212 and the quantization step 214 can be combined into a single step. In some embodiments, process 200A can include additional steps. In some embodiments, process 200A can omit one or more steps of Figure 2A.
[0063]
[0091] Figure 2B shows a schematic diagram of another example of an encoding process according to some embodiments of the present disclosure. As shown in Figure 2B, process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0064]
[0092] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict the current BPU using pixels from one or more adjacent BPUs that have already been encoded within the same picture. That is, the prediction criterion 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict the current BPU using regions from one or more already encoded pictures. That is, the prediction criterion 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the temporal redundancy inherent in a picture.
[0065]
[0093] Referring to process 200B, in the forward path, the encoder performs prediction operations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra-prediction. In the original BPU of the picture being encoded, the prediction criterion 224 may include one or more adjacent BPUs within the same picture that are encoded (in the forward path) and reconstructed (in the reconstructed path). The encoder may generate a predicted BPU 208 by extrapolating adjacent BPUs. Extrapolation techniques may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the values of the corresponding pixels for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., to the lower left, lower right, upper left, or upper right of the original BPU), or in any direction defined by the video encoding standard used. In intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the extrapolation parameters, and the orientation of the adjacent BPUs used relative to the original BPU.
[0066]
[0094] In another example, during the temporal prediction stage 2044, the encoder can perform interpretation. In the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstructed path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the encoder can generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the predicted BPU 208. Once all the reconstructed BPUs for the same picture have been generated, the encoder can generate a reconstructed picture as a reference picture. The encoder can perform a "motion estimation" operation to search for a matching region within the scope of the reference picture (referred to as a "search window"). The location of the search window in the reference picture can be determined based on the location of the original BPU in the current picture. For example, the search window can be centered in the reference picture at a location with the same coordinates as the original BPU in the current picture and can be extended over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (for example, by using a pixel recursive algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a match region. The match region may have different dimensions from the original BPU (e.g., smaller, equal to, larger, or a different shape). Since the reference picture and the current picture are separated in time on the timeline (for example, as shown in Figure 1), the match region can be considered to "move" to the original BPU's position over time. The encoder can record the direction and distance of such movement as a "motion vector". When multiple reference pictures are used (for example, as picture 106 in Figure 1), the encoder can search for a match region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the match region for each matching reference picture.
[0067]
[0095] Motion estimation can be used to identify various types of motion, such as translation, rotation, and scaling. In interpretation, the prediction data 206 may include, for example, the location of the matched region (e.g., coordinates), the motion vector associated with the matched region, the number of reference pictures, the weights associated with the reference pictures, and so on.
[0068]
[0096] To generate the predicted BPU 208, the encoder can perform a “motion compensation” operation. Using motion compensation, the predicted BPU 208 can be reconstructed based on the prediction data 206 (e.g., motion vectors) and the prediction criterion 224. For example, the encoder can move the matching region of a reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 in Figure 1), the encoder can move the matching region of each reference picture according to its respective motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder can add up the weighted sum of the pixel values of the moved matching regions.
[0069]
[0097] In some embodiments, interpretation can be unidirectional or bidirectional. Unidirectional interpretation can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in Figure 1 is a unidirectional interpretation-predicted picture, with the reference picture (i.e., picture 102) preceding picture 104. Bidirectional interpretation can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in Figure 1 is a bidirectional interpretation-predicted picture, with the reference pictures (i.e., pictures 104 and 108) being in both temporal directions relative to picture 104.
[0070]
[0098] Continuing to refer to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate distortion optimization technique, in which case the encoder may select a prediction mode to minimize the value of the cost function, depending on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate the corresponding predicted BPU 208 and predicted data 206.
[0071]
[0099] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed within the current picture), the encoder can directly feed the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). The encoder can feed the prediction criterion 224 to the loop filtering stage 232, where the encoder can apply loop filters to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused during the encoding of the prediction criterion 224. In the loop filtering stage 232, the encoder can apply various loop filtering techniques, such as deblocking, sample-adaptive offset, adaptive loop filtering, etc. The loop-filtered reference picture can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for subsequent pictures in video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the loop filter parameters (e.g., the loop filter strength) in the binary coding stage 226, along with the quantized transformation coefficients 216, the prediction data 206, and other information.
[0072]
[0100] Figure 3A shows a schematic diagram of an example of a decoding process according to some embodiments of the present disclosure. As shown in Figure 3A, process 300A can be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder can execute process 300A at the level of a basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can iterate through process 300A, in which case the decoder can decode a basic processing unit in a single iteration of process 300A. In some embodiments, the decoder can execute process 300A in parallel for each region of picture (e.g., regions 114-118) to be encoded within the video bitstream 228.
[0073]
[0101] In Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (called the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantized transformation coefficients 216. The decoder can supply the quantized transformation coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the predicted criterion 224. In some embodiments, the predicted criterion 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the predicted criterion 224 to the prediction stage 204 for performing the prediction operation in the next iteration of process 300A.
[0074]
[0102] The decoder can iteratively perform process 300A to decode each encoded BPU of the encoded picture and generate a predicted criterion 224 for encoding the next encoded BPU of the encoded picture. After decoding all encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0075]
[0103] In the binary decoding stage 302, the decoder can perform the inverse transform of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transformation coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, type of transformation, parameters of the quantization process (e.g., quantization scale factor), encoder control parameters (e.g., bitrate control parameter), and so on. In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder can depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0076]
[0104] Figure 3B shows a schematic diagram of another example of a decoding process according to some embodiments of the present disclosure. As shown in Figure 3B, process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0077]
[0105] In process 300B, during decoding, the encoded basic processing unit ("current BPU") of the encoded picture ("current picture") can contain various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction was used by the encoder to encode the current BPU, the prediction data 206 can contain prediction mode indicators (e.g., flag values) indicating the intra-prediction, parameters of the intra-prediction operation, etc. Parameters of the intra-prediction operation can include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as a reference, the size of the adjacent BPUs, extrapolation parameters, the orientation of the adjacent BPU relative to the original BPU, etc. In another example, if inter-prediction was used by the encoder to encode the current BPU, the prediction data 206 can contain prediction mode indicators (e.g., flag values) indicating the inter-prediction, parameters of the inter-prediction operation, etc. The parameters for the interpretation operation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the locations (e.g., coordinates) of one or more matching regions within each reference picture, and one or more motion vectors associated with each matching region.
[0078]
[0106] Based on the prediction mode index, the decoder can determine whether to perform a spatial prediction (e.g., intra-prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter-prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are explained in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate the predicted BPU 208. As explained in Figure 3A, the decoder can generate the prediction criterion 224 by adding the predicted BPU 208 and the reconstructed residual BPU 222.
[0079]
[0107] In process 300B, the decoder can supply the predicted criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing the prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder can supply the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture from which all BPUs have been decoded), the decoder can supply the prediction criterion 224 to the loop filter stage 232 to reduce or remove distortion (e.g., blocking artifacts). The decoder can apply the loop filter to the prediction criterion 224 in the manner described in Figure 2B. The loop-filtered reference picture can be stored in buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for subsequent encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, the prediction data includes loop filter parameters when the prediction mode index of the prediction data 206 indicates that inter-prediction was used to encode the current BPU.
[0080]
[0108] There can be four types of loop filters. For example, loop filters can include deblocking filters, sample adaptive offset (SAO) filters, luma mapping with chroma scaling (LMCS) filters, and adaptive loop filters (ALF). The order in which the four types of loop filters are applied can be LMCS filter, deblocking filter, SAO filter, and then ALF. An LMCS filter can have two principal components. The first component can be an in-loop mapping of the luma component based on an adaptive piecewise linear model. The second component can be for the chroma component and can apply luma-dependent chroma residual scaling.
[0081]
[0109] Figure 4 shows a block diagram of an example of an apparatus for encoding or decoding video according to some embodiments of the present disclosure. As shown in Figure 4, the apparatus 400 may include a processor 402. When the processor 402 executes instructions described herein, the apparatus 400 can become a dedicated device for video encoding or decoding. The processor 402 can be any type of circuitry capable of handling or processing information. For example, the processor 402 may include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP (Intellectual Property) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), composite programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), systems on a chip (SoCs), application-specific integrated circuits (ASICs), and so on. In some embodiments, the processor 402 may also be a set of processors grouped together as a single logical component. For example, as shown in Figure 4, the processor 402 may include a plurality of processors, including processor 402a, processor 402b, and processor 402n.
[0082]
[0110] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for implementing stages in processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and processing data (e.g., via the bus 410), execute the program instructions, and perform operations or handling on the processing data. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number of random-access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash® (CF) cards, and any combination thereof. Memory 404 can also be a group of memories (not shown in Figure 4) grouped together as a single logical component.
[0083]
[0111] Bus 410 can be a communication device that transfers data between components within the device 400, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a USB (Universal Serial Bus) port, a PCI (Peripheral Component Interconnect) express port).
[0084]
[0112] To ensure clarity and avoid ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits” in this disclosure. The data processing circuits can be implemented entirely in hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits can be a single, independent module, or they can be fully or partially integrated into any other component of the device 400.
[0085]
[0113] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth® adapters, infrared adapters, near-field communication ("NFC") adapters, cellular network chips, and any combination thereof.
[0086]
[0114] In some embodiments, the apparatus 400 may further include a peripheral device interface 408 to provide connectivity to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces communicatively coupled to a video archive), and the like.
[0087]
[0115] It should be noted that the video codec (for example, the codec that runs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. In another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0088]
[0116] In quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in Figure 2A or 2B, and inverse quantization 218 in Figure 3A or 3B), the quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residuals. The initial QP value used for encoding the picture or slice can be signaled at a high level, for example, using the syntax element init_qp_minus26 in the Picture Parameter Set (PPS) and the syntax element slice_qp_delta in the slice header. Furthermore, the QP value can be fitted at a local level for each CU using the delta QP value sent in the granularity of the quantization group.
[0089]
[0117] In VVC Transform-Skip mode, residual blocks (e.g., the difference between the original block and the predicted block) can be directly quantized and entropy-encoded. The transformation process can be bypassed in Transform-Skip (TS) mode. For example, the variable transform_skip_flag can be signaled at the transformation block level to indicate whether TS mode has been selected for processing. TS mode can be efficient for lossless compression. For example, TS mode can be efficient for camera capture or screen content sequences. In the case of lossy compression, TS mode can also improve the compression process for certain types of video content, such as computer-generated images or graphics mixed with camera image content (e.g., scrolling text). A transformation block is a block of samples resulting from a transformation in the decoding process, where the transformation is the process by which blocks of transformation coefficients are transformed into blocks of spatial domain values.
[0090]
[0118] In addition to TS mode, VVC also employs Block Differential Pulse-Code Modulation (BDPCM) mode. In BDPCM mode, residual blocks can be directly quantized, and the delta between the quantized residual and its predicted quantized value can be entropy encoded. The predicted quantized value may be horizontal or vertical. The variable bdpcm_flag can be transmitted at the CU level to indicate whether BDPCM is applied. If BDPCM is applied, another flag can be sent to signal the direction of the BDPCM mode (e.g., horizontal or vertical). In some examples, if BDPCM mode is selected, the value of transform_skip_flag can be inferred to be 1, signaling that the transformation process is bypassed for the current block.
[0091]
[0119] In VVC (for example, VVC draft 8), state-dependent scalar quantization can be used in addition to scalar quantization. In state-dependent scalar quantization, the set of acceptable reconstruction values for a transformation coefficient is determined by the values of the transformation coefficient levels that precede the current transformation coefficient level in the reconstruction order. State-dependent quantization ("DQ") can be enabled at the sequence level using the Sequence Parameter Set ("SPS") level variable sps_dep_quant_enabled_flag. If the variable sps_dep_quant_enabled_flag is equal to 1, another picture-level variable ph_dep_quant_enabled_flag can be sent to indicate that scalar quantization is applied to the picture.
[0092]
[0120] Sign Data Hiding (SDH) is a mechanism in HEVC, or VVC (e.g., VVC Draft 8), to reduce the number of positive and negative signs that are encoded. For each coefficient group (CG), the encoding of the positive or negative sign of the last non-zero coefficient (e.g., in reverse scan order) can simply be omitted when SDH is enabled. Instead, the positive or negative sign values can be embedded in the parity of the sum of non-zero coefficient levels within the CG using predefined rules. For example, the sum of even numbers may correspond to positive parity (e.g., "+"), and the sum of odd numbers may correspond to negative parity (e.g., "-"). One metric for using SDH is the distance between the first and last non-zero coefficients of a CG in scan order. For example, if this distance is 4 or greater, SDH is used for that CG. In VVC (for example, VVC Draft 8), there is an SPS level gating variable sps_sign_data_hiding_enabled_flag that determines whether SDH is enabled for the current video sequence. If the variable sps_sign_data_hiding_enabled_flag is equal to 1, another picture level variable pic_sign_data_hiding_enabled_flag can be signaled in the picture header to indicate whether SDH is enabled for that picture.
[0093]
[0121] DQ and SDH can be incompatible with each other. Therefore, the VVC specification (e.g., VVC Draft 8) does not allow both DQ and SDH to be enabled for the same video sequence (e.g., sps_dep_quant_enabled_flag is equal to 1 and sps_sign_data_hiding_enabled_flag is equal to 1). For example, sps_sign_data_hiding_enabled_flag can only be signaled if sps_dep_quant_enabled_flag is equal to 0. If sps_dep_quant_enabled_flag is equal to 1, then sps_sign_data_hiding_enabled_flag is presumed to be 0.
[0094]
[0122] VVC coding (e.g., VVC Draft 8) offers two residual coding methods: the standard residual coding method (e.g., residual_coding) and the transform-skip residual coding method (residual_ts_coding). In standard residual coding, the sign of each non-zero coefficient is coded in bypass mode in the third scan. The last sign in the CG can be coded or hidden, depending on whether SDH is enabled for the CG. Both TS and BDPCM blocks can choose between standard residual coding and TS residual coding. If the slice level flag or variable slice_ts_residual_coding_disabled_flag has a value equal to 0, blocks coded in TS mode and BDPCM mode in that slice will choose residual_ts_coding as the residual coding process for the block. If the value of the slice level flag slice_ts_residual_coding_disabled_flag is equal to 1, the TS-coded and BDPCM-coded blocks of that slice will select the normal residual coding method (e.g., residual_coding) as the residual coding process for the block.
[0095]
[0123] If both of the following conditions are met, i.e., 1) the variable slice_ts_residual_coding_disabled_flag is equal to 1, and 2) the variable pic_sign_data_hiding_enabled_flag is equal to 1, then TS blocks using non-BDPCM and TS blocks using BDPCM are able to use SDH. When SDH is enabled, each CG of an entropy-encoded or decoded block must satisfy at least one of the following two conditions: i.e., 1) the sum of the absolute values of the coefficients is even and the sign of the upper left coefficient is positive, or 2) the sum of the absolute values of the coefficients is odd and the sign of the upper left coefficient is negative. If none of the CGs satisfy either of the above conditions, the encoder may adjust the absolute value of one of the coefficients in the CG so that at least one of the above conditions is always met.
[0096]
[0124] Figure 5 shows an exemplary table containing supporting conditions for allowing or disabling SDH in TS and BDPCM blocks according to some embodiments of the present disclosure. As shown in Figure 5, SDH is disabled when pic_sign_data_hiding_enabled_flag has a value of 1 and slice_ts_residual_coding_disabled_flag has a value of 0. SDH is enabled when pic_sign_data_hiding_enabled_flag has a value of 1 and slice_ts_residual_coding_disabled_flag has a value of 1.
[0097]
[0125] There are many challenges with the current design (e.g., VVC Draft 8). Firstly, the VVC design ensures that the BDP blocks use SDH to guarantee the aforementioned SDH condition, even without an efficient coding algorithm to adjust the coefficient values of the BDP blocks. Figure 6A shows exemplary encoder adjustment of a BDP block before adjustment, according to some embodiments of the present disclosure. Figure 6B shows exemplary encoder adjustment of a BDP block after adjustment, according to some embodiments of the present disclosure. The coefficients before adjustment are shown in Figure 6A. As shown in Figure 6A, the sum of the coefficients of the horizontal BDPCM block is odd (e.g., 211), and the sign of the upper left coefficient (e.g., the sign of the number 14) is positive. This does not satisfy the requirements for SDH. As a result, adjustment is required during coding. The coefficients after adjustment are shown in Figure 6B. As shown in Figure 6B, the encoder is adjusted so that the sum of the absolute values is even (e.g., 212), and the value -21 (shown in bold) is changed to -22. As a result, CG can satisfy the requirements for SDH. However, changing one coefficient value can potentially affect many more coefficients. As shown in Figure 6B, the values -12, -2, 4, -1, -1, and -1 shown in Figure 6A have all been changed (shown in bold), indicating that the error has propagated. This error propagation makes BDPCM with SDH less efficient in terms of compression performance.
[0098]
[0126] Another challenge with the current design of VVC (e.g., VVC Draft 8) is its ability to perform lossless compression. In lossless compression, normal residual coding (e.g., the variable slice_ts_residual_coding_disabled_flag equals 1) can achieve a higher compression gain than TS residual coding (e.g., the variable slice_ts_residual_coding_disabled_flag equals 0). As a result, one important use of the condition slice_ts_residual_coding_disabled_flag equals 1 is lossless compression. Since SDH is a lossy coding tool, it cannot always produce lossless results. To achieve lossless compression, it may be necessary to prohibit sign data hiding. As an alternative, the combination of slice_ts_residual_coding_disabled_flag == 1 and sign data hiding may be prohibited.
[0099]
[0127] Syntax redundancy is another issue. The VVC specification (e.g., VVC Draft 8) supports the combination of the conditions that slice_ts_residual_coding_disabled_flag is equal to 1 and pic_sign_data_hiding_enabled_flag is equal to 1. However, based on the aforementioned shortcomings, the combination of normal residual coding (e.g., slice_ts_residual_coding_disabled_flag is equal to 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag is equal to 1) is not a useful configuration. Consequently, syntax redundancy can be reduced by prohibiting this combination.
[0100]
[0128] Embodiments of this disclosure provide methods for addressing the problems described above. In some embodiments, SDH is disabled for both non-BDPCM and BDPCM TS blocks, regardless of the value of slice_ts_residual_coding_disabled_flag. Figure 7 shows an exemplary table containing conditions for disabling sign data hiding according to some embodiments of this disclosure. As shown in Figure 7, the variable transform_skip_flag is set to 1 regardless of the value of the variable slice_ts_residual_coding_disabled_flag. It is asserted that if the block is encoded in BDPCM mode, it is inferred that the variable transform_skip_flag is 1.
[0101]
[0129] Figure 8 shows exemplary syntax, including some of the residual coding syntax, according to several embodiments of the present disclosure. As shown in Figure 8, changes from the previous VVC are shown in bold italics, and syntax to be removed is further indicated by strikethrough. For example, sign data hiding is disabled when the variable transform_skip_flag is equal to 1. Furthermore, the redundancy condition check for the variable ph_dep_quant_enabled_flag can be eliminated. According to the VVC specification, there is no valid case where both flags pic_sign_data_hiding_enabled_flag and ph_dep_quant_enabled_flag are valued at 1.
[0102]
[0130] In some embodiments, SDH is disabled if the block is encoded in BDPCM mode (e.g., the variable BdpcmFlag is equal to 1), regardless of the value of the variable slice_ts_residual_coding_disabled_flag. However, if the variable slice_ts_residual_coding_disabled_flag is equal to 1, SDH for TS blocks using non-BDPCM mode can be permitted. If slice_ts_residual_coding_disabled_flag is equal to 0, SDH for TS blocks is disabled (with or without BDPCM). Figure 9 shows an exemplary table containing the conditions for permitting sign data hiding for conversion skip mode and block difference pulse code modulation mode according to some embodiments of the present disclosure. As shown in Figure 9, in non-BDPCM mode (e.g., the variable BdpcmFlag is equal to 0), SDH can be enabled if the variable slice_ts_residual_coding_disabled_flag is equal to 1.
[0103]
[0131] Figure 10 shows exemplary syntax, including a portion of the residual coding syntax for the conditions shown in Figure 9, according to several embodiments of the present disclosure. As shown in Figure 10, changes from the previous VVC are shown in bold italics, and syntax to be removed is further indicated by strikethrough. For example, SDH is disabled if BdpcmFlag is equal to 0 (e.g., signHidden is equal to 0). Furthermore, the redundancy condition check for the variable ph_dep_quant_enabled_flag can be removed from the syntax.
[0104]
[0132] In some embodiments, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, sign data hiding is disabled for all encoded blocks (e.g., both TS blocks and non-TS blocks). This is because when the variable slice_ts_residual_coding_disabled_flag is equal to 1, it is more likely that the slice will be attempted to be encoded in reversible mode, in which case SDH may not be suitable. Figure 11 shows exemplary syntax, including some of the syntax for residual coding to disable sign data hiding according to some embodiments of the present disclosure. As shown in Figure 11, changes from the previous VVC are shown in bold italics, and syntax to be removed is further shown in strikethrough. For example, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, SDH is disabled.
[0105]
[0133] In some embodiments, a slice-level sign data hiding flag (e.g., the variable slice_sign_data_hiding_enabled_flag) can be introduced to control sign data hiding for a slice. For example, if the variable slice_sign_data_hiding_enabled_flag is equal to 0, sign bit hiding is disabled for the current slice. If the variable slice_sign_data_hiding_enabled_flag is equal to 1, sign bit hiding is enabled for the current slice. In some embodiments, if the variable slice_sign_data_hiding_enabled_flag does not exist, it is inferred to be equal to 0.
[0106]
[0134] In some embodiments, both sign data hiding and slice_ts_residual_coding_disabled_flag == 1 are disabled for a given slice. More specifically, the following combinations, namely: a.slice_sign_data_hiding_enabled_flag is equal to 1, and The combination of b.slice_ts_residual_coding_disabled_flag being equal to 1 is not permitted.
[0107]
[0135] In some embodiments, the variable slice_sign_data_hiding_enabled_flag is signaled if both of the following conditions are met: 1) the variable pic_sign_data_hiding_enabled_flag is equal to 1, i.e., the current picture allows SDH, and 2) slice_ts_residual_coding_disabled_flag is equal to 0, i.e., the slice is likely encoded in an irreversible mode in which SDH can be a useful tool.
[0108]
[0136] Figure 12 shows exemplary syntax, including a portion of the slice header syntax for controlling the slice-level sign data hiding flag, according to several embodiments of the present disclosure. Changes from the previous VVC are shown in bold italics, as shown in Figure 10. For example, a new variable, slice_sign_data_hiding_enabled_flag, has been added, as shown in Figure 10.
[0109]
[0137] In some embodiments, if slice_sign_data_hiding_enabled_flag is not equal to 1, the slice level slice_ts_residual_coding_disabled_flag is signaled. This is one way of preventing the activation of both slice_ts_residual_coding_disabled_flag and slice_sign_data_hiding_enabled_flag in the same slice.
[0110]
[0138] Figure 13 shows exemplary syntax, including some of the syntax for residual coding for controlling slice-level sign data hiding flags according to several embodiments of the present disclosure. As shown in Figure 13, changes from the previous VVC are shown in bold, and syntax to be removed is further indicated in strikethrough. For example, when the variable slice_sign_data_hiding_enabled_flag is equal to 0, SDH is disabled regardless of the value of the variables ph_dep_quant_enabled_flag or pic_sign_data_hiding_enabled_flag.
[0111]
[0139] In some embodiments, the variable slice_sign_data_hiding_enabled_flag can be signaled before the signaling of the variable slice_ts_residual_coding_disabled_flag, and the variable slice_ts_residual_coding_disabled_flag can be conditionally signaled if the variable slice_sign_data_hiding_enabled_flag is equal to 0. Figure 14 shows exemplary syntax, including a portion of the syntax of a slice header for a slice-level sign data hiding flag according to some embodiments of the present disclosure. Changes from the previous VVC are shown in bold, as shown in Figure 14. For example, slice_sign_data_hiding_enabled_flag can be set if pic_sign_data_hiding_enabled_flag is equal to 1. As shown in Figure 14, the signaling of the variable `slice_sign_data_hiding_enabled_flag` at the slice level is determined by the variable `pic_sign_data_hiding_enabled_flag` at the picture level. Furthermore, the signaling of the variable `slice_ts_residual_coding_disabled_flag` is determined by the variable `slice_sign_data_hiding_enabled_flag`.
[0112]
[0140] In some embodiments, the signaling for the variable `slice_sign_data_hiding_enabled_flag` and the signaling for the variable `slice_ts_residual_coding_disabled_flag` can be processed independently of each other. In other words, one signaling is independent of the other. In some embodiments, this may necessitate the encoder sending a value of 0 for the variable `slice_sign_data_hiding_enabled_flag` even if the encoder is already in reversible mode. This can result in syntax redundancy. Nevertheless, this still provides the flexibility to use normal residual coding and SDH together for TS and BDPCM blocks to improve the efficiency of irreversible coding in the future.
[0113]
[0141] In some embodiments, the slice-level variable slice_ts_residual_coding_disabled_flag can be replaced with a slice-level reversible variable, i.e., slice_lossless_flag. In some embodiments, a value of 1 for the variable slice_lossless_flag indicates that the current slice is reversibly coded and all residual blocks in that slice analyze the residual samples using the residual_coding() syntax. A value of 0 for the variable slice_lossless_flag indicates that the current slice is not reversibly coded. In some embodiments, when the variable slice_lossless_flag does not exist, it is presumed to be equal to 0.
[0114]
[0142] Figure 15 shows exemplary syntax, including a portion of the slice header syntax for a slice-level lossless flag, according to several embodiments of the present disclosure. As shown in Figure 15, changes from the previous VVC are shown in bold italics, and syntax to be removed is further indicated by strikethrough. For example, the variable slice_lossless_flag is added to and incorporated into the syntax. The variable slice_lossless_flag is signaled after slice type signaling. If the variable slice_lossless_flag is equal to 1, some or all of the loop filters (e.g., adaptive loop filters, sample-adaptive offsets, deblocking filters, and luma mapping with chroma scaling) can be disabled for lossless slices. For example, if the variable slice_lossless_flag is equal to 1, the variables slice_alf_enabled_flag, slice_sao_luma_flag, slice_deblocking_filter_override_flag, and slice_lmcs_enabled_flag will not be signaled.
[0115]
[0143] Figure 16 shows exemplary syntax, including a portion of the syntax for residual coding for slice-level lossless flags, according to several embodiments of the present disclosure. As shown in Figure 16, changes from the previous VVC are shown in bold italics, and syntax to be removed is further indicated by strikethrough. For example, the variable slice_lossless_flag is added to and incorporated into the syntax. In some embodiments, as shown in Figure 16, the variable slice_lossless_flag can be replaced with the variable slice_ts_residual_coding_disabled_flag when determining the residual coding condition.
[0116]
[0144] In some embodiments, it is impossible for both the values of the variables pic_sign_data_hiding_enabled_flag and slice_ts_residual_coding_disabled_flag to be equal to 1. When pic_sign_data_hiding_enabled_flag is equal to 1, the codec is almost certainly operating in lossy mode, and the probability of using slice_ts_residual_coding_disabled_flag being equal to 0 is very high in lossy compression. As a result, to reduce syntax redundancy, it can be inferred that if the variable pic_sign_data_hiding_enabled_flag is equal to 1, then the variable slice_ts_residual_coding_disabled_flag is 0. The variable slice_ts_residual_coding_disabled_flag is signaled only if the variable pic_sign_data_hiding_enabled_flag is equal to 0. Figure 17 shows exemplary syntax, including some of the syntax for slice headings with reduced syntax redundancy, according to several embodiments of the present disclosure. As shown in Figure 17, changes from the previous VVC are indicated in bold italics. For example, when the variable pic_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is equal to 1.
[0117]
[0145] In some embodiments, it is not possible to simultaneously support sign data hiding and state-dependent quantization for a given picture. As a result, the variables ph_dep_quant_enabled_flag and pic_sign_data_hiding_enabled_flag may not be equal to 1 at the same time. To avoid this combination, the variable slice_ts_residual_coding_disabled_flag is signaled when the variable ph_dep_quant_enabled_flag is equal to 1. Figure 18 shows exemplary syntax, including some of the syntax for slice headers for the conditions of sign data hiding and state-dependent quantization according to some embodiments of the present disclosure. As shown in Figure 18, changes from the previous VVC are indicated in bold italics. For example, as shown in Figure 18, the variable slice_ts_residual_coding_disabled_flag is signaled when the variable ph_dep_quant_enabled_flag is equal to 1.
[0118]
[0146] In some embodiments, state-dependent quantization and slice_ts_residual_coding_disabled_flag == 1 cannot be supported simultaneously for a given slice. To avoid this combination, the variable slice_ts_residual_coding_disabled_flag is not signaled when the variable ph_dep_quant_enabled_flag is equal to 1.
[0119]
[0147] In some embodiments, slice_ts_residual_coding_disabled_flag is signaled if both slice_sign_data_hiding_enabled_flag and ph_dep_quant_enabled_flag are non-zero values. More specifically, these are described below. if (!slice_sign_data_hiding_enabled_flag && ! ph_dep_quant_enabled_flag) signal _residual_coding_disabled_flag
[0120]
[0148] Embodiments of this disclosure further provide methods for performing video coding. Figure 19 shows a flowchart of an example of a video coding method using a conversion skip mode and sign data hiding according to some embodiments of this disclosure. In some embodiments, the method 19000 shown in Figure 19 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 19000 shown in Figure 19 can be performed according to the syntax shown in Figure 8. In some embodiments, the method 19000 shown in Figure 19 is performed according to the VVC standard.
[0121]
[0149] In step S19010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0122]
[0150] In step S19020, it is determined whether the video frame has been encoded according to the transform skip mode at the transform block level. For example, as shown in Figure 8, the variable transform_skip_flag can be used to determine whether the video frame has been encoded according to the transform skip mode at the transform block level.
[0123]
[0151] In step S19030, in response to the determination that the video frame has been encoded according to the transform skip mode at the transform block level, sign data hiding is turned off for residual coding. For example, as shown in Figure 8, when the variable transform_skip_flag is equal to 0, the variable signHidden is set to 0. As a result, sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding does not depend on whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 8, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden does not depend on the variable ph_dep_quant_enabled_flag. In some embodiments, turning off sign data hiding for both TS blocks using non-BDPCM and TS blocks using BDPCM can improve the efficiency of compression performance. For example, the error propagation shown in Figure 6 may be eliminated by turning off sign data hiding.
[0124]
[0152] Figure 20 shows a flowchart of an example of a video coding method using block-difference pulse code modulation mode and sign data hiding according to some embodiments of the present disclosure. In some embodiments, the method 20000 shown in Figure 20 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 20000 shown in Figure 20 can be performed according to the syntax shown in Figure 10. In some embodiments, the method 20000 shown in Figure 20 is performed according to the VVC standard.
[0125]
[0153] In step S20010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0126]
[0154] In step S20020, it is determined whether the video frame is encoded according to a block-difference pulse code modulation mode. In some embodiments, it is determined whether the video frame is encoded according to a block-difference pulse code modulation mode at the block level. For example, as shown in Figure 10, the variable BdpcmFlag can be used to determine whether the video frame is encoded according to a block-difference pulse code modulation mode at the block level.
[0127]
[0155] In step S20030, in response to the determination that the video frame is encoded according to block-difference pulse code modulation mode, sign data hiding is turned off for residual coding. For example, as shown in Figure 10, when the variable BdpcmFlag is equal to 0, the variable signHidden is set to 0. As a result, sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding does not depend on whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 10, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden does not depend on the variable ph_dep_quant_enabled_flag. In some embodiments, turning off sign data hiding for TS blocks using BDPCM can improve the efficiency of compression performance. For example, the error propagation shown in Figure 6 may be eliminated by turning off sign data hiding.
[0128]
[0156] Figure 21 shows a flowchart of an example of a video coding method using transform-skip residual coding and sign data hiding according to some embodiments of the present disclosure. In some embodiments, the method 21000 shown in Figure 21 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 21000 shown in Figure 21 can be performed according to the syntax shown in Figure 11. In some embodiments, the method 21000 shown in Figure 21 is performed according to the VVC standard.
[0129]
[0157] In step S21010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0130]
[0158] In step S21020, it is determined whether the video frame is encoded at the slice level according to the transform-skip residual coding mode. For example, as shown in Figure 11, the variable slice_ts_residual_coding_disabled_flag can be used to determine whether the video frame is encoded at the slice level according to the transform-skip residual coding mode.
[0131]
[0159] In step S21030, in response to the determination that the video frame is not encoded according to the transform-skip residual coding mode at the slice level, sign data hiding is turned off for residual coding. For example, as shown in Figure 11, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, it is determined that the video frame is not encoded according to the transform-skip residual coding mode at the slice level. As a result, the variable signHidden is set to 0, and sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding does not depend on whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 11, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden does not depend on the variable ph_dep_quant_enabled_flag. In some embodiments, sign data hiding is a tool for lossy coding. Standard residual coding (for example, when slice_ts_residual_coding_disabled_flag is equal to 1) can achieve a higher compression gain than TS residual coding. Therefore, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, the compression being performed can be lossless. As a result, sign data hiding can be turned off.
[0132]
[0160] Figure 22 shows a flowchart of an example of a video coding method using conversion-skip residual coding and sign data hiding at the picture level, according to some embodiments of the present disclosure. In some embodiments, the method 22000 shown in Figure 22 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 22000 shown in Figure 22 can be performed according to the syntax shown in Figure 12 or the syntax shown in Figure 13. In some embodiments, the method 22000 shown in Figure 22 is performed according to the VVC standard.
[0133]
[0161] In step S22010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0134]
[0162] In step S22020, it is determined whether sign data hiding is enabled at the picture level of the video frame and whether transform skip residual coding is disabled at the slice level of the video frame. For example, as shown in Figure 12, the variable pic_sign_data_hiding_enabled_flag can be used to determine whether sign data hiding is enabled at the picture level of the video frame, and the variable slice_ts_residual_coding_disabled_flag can be used to determine whether the video frame is coded according to the transform skip residual coding mode at the slice level of the video frame.
[0135]
[0163] In step S22030, in response to the determination that sign data hiding is enabled at the picture level of the video frame and transform-skip residual coding is enabled at the slice level of the video frame, sign data hiding is turned on at the slice level of the video frame. For example, as shown in Figure 12, when the variable pic_sign_data_hiding_enabled_flag is equal to 1, it is determined that sign data hiding is enabled at the picture level of the video frame. Furthermore, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, it is determined that the video frame is not coded according to the transform-skip residual coding mode at the slice level. As a result, the variable slice_sign_data_hiding_enabled_flag is set to 1, and sign data hiding is turned on at the slice level. In some embodiments, sign data hiding is a tool for lossy coding. Normal residual coding (e.g., slice_ts_residual_coding_disabled_flag is equal to 1) can achieve a higher compression gain than transform-skip residual coding. Therefore, transformation-skipping residual coding can be used for lossy compression. As a result, sign data hiding can be enabled.
[0136]
[0164] In some embodiments, method 22000 further includes steps S22040 and S22050. In step S22040, it is determined whether sign data hiding is turned off at the slice level of the video frame. For example, as shown in Figure 13, the variable slice_sign_data_hiding_enabled_flag can be examined to determine whether sign data hiding is turned off at the slice level of the video frame.
[0137]
[0165] In step S22050, in response to the determination that sign data hiding is turned off at the slice level of the video frame, sign data hiding is turned off for residual coding. For example, as shown in Figure 13, when the variable slice_sign_data_hiding_enabled_flag is equal to 0, the variable signHidden is set to 0, and sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding does not depend on whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 13, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden does not depend on the variable ph_dep_quant_enabled_flag.
[0138]
[0166] Figure 23 shows a flowchart of an example of a video coding method using sign data hiding at the picture level, sign data hiding at the slice level, and transform skip residual coding at the slice level, according to some embodiments of the present disclosure. In some embodiments, the method 23000 shown in Figure 23 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 23000 shown in Figure 23 can be performed according to the syntax shown in Figure 14. In some embodiments, the method 23000 shown in Figure 23 is performed according to the VVC standard.
[0139]
[0167] In step S23010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0140]
[0168] In step S23020, it is determined whether sign data hiding is enabled at the picture level of the video frame. For example, as shown in Figure 14, the variable pic_sign_data_hiding_enabled_flag can be used to determine whether sign data hiding is enabled at the picture level of the video frame.
[0141]
[0169] In step S23030, in response to the determination that sign data hiding is enabled at the picture level of the video frame, sign data hiding is turned on at the slice level of the video frame. For example, as shown in Figure 14, when the variable pic_sign_data_hiding_enabled_flag is equal to 1, it is determined that sign data hiding is enabled at the picture level of the video frame. As a result, the variable slice_sign_data_hiding_enabled_flag is set to 1, and sign data hiding is turned on at the slice level.
[0142]
[0170] In step S23040, it is determined whether sign data hiding is turned off at the slice level of the video frame. For example, as shown in Figure 14, the variable slice_sign_data_hiding_enabled_flag is checked to determine whether sign data hiding is turned off at the slice level of the video frame.
[0143]
[0171] In step S23050, in response to the determination that sign data hiding is turned off at the slice level of the video frame, transform-skip residual coding is turned off at the slice level of the video frame. For example, as shown in Figure 14, if the variable slice_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is set to 1, and transform-skip residual coding is turned off at the slice level of the video frame. In some embodiments, sign data hiding is a tool for lossy coding. Normal residual coding (e.g., slice_ts_residual_coding_disabled_flag is equal to 1) can achieve a higher compression gain than transform-skip residual coding. Therefore, transform-skip residual coding can be used for lossy compression. When sign data hiding is turned off, transform-skip residual coding at the slice level can also be turned off.
[0144]
[0172] Figure 24 shows a flowchart of an example of a video coding method using a lossless coding mode and sign data hiding according to some embodiments of the present disclosure. In some embodiments, the method 24000 shown in Figure 24 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 24000 shown in Figure 24 can be performed according to the syntax shown in Figure 15 or the syntax shown in Figure 16. In some embodiments, the method 24000 shown in Figure 24 is performed according to the VVC standard.
[0145]
[0173] In step S24010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0146]
[0174] In step S24020, it is determined whether the video frame is encoded in lossless mode. In some embodiments, it is determined at the slice level whether the video frame is encoded in lossless mode. For example, as shown in Figure 15, the variable slice_lossless_flag can be used to determine whether the video frame is encoded at the slice level in lossless mode.
[0147]
[0175] In step S24030, in response to the determination that the video frame is encoded in lossless mode at the slice level, one or more loop filters are turned off at the slice level. For example, as shown in Figure 15, if the variable slice_lossless_flag is equal to 1, the variables slice_alf_enabled_flag, slice_sao_luma_flag, slice_deblocking_filter_override_flag, and slice_lmcs_enabled_flag are not signaled. In some embodiments, turning off one or more loop filters can make the video encoding more efficient.
[0148]
[0176] Figure 25 shows a flowchart of an example of a video coding method using sign data hiding at the picture level and transform skip residual coding at the slice level, according to some embodiments of the present disclosure. In some embodiments, the method 25000 shown in Figure 25 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 25000 shown in Figure 25 can be performed according to the syntax shown in Figure 17. In some embodiments, the method 25000 shown in Figure 25 is performed according to the VVC standard.
[0149]
[0177] In step S25010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0150]
[0178] In step S25020, it is determined whether sign data hiding is turned off at the picture level of the video frame. For example, as shown in Figure 17, the variable pic_sign_data_hiding_enabled_flag can be used to determine whether sign data hiding is turned off at the picture level of the video frame.
[0151]
[0179] In step S25030, in response to the determination that sign data hiding is turned off at the picture level of the video frame, transform skip residual coding is turned off at the slice level of the video frame. For example, as shown in Figure 17, the variable pic_sign_data_hiding_enabled_flag is checked. If the variable pic_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is signaled. As a result, transform skip residual coding is turned off at the slice level of the video frame. In some embodiments, the combination of normal residual coding (e.g., slice_ts_residual_coding_disabled_flag == 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag == 1) is not a useful configuration. Therefore, syntax redundancy can be reduced by prohibiting this combination.
[0152]
[0180] Figure 26 shows a flowchart of an example of a video coding method using state-dependent quantization and slice-level transform-skip residual coding, according to some embodiments of the present disclosure. In some embodiments, the method 26000 shown in Figure 26 can be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 26000 shown in Figure 26 can be performed according to the syntax shown in Figure 18. In some embodiments, the method 26000 shown in Figure 26 is performed according to the VVC standard.
[0153]
[0181] In step S26010, the video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0154]
[0182] In step S26020, it is determined whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 18, the variable ph_dep_quant_enabled_flag can be used to determine whether state-dependent quantization is enabled for the video frame.
[0155]
[0183] In step S26030, in response to the determination that state-dependent quantization is enabled for the video frame, transform-skip residual coding is turned off at the slice level of the video frame. For example, as shown in Figure 18, the variable ph_dep_quant_enabled_flag is checked. If the variable ph_dep_quant_enabled_flag is equal to 1, the variable slice_ts_residual_coding_disabled_flag is signaled. As a result, transform-skip residual coding is turned off at the slice level of the video frame. In some embodiments, the combination of normal residual coding (e.g., slice_ts_residual_coding_disabled_flag == 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag == 1) is not a useful configuration. Therefore, syntax redundancy can be reduced by prohibiting this combination.
[0156]
[0184] In some embodiments, non-temporary computer-readable storage media containing instructions are also provided, and the instructions may be executed by devices (such as disclosed encoders and decoders) for performing the methods described above. Common forms of non-temporary media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, and EPROMs, FLASH®-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. Devices may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0157]
[0185] In this specification, relational terms such as “first” and “second” are used solely to distinguish one entity or action from another, and do not imply or require any actual relationship or order between these entities or actions. Furthermore, “comprising,” “having,” “containing,” and “including,” as well as other similar forms of words, are intended to be equivalent in meaning and are intended to be non-restrictive, not intended to be an exhaustive list of such items, nor to be limited to only the listed items.
[0158]
[0186] As used herein, unless otherwise specified, the term “or” encompasses all possible combinations, except in cases where it is impractical. For example, if we say that a database may contain A or B, then unless otherwise specified or impractical, that database may contain A, B, A and B. As a second example, if we say that a database may contain A, B, or C, then unless otherwise specified or impractical, that database may contain A, B, C, A and B, A and C, B and C, A and B and C.
[0159]
[0187] It will be understood that the embodiments described above can be implemented by hardware or software (program code), or a combination of hardware and software. When implemented by software, the software can be stored in the computer-readable medium described above. When executed by a processor, the software can perform the methods disclosed. The computing units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into a single module / unit, and that each of the above modules / units can be further divided into multiple submodules / subunits.
[0160]
[0188] The above specification has described embodiments with respect to numerous specific details that may vary depending on the implementation. Certain adaptations and modifications can be made to the embodiments described. Other embodiments may become apparent to those skilled in the art by examining this specification and practicing the invention disclosed herein. This specification and examples are to be considered merely illustrative, and the true scope and spirit of the invention are intended to be shown by the following claims. The order of steps shown in the figures is also intended to be for illustrative purposes only and is not intended to limit the sequence of steps to any particular order. Therefore, those skilled in the art will understand that these steps can be performed in different orders while implementing the same method.
[0161]
[0189] The embodiments can be further described using the following clauses. 1. A video encoding method, Receiving video frames for residual coding, To determine whether the video frame is encoded according to the conversion skip mode at the conversion block level, In response to the determination that the video frame is encoded according to the conversion skip mode, sign data hiding is turned off for residual coding, A video encoding method that includes this. 2. The video encoding method described in Clause 1, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 3. The video encoding method described in Clause 1, wherein the video frames are contained within the bitstream. 4. The video encoding method described in Clause 1, performed in accordance with the VVC (versatile video coding) standard. 5. A video encoding method, Receiving video frames for residual coding, To determine whether the video frame is encoded according to block difference pulse code modulation mode, In response to the determination that the video frame is encoded according to block difference pulse code modulation mode, sign data hiding is turned off for residual coding, A video encoding method that includes this. 6. Determining whether a video frame is encoded according to a block-difference pulse code modulation mode is The video encoding method described in Clause 5 further includes determining whether a video frame is encoded at the block level according to a block-difference pulse code modulation mode. 7. The video encoding method described in Clause 5, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 8. A video encoding method as described in Clause 5, wherein the video frames are contained within the bitstream. 9. The video encoding method described in Clause 5, performed in accordance with the VVC (versatile video coding) standard. 10. A video encoding method, Receiving video frames for residual coding, To determine whether a video frame is encoded at the slice level according to the conversion skip residual coding mode, In response to the determination that a video frame has not been encoded according to the transform skip residual coding mode at the slice level, sign data hiding is turned off for residual coding, A video encoding method that includes this. 11. The video coding method described in Clause 10, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 12. The video encoding method described in Clause 10, wherein the video frames are contained within the bitstream. 13. The video encoding method described in Clause 10, performed in accordance with the VVC (versatile video coding) standard. 14. A video encoding method, Receiving video frames for residual coding, Determine whether sign data hiding is enabled at the picture level of the video frame, and whether conversion skip residual coding is disabled at the slice level of the video frame. In response to the determination that sign data hiding is enabled at the picture level of the video frame and transform skip residual coding is enabled at the slice level of the video frame, sign data hiding is turned on at the slice level of the video frame, A video encoding method that includes this. 15. Determine whether sign data hiding is turned off at the slice level of the video frame, In response to the determination that sign data hiding is turned off at the slice level of the video frame, sign data hiding is turned off for residual coding, The video encoding method described in Clause 14, which further includes the following: 16. The video encoding method described in Clause 14, wherein the video frames are contained within the bitstream. 17. A video encoding method as described in Clause 14, performed in accordance with the VVC (versatile video coding) standard. 18. A video encoding method, Receiving video frames for residual coding, Determine whether sign data hiding is enabled at the picture level of the video frame, In response to the determination that sign data hiding is enabled at the picture level of the video frame, sign data hiding is turned on at the slice level of the video frame, Determine whether sign data hiding is turned off at the slice level of the video frame, In response to the determination that sign data hiding is turned off at the video frame slice level, transform skip residual coding is turned off at the video frame slice level, A video encoding method that includes this. 19. A video encoding method as described in Clause 18, wherein the video frames are contained within a bitstream. 20. A video encoding method as described in Clause 18, performed in accordance with the VVC (versatile video coding) standard. 21. A video encoding method, Receiving video frames for residual coding, To determine whether the video frame is encoded in lossless mode at the slice level, In response to the determination that a video frame is encoded in lossless mode at the slice level, one or more loop filters are turned off at the slice level. A video encoding method that includes this. 22. The video encoding method described in Clause 21, wherein the video frames are contained within the bitstream. 23. A video encoding method as described in Clause 21, performed in accordance with the VVC (versatile video coding) standard. 24. A video encoding method, Receiving video frames for residual coding, Determine whether sign data hiding is turned off at the picture level of the video frame, In response to the determination that sign data hiding is turned off at the picture level of the video frame, transform skip residual coding is turned off at the slice level of the video frame, A video encoding method that includes this. 25. The video encoding method described in Clause 24, wherein the video frames are contained within the bitstream. 26. A video encoding method as described in Clause 24, performed in accordance with the VVC (versatile video coding) standard. 27. A video encoding method, Receiving video frames for residual coding, Determining whether state-dependent quantization is enabled for video frames, In response to the determination that state-dependent quantization is enabled for the video frame, transform skip residual coding is turned off at the slice level of the video frame, A video encoding method that includes this. 28. A video encoding method as described in Clause 27, wherein the video frames are contained within a bitstream. 29. A video encoding method as described in Clause 27, performed in accordance with the VVC (versatile video coding) standard. 30. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether video frames are encoded according to the conversion skip mode at the conversion block level, and In response to the determination that the video frame is encoded according to the conversion skip mode, turn off sign data hiding for residual coding. A system configured to perform the following action. 31. The system described in Clause 30, in which turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 32. A system as described in Clause 30, in which video frames are contained within a bitstream. 33. The system described in Clause 30, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 34. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether a video frame is encoded according to a block difference pulse code modulation mode, and In response to the determination that the video frame is encoded according to block-difference pulse code modulation mode, sign data hiding is turned off for residual coding. A system configured to perform the following action. 35. The processor executes a set of instructions to the system. The system described in Clause 34, further configured to perform a determination of whether a video frame is encoded at the block level according to a block-difference pulse code modulation mode. 36. The system described in Clause 34, in which turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 37. A system as described in Clause 34, in which video frames are contained within a bitstream. 38. A system as described in Clause 34, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 39. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether a video frame is encoded at the slice level according to the conversion skip residual coding mode, and In response to the determination that a video frame has not been encoded at the slice level according to the conversion skip residual coding mode, sign data hiding is turned off for residual coding. A system configured to perform the following action. 40. The system described in Clause 39, in which turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 41. A system as described in Clause 39, in which video frames are contained within a bitstream. 42. A system as described in Clause 39, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 43. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, Determine whether sign data hiding is enabled at the picture level of the video frame, and whether conversion skip residual coding is disabled at the slice level of the video frame, and In response to the determination that sign data hiding is enabled at the picture level of the video frame and that conversion skip residual coding is enabled at the slice level of the video frame, turn on sign data hiding at the slice level of the video frame. A system configured to perform the following action. 44. The processor executes a set of instructions to the system. To determine whether sign data hiding is turned off at the slice level of the video frame, and In response to the determination that sign data hiding is turned off at the slice level of the video frame, turn off sign data hiding for residual coding. The system described in Clause 43, which is further configured to perform the following actions. 45. A system as described in Clause 43, in which video frames are contained within a bitstream. 46. A system as described in Clause 43, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 47. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether sign data hiding is enabled at the picture level of the video frame. In response to the determination that sign data hiding is enabled at the picture level of the video frame, sign data hiding is turned on at the slice level of the video frame. To determine whether sign data hiding is turned off at the slice level of the video frame, and In response to the determination that sign data hiding is turned off at the video frame slice level, turn off transform skip residual coding at the video frame slice level. A system configured to perform the following action. 48. A system as described in Clause 47, in which video frames are contained within a bitstream. 49. A system as described in Clause 47, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 50. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether the video frame is encoded in lossless mode at the slice level, and In response to the determination that a video frame is encoded in lossless mode at the slice level, one or more loop filters are turned off at the slice level. A system configured to perform the following action. 51. A system as described in Clause 50, in which video frames are contained within a bitstream. 52. The system described in Clause 50, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 53. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, To determine whether sign data hiding is turned off at the picture level of the video frame. In response to the determination that sign data hiding is turned off at the picture level of the video frame, turn off transform skip residual coding at the slice level of the video frame. A system configured to perform the following action. 54. A system as described in Clause 53, in which video frames are contained within a bitstream. 55. A system as described in Clause 54, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 56. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, and the processor executes a set of instructions to the system. Receiving video frames for residual coding, Determining whether state-dependent quantization is enabled for a video frame. In response to the determination that state-dependent quantization is enabled for the video frame, the transformation skip residual coding is turned off at the slice level of the video frame. A system configured to perform the following action. 57. A system as described in Clause 56, in which video frames are contained within a bitstream. 58. A system as described in Clause 56, in which residual coding is performed in accordance with the VVC (versatile video coding) standard. 59. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, To determine whether the video frame is encoded according to the conversion skip mode at the conversion block level, In response to the determination that the video frame is encoded according to the conversion skip mode, sign data hiding is turned off for residual coding, Non-temporary computer-readable media, including [specific examples of such media]. 60. Non-transient computer-readable media as described in Clause 59, where turning off sign data hiding for residual coding is not dependent on whether state-dependent quantization is enabled for the video frame. 61. A non-temporary computer-readable medium as defined in Clause 59, in which video frames are contained within a bitstream. 62. Non-temporary computer-readable media as described in Clause 59, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 63. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, To determine whether the video frame is encoded according to block difference pulse code modulation mode, In response to the determination that the video frame is encoded according to block difference pulse code modulation mode, sign data hiding is turned off for residual coding, Non-temporary computer-readable media, including [specific examples of such media]. 64. A set of instructions is provided to a computer system. To further perform the determination of whether the video frame is encoded at the block level according to the block-difference pulse code modulation mode, A non-temporary computer-readable medium as described in Clause 63, which is executable by at least one processor of a computer system. 65. Non-transient computer-readable media as described in Clause 63, where turning off sign data hiding for residual coding is not dependent on whether state-dependent quantization is enabled for the video frame. 66. A non-temporary computer-readable medium as defined in Clause 63, in which video frames are contained within a bitstream. 67. Non-temporary computer-readable media as described in Clause 63, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 68. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, To determine whether a video frame is encoded at the slice level according to the conversion skip residual coding mode, In response to the determination that a video frame has not been encoded according to the transform skip residual coding mode at the slice level, sign data hiding is turned off for residual coding, Non-temporary computer-readable media, including [specific examples of such media]. 69. Non-transient computer-readable media as described in Clause 68, where turning off sign data hiding for residual coding is not dependent on whether state-dependent quantization is enabled for the video frame. 70. A non-temporary computer-readable medium as defined in Clause 68, in which video frames are contained within a bitstream. 71. Non-temporary computer-readable media as described in Clause 68, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 72. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, Determine whether sign data hiding is enabled at the picture level of the video frame, and whether conversion skip residual coding is disabled at the slice level of the video frame. In response to the determination that sign data hiding is enabled at the picture level of the video frame and transform skip residual coding is enabled at the slice level of the video frame, sign data hiding is turned on at the slice level of the video frame, Non-temporary computer-readable media, including [specific examples of such media]. 73. A set of instructions is provided to a computer system. To determine whether sign data hiding is turned off at the slice level of the video frame, and In response to the determination that sign data hiding is turned off at the slice level of the video frame, further action is taken to turn off sign data hiding for residual coding. A non-temporary computer-readable medium as described in Clause 72, which is executable by at least one processor of a computer system. 74. A non-temporary computer-readable medium as defined in Clause 72, in which video frames are contained within a bitstream. 75. Non-temporary computer-readable media as described in Clause 72, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 76. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, Determine whether sign data hiding is enabled at the picture level of the video frame, In response to the determination that sign data hiding is enabled at the picture level of the video frame, sign data hiding is turned on at the slice level of the video frame, Determine whether sign data hiding is turned off at the slice level of the video frame, In response to the determination that sign data hiding is turned off at the video frame slice level, transform skip residual coding is turned off at the video frame slice level, Non-temporary computer-readable media, including [specific examples of such media]. 77. A non-temporary computer-readable medium as defined in Clause 76, in which video frames are contained within a bitstream. 78. Non-temporary computer-readable media as described in Clause 76, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 79. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, To determine whether the video frame is encoded in lossless mode at the slice level, In response to the determination that a video frame is encoded in lossless mode at the slice level, one or more loop filters are turned off at the slice level. Non-temporary computer-readable media, including [specific examples of such media]. 80. A non-temporary computer-readable medium as defined in Clause 79, in which video frames are contained within a bitstream. 81. Non-temporary computer-readable media as described in Clause 79, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 82. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, Determine whether sign data hiding is turned off at the picture level of the video frame, In response to the determination that sign data hiding is turned off at the picture level of the video frame, transform skip residual coding is turned off at the slice level of the video frame, Non-temporary computer-readable media, including [specific examples of such media]. 83. A non-temporary computer-readable medium as described in Clause 82, in which video frames are contained within a bitstream. 84. Non-temporary computer-readable media as described in Clause 82, wherein the method is performed in accordance with the VVC (versatile video coding) standard. 85. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, Determining whether state-dependent quantization is enabled for video frames, In response to the determination that state-dependent quantization is enabled for the video frame, transform skip residual coding is turned off at the slice level of the video frame, Non-temporary computer-readable media, including [specific examples of such media]. 86. A non-temporary computer-readable medium as described in Clause 85, in which video frames are contained within a bitstream. 87. Non-temporary computer-readable media as described in Clause 85, wherein the method is performed in accordance with the VVC (versatile video coding) standard.
[0162]
[0190] The drawings and this specification disclose exemplary embodiments. However, many variations and modifications can be made to these embodiments. Accordingly, although certain terms are used, they are used only in a general and descriptive sense and not for limiting purposes.
Claims
1. A video encoding method, Receiving video frames for residual coding, Determining whether the video frame is encoded according to a first encoding mode, In response to determining whether the video frame is encoded according to the first encoding mode, sign data hiding is turned off for the residual encoding, A video encoding method that includes this.
2. The first coding mode is a transform skip mode at the transform block level, and the sign data hiding is turned off for the residual coding. The video encoding method according to claim 1, comprising turning off sign data hiding for residual encoding in response to a determination that the video frame has been encoded according to the conversion skip mode at the conversion block level.
3. The first coding mode is a block difference pulse code modulation mode, and the sign data hiding is turned off for the residual coding. The video encoding method according to claim 1, comprising turning off sign data hiding for residual encoding in response to a determination that the video frame is encoded according to the block difference pulse code modulation mode.
4. Determining whether the video frame is encoded according to the first encoding mode is: The video encoding method according to claim 3, further comprising determining whether the video frame is encoded at the block level according to the block difference pulse code modulation mode.
5. The first coding mode is a slice-level transform-skip residual coding mode, and the sign data hiding is turned off for the residual coding. The video encoding method according to claim 1, comprising turning off sign data hiding for residual encoding in response to a determination that the video frame has not been encoded according to the conversion skip residual encoding mode at the slice level.
6. The video coding method according to claim 1, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame.
7. The video encoding method according to claim 1, wherein the video frame is contained within a bitstream.
8. The video encoding method according to claim 1, performed in accordance with the VVC (versatile video coding) standard.
9. A system for performing video data processing, Memory that stores a set of instructions, The system includes a processor, the processor executes the set of instructions to the system, Receiving video frames for residual coding, To determine whether the video frame is encoded according to a first encoding mode, and In response to determining whether the video frame is encoded according to the first encoding mode, sign data hiding is turned off for the residual encoding. A system configured to perform the following action.
10. The first encoding mode is a conversion skip mode at the conversion block level, and the processor executes the set of instructions to the system, The system according to claim 9, configured to perform the action of turning off sign data hiding for residual coding in response to a determination that the video frame has been encoded according to the conversion skip mode at the conversion block level.
11. The first encoding mode is a block difference pulse code modulation mode, and the processor executes the set of instructions to the system, The system according to claim 9, configured to perform the action of turning off the sign data hiding for residual coding in response to a determination that the video frame is encoded according to the block difference pulse code modulation mode.
12. Determining whether the video frame is encoded according to the first encoding mode is: The system according to claim 11, further comprising determining whether the video frame is encoded at the block level according to the block difference pulse code modulation mode.
13. The first coding mode is a slice-level transform-skip residual coding mode, and the processor executes the set of instructions to the system, The system according to claim 9, configured to perform the action of turning off sign data hiding for residual coding in response to a determination that the video frame has not been coded according to the conversion skip residual coding mode at the slice level.
14. The system according to claim 9, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame.
15. The system according to claim 9, wherein the aforementioned video frame is contained within a bitstream.
16. The system according to claim 9, wherein the residual coding is performed in accordance with the VVC (versatile video coding) standard.
17. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to start a method for performing video data processing, Receiving video frames for residual coding, Determining whether the video frame is encoded according to a first encoding mode, In response to determining whether the video frame is encoded according to the first encoding mode, sign data hiding is turned off for the residual encoding. Non-temporary computer-readable media, including [specific examples of such media].
18. The first encoding mode is a conversion skip mode at the conversion block level, and the set of instructions is configured on the computer system. In response to the determination that the video frame is encoded according to the conversion skip mode at the conversion block level, further cause the sign data hiding to be turned off for the residual coding, The non-temporary computer-readable medium according to claim 17, which is executable by at least one processor of the computer system.
19. The first encoding mode is a block difference pulse code modulation mode, and the set of instructions is transmitted to the computer system. In response to the determination that the video frame is encoded according to the block difference pulse code modulation mode, further cause the sign data hiding to be turned off for the residual coding, The non-temporary computer-readable medium according to claim 17, which is executable by at least one processor of the computer system.
20. The first coding mode is a slice-level transform-skip residual coding mode, and the set of instructions is configured on a computer system. In response to the determination that the video frame is not encoded according to the conversion skip residual coding mode at the slice level, further cause the sign data hiding to be turned off for the residual coding, The non-temporary computer-readable medium according to claim 17, which is executable by at least one processor of the computer system.