Sign data hiding in video recording
By selectively disabling sign data hiding in certain encoding modes, the method optimizes video compression efficiency, addressing inefficiencies in advanced coding standards like VVC.
Patent Information
- Application Number
- JP2022554702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-24
- Filing Date
- 2021-03-24
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-03-24
AI Technical Summary
Existing video coding standards face challenges in achieving efficient compression while maintaining quality, particularly with the increasing complexity of standards like VVC, where techniques like sign data hiding can introduce inefficiencies.
The method involves selectively turning off sign data hiding for residual encoding in specific modes such as transform skip, block differential pulse code modulation, and lossless modes, based on encoding conditions, to optimize compression efficiency.
This approach enhances compression efficiency by reducing redundancy and improving encoding performance in video data processing, aligning with the goals of advanced standards like VVC.
Smart Images

Figure 0007807382000001 
Figure 0007807382000002 
Figure 0007807382000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 994,239, filed March 24, 2020. The provisional application is incorporated herein by reference in its entirety.
[0002] Technical Field
[0002] The present disclosure relates generally to video data processing, and more particularly to residual coding of video data. [Background technology]
[0003] background
[0003] A video is a set of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is typically called encoding, and the decompression process is typically called decoding. There are various video coding formats that use standardized video coding techniques, most commonly techniques based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC) standard (e.g., HEVC / H.265), the Versatile Video Coding (VVC) standard (e.g., VVC / H.266), and the AVS standard, which specify specific video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards is increasing. Summary of the Invention [Means for solving the problem]
[0004] Disclosure Overview
[0004] An embodiment of the present disclosure provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a transform skip mode at a transform block level, and turning off sign data hiding for the residual encoding in response to determining that the video frame is encoded according to the transform skip mode.
[0005]
[0005] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a block differential pulse code modulation mode, and turning off sign data hiding for the residual encoding in response to determining that the video frame is encoded according to the block differential pulse code modulation mode.
[0006]
[0006] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a transform skip residual encoding mode at a slice level, and turning off sign data hiding for the residual encoding in response to determining that the video frame is not encoded according to a transform skip residual encoding mode at the slice level.
[0007]
[0007] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether sign data hiding is enabled at a picture level of the video frame and whether transform skip residual encoding is disabled at a slice level of the video frame, and in response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual encoding is enabled at the slice level of the video frame, turning on sign data hiding at the slice level of the video frame.
[0008]
[0008] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether sign data hiding is enabled at a picture level for the video frame, turning on sign data hiding at a slice level for the video frame in response to determining that sign data hiding is enabled at the picture level for the video frame, determining whether sign data hiding is turned off at the slice level for the video frame, and turning off transform skip residual encoding at the slice level for the video frame in response to determining that sign data hiding is turned off at the slice level for the video frame.
[0009]
[0009] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded in a lossless mode at a slice level, and turning off one or more loop filters at the slice level in response to determining that the video frame is encoded in a lossless mode at the slice level.
[0010]
[0010] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether sign data hiding is turned off at the picture level of the video frame, and in response to determining that sign data hiding is turned off at the picture level of the video frame, turning off transform skip residual encoding at the slice level of the video frame.
[0011]
[0011] An embodiment of the present disclosure further provides a method for video data encoding, the method including receiving a video frame for residual encoding, determining whether state-dependent quantization is enabled for the video frame, and in response to determining that state-dependent quantization is enabled for the video frame, turning off transform skip residual encoding at a slice level for the video frame.
[0012]
[0012] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether the video frame is encoded according to a transform skip mode at the transform block level, and in response to determining that the video frame is encoded according to the transform skip mode, turn off sign data hiding for residual encoding.
[0013]
[0013] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether the video frame is encoded according to a block differential pulse code modulation mode, and in response to determining that the video frame is encoded according to the block differential pulse code modulation mode, turn off sign data hiding for the residual encoding.
[0014]
[0014] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory that stores a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether the video frame is encoded according to a transform skip residual encoding mode at the slice level, and in response to determining that the video frame is not encoded according to the transform skip residual encoding mode at the slice level, turn off sign data hiding for the residual encoding.
[0015]
[0015] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory that stores a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether sign data hiding is enabled at the picture level of the video frame and whether transform skip residual encoding is disabled at the slice level of the video frame, and in response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual encoding is enabled at the slice level of the video frame, turn on sign data hiding at the slice level of the video frame.
[0016]
[0016] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory that stores a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding; determine whether sign data hiding is enabled at the picture level of the video frame; in response to determining that sign data hiding is enabled at the picture level of the video frame, turn on sign data hiding at the slice level of the video frame; determine whether sign data hiding is turned off at the slice level of the video frame; and in response to determining that sign data hiding is turned off at the slice level of the video frame, turn off transform skip residual encoding at the slice level of the video frame.
[0017]
[0017] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether the video frame is encoded in a lossless mode at the slice level, and turn off one or more loop filters at the slice level in response to determining that the video frame is encoded in a lossless mode at the slice level.
[0018]
[0018] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether sign data hiding is turned off at the picture level of the video frame, and in response to determining that sign data hiding is turned off at the picture level of the video frame, turn off transform skip residual encoding at the slice level of the video frame.
[0019]
[0019] An embodiment of the present disclosure further provides a system for performing video data processing, the system including a memory storing a set of instructions and a processor, the processor being configured to execute the set of instructions to cause the system to receive a video frame for residual encoding, determine whether state-dependent quantization is enabled for the video frame, and in response to determining that state-dependent quantization is enabled for the video frame, turn off transform skip residual encoding at the slice level of the video frame.
[0020]
[0020] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a transform skip mode at a transform block level, and turning off sign data hiding for the residual encoding in response to determining that the video frame is encoded according to the transform skip mode.
[0021]
[0021] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a block differential pulse code modulation mode, and turning off sign data hiding for the residual encoding in response to determining that the video frame is encoded according to the block differential pulse code modulation mode.
[0022]
[0022] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded according to a transform skip residual encoding mode at the slice level, and turning off sign data hiding for the residual encoding in response to determining that the video frame is not encoded according to the transform skip residual encoding mode at the slice level.
[0023]
[0023] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether sign data hiding is enabled at the picture level of the video frame and whether transform skip residual encoding is disabled at the slice level of the video frame, and in response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual encoding is enabled at the slice level of the video frame, turning on sign data hiding at the slice level of the video frame.
[0024]
[0024] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether sign data hiding is enabled at a picture level of the video frame, turning on sign data hiding at a slice level of the video frame in response to determining that sign data hiding is enabled at the picture level of the video frame, determining whether sign data hiding is turned off at the slice level of the video frame, and turning off transform skip residual encoding at the slice level of the video frame in response to determining that sign data hiding is turned off at the slice level of the video frame.
[0025]
[0025] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether the video frame is encoded in a lossless mode at the slice level, and turning off one or more loop filters at the slice level in response to determining that the video frame is encoded in a lossless mode at the slice level.
[0026]
[0026] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether sign data hiding is turned off at the picture level of the video frame, and in response to determining that sign data hiding is turned off at the picture level of the video frame, turning off transform skip residual encoding at the slice level of the video frame.
[0027]
[0027] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including receiving a video frame for residual encoding, determining whether state-dependent quantization is enabled for the video frame, and in response to determining that state-dependent quantization is enabled for the video frame, turning off transform skip residual encoding at a slice level for the video frame.
[0028] BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0029] [Figure 1] 1 illustrates an example structure of a video sequence according to some embodiments of the present disclosure. [Figure 2A]
[0030] 1 illustrates a schematic diagram of an example encoding process according to some embodiments of the present disclosure. [Figure 2B]
[0031] 10 shows a schematic diagram of another example of an encoding process according to some embodiments of the present disclosure. [Figure 3A]
[0032] 1 illustrates a schematic diagram of an example of a decoding process according to some embodiments of the present disclosure. [Figure 3B]
[0033] 10 shows a schematic diagram of another example of a decoding process according to some embodiments of the present disclosure. [Figure 4]
[0034] 1 shows a block diagram of an example of an apparatus for encoding or decoding video, according to some embodiments of the present disclosure. [Figure 5]
[0035] 1 illustrates an example table including support conditions that allow or prohibit SDH of TS and BDPCM blocks, according to some embodiments of the present disclosure. [Figure 6A]
[0036] 1 illustrates an example encoder adjustment of a BDPCM block before adjustment, according to some embodiments of the present disclosure. [Figure 6B]
[0037] 10 illustrates an example encoder adjustment of a BDPCM block after adjustment, according to some embodiments of the present disclosure. [Figure 7]
[0038] 1 illustrates an example table containing conditions for disabling signature data hiding, according to some embodiments of the present disclosure. [Figure 8]
[0039] 1 illustrates an example syntax, including a portion of the syntax for residual coding, according to some embodiments of the present disclosure. [Figure 9]
[0040] 1 illustrates an example table including conditions that allow sign data hiding for transform skip mode and block differential pulse code modulation mode, according to some embodiments of the present disclosure. [Figure 10]
[0041] 10 illustrates an example syntax including a portion of the syntax for residual coding for the conditions illustrated in FIG. 9 according to some embodiments of the present disclosure. [Figure 11]
[0042] 10 illustrates an example syntax including a portion of a syntax for residual coding for disabling sign data hiding, according to some embodiments of the present disclosure. [Figure 12]
[0043] 10 illustrates an example syntax including a portion of the slice header syntax for control of slice-level sign data hiding flags according to some embodiments of the present disclosure. [Figure 13]
[0044] 10 illustrates example syntax, including a portion of the residual coding syntax, for control of slice-level sign data hiding flags, according to some embodiments of the present disclosure. [Figure 14]
[0045] 10 illustrates an example syntax including a portion of the syntax of a slice header for a slice-level sign data hiding flag, according to some embodiments of the present disclosure. [Figure 15A]
[0046] 10 illustrates example syntax, including a portion of the syntax of a slice header, for a slice-level lossless flag, in accordance with some embodiments of the present disclosure. [Figure 15B] 10 illustrates an example syntax including a portion of the syntax of a slice header for a slice-level lossless flag, according to some embodiments of the present disclosure. [Figure 16]
[0047] 10 illustrates an example syntax including a portion of the residual coding syntax for slice-level lossless flags, according to some embodiments of the present disclosure. [Figure 17]
[0048] 1 illustrates an example syntax including a portion of the syntax for slice headings with reduced syntactic redundancy, according to some embodiments of the present disclosure. [Figure 18]
[0049] 10 illustrates example syntax, including portions of slice header syntax, for sign data hiding and state-dependent quantization conditions, according to some embodiments of the present disclosure. [Figure 19]
[0050] 1 shows a flowchart of an example method for video encoding using transform skip mode and sign data hiding, according to some embodiments of the present disclosure. [Figure 20]
[0051] 1 shows a flowchart of an example of a method for video encoding using block differential pulse code modulation mode and sign data hiding, according to some embodiments of the present disclosure. [Figure 21]
[0052] 1 shows a flowchart of an example method for video encoding using transform skip residual coding and sign data hiding, in accordance with some embodiments of this disclosure. [Figure 22]
[0053] 1 shows a flowchart of an example video encoding method with transform skip residual coding and picture-level sign data hiding, in accordance with some embodiments of this disclosure. [Figure 23]
[0054] 1 shows a flowchart of an example video coding method using sign data hiding at the picture level, sign data hiding at the slice level, and transform skip residual coding at the slice level, in accordance with some embodiments of this disclosure. [Figure 24]
[0055] 1 shows a flowchart of an example method for video encoding using a lossless encoding mode and sign data hiding, according to some embodiments of the present disclosure. [Figure 25]
[0056] 1 shows a flowchart of an example video coding method using sign data hiding at the picture level and transform skip residual coding at the slice level, in accordance with some embodiments of this disclosure. [Figure 26]
[0057] 1 shows a flowchart of an example video coding method using state-dependent quantization and slice-level transform skip residual coding, in accordance with some embodiments of this disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0030] Detailed Description
[0058] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event that terms and definitions set forth herein conflict with terms and / or definitions incorporated by reference, the descriptions in this specification shall control.
[0031]
[0059] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC) (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC) (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.
[0032]
[0060] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, the Joint Video Experts Team ("JVET") has been developing technologies that surpass HEVC using the Joint Exploration Model ("JEM") reference software. Because the coding technology has been incorporated into JEM, JEM has achieved significantly improved coding performance compared to HEVC. VCEG and MPEG have also officially begun development of next-generation video compression standards beyond HEVC.
[0033]
[0061] The VVC standard has recently evolved and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0034]
[0062] A video is a set of still pictures (or frames) arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display these pictures in chronological order. In some applications, such as for surveillance, conferencing, or live broadcasting, the video capture device can transmit the captured video in real time to a video playback device (e.g., a computer with a monitor).
[0035]
[0063] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed. For example, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. Modules or circuitry for compression are commonly referred to as "encoders," and modules or circuitry for decompression are commonly referred to as "decoders." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard, in which case the codec can be called a "transcoder."
[0036]
[0064] A video coding process can identify and retain useful information that can be used to reconstruct a picture. If the information ignored in a video coding process cannot be perfectly reconstructed, the coding process can be called "lossy." Otherwise, the video coding process can be called "lossless." Most coding processes are lossy; this is a trade-off to reduce the required storage space and transmission bandwidth.
[0037]
[0065] Often, useful information about the picture being coded (called the "current picture") may include changes relative to a reference picture (e.g., a previously coded or reconstructed picture). Such changes may include pixel position changes, luminance changes, or color changes. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.
[0038]
[0066] A picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is called an "I-picture." If some or all of the blocks in a picture (e.g., blocks, which generally refer to portions of a video picture) are predicted using intra- or inter-prediction with one reference picture (e.g., unidirectional prediction), the picture is called a "P-picture." If at least one block in a picture is predicted using two reference pictures (e.g., bidirectional prediction), the picture is called a "B-picture."
[0039]
[0067] 1 illustrates an example structure of a video sequence according to some embodiments of the present disclosure. As illustrated in FIG. 1, the video sequence 100 can be live video or captured and archived video. The video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0040]
[0068] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the example shown in FIG.
[0041]
[0069] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, a basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit may have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit may be selected for each picture based on a balance between coding efficiency and the level of detail to be maintained within the basic processing unit.
[0042]
[0070] A basic processing unit may be a logical unit that may include various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements where the luma and chroma components may have the same size basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" (CTBs). Any operation performed on a basic processing unit may be repeatedly performed on each of its luma and chroma components.
[0043]
[0071] Video coding involves multiple operational stages, examples of which are shown in FIGS. 2A-2B and 3A-3B. Because the size of a basic processing unit at each stage may still be too large to process, it can be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, a basic processing sub-unit may be referred to as a "block" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as a "coding unit" (CU) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing sub-unit may have the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing sub-unit is also a logical unit and may include a group of various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on a basic processing sub-unit can be repeatedly performed on each of its luma and chroma components. Note that such division can be performed to further levels, depending on processing needs. Note also that different stages can divide the basic processing unit using different schemes.
[0044]
[0072] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), an encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs as in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.
[0045]
[0073] In another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, the elementary processing sub-units may still be too large to process. The encoder may further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.
[0046]
[0074] In another example, in the transform stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform a transform operation on a residual elementary processing sub-unit (e.g., a CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder may further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. Note that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0047]
[0075] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different ways.
[0048]
[0076] In some implementations, to provide parallel processing and error resilience for video encoding and decoding, a picture can be divided into regions for processing, allowing the encoding or decoding process for one region of a picture to not depend on information from any other region of the picture. In other words, each region of a picture can be processed independently. In this way, a codec can process various regions of a picture in parallel, thereby improving coding efficiency. Also, when data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience. Some video coding standards allow a picture to be divided into various types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that various pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.
[0049]
[0077] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are indicated by solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.
[0050]
[0078] FIG. 2A illustrates a schematic diagram of an example encoding process according to some embodiments of the present disclosure. For example, the encoding process 200A illustrated in FIG. 2A may be performed by an encoder. As illustrated in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform the process 200A at the level of a basic processing unit for each original picture of the video sequence 202. For example, the encoder may perform the process 200A iteratively, in which case the encoder may encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for each original picture region of video sequence 202 (eg, regions 114-118).
[0051]
[0079] 2A , an encoder may provide a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0052]
[0080] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward pass) and generate a predicted reference 224 for encoding the next original BPU of the original picture (in the reconstruction pass). After encoding all of the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0053]
[0081] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any act of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or in any manner for inputting data.
[0054]
[0082] In the prediction stage 204, in the current iteration, the encoder receives the original BPU and a prediction reference 224 and may perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting the prediction data 206, which can be used to reconstruct the original BPU from the prediction data 206 and the prediction reference 224 as a predicted BPU 208.
[0055]
[0083] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. Each pixel of residual BPU 210 may have a residual value that is the result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.
[0056]
[0084] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. No basis pattern can be reproduced from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.
[0057]
[0085] Different transform algorithms may use different basis patterns. The transform stage 212 may use various transform algorithms, such as a discrete cosine transform, a discrete sine transform, or the like. The transform in the transform stage 212 may be invertible. That is, the encoder may reconstruct the residual BPU 210 by inverting the transform (called an "inverse transform"). For example, to reconstruct a pixel of the residual BPU 210, the inverse transform may multiply the value of the corresponding pixel in the basis pattern by each associated coefficient and add the products to produce a weighted sum. In video coding standards, both the encoder and the decoder may use the same transform algorithm (and therefore the same basis pattern). Thus, the encoder may record only the transform coefficients, and the decoder may reconstruct the residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.
[0058]
[0086] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye generally perceives low-frequency fluctuations better, the encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scale parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and some transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients through the inverse operation of quantization (called "dequantization").
[0059]
[0087] Because the encoder ignores any remainder of such division in a rounding operation, quantization stage 214 may be lossy. Typically, quantization stage 214 may be responsible for the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different levels of information loss, the encoder may use different values of the quantization scale factor or any other parameter of the quantization process.
[0060]
[0088] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization scale factor), encoder control parameters (e.g., bitrate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0061]
[0089] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0062]
[0090] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages of FIG. 2A.
[0063]
[0091] 2B shows a schematic diagram of another example of an encoding process according to some embodiments of the present disclosure. As shown in FIG. 2B, process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0064]
[0092] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict a current BPU using pixels from one or more neighboring BPUs already coded in the same picture. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce spatial redundancy inherent in a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict a current BPU using regions from one or more already coded pictures. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce temporal redundancy inherent in a picture.
[0065]
[0093] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstructed path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., to the bottom left, bottom right, top left, or top right of the original BPU), or in any direction defined by the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0066]
[0094] In another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of a current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstructed path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within the reference picture (called a "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pixel-recursive algorithm, a block-matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region can have different dimensions (e.g., smaller, equal, larger, or a different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can search for the matching region in each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region in each matching reference picture.
[0067]
[0095] Motion estimation may be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.
[0068]
[0096] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference picture according to the motion vector, in which case the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as in picture 106 of FIG. 1), the encoder may move the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values of the moved matching region.
[0069]
[0097] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture, with a reference picture (i.e., picture 102) preceding picture 104. Bidirectional inter-prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture, with reference pictures (i.e., pictures 104 and 108) in both temporal directions relative to picture 104.
[0070]
[0098] Continuing with the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.
[0071]
[0099] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating prediction reference 224 (e.g., the current BPU being encoded and reconstructed in the current picture), the encoder may provide prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). The encoder may provide prediction reference 224 to loop filter stage 232, where the encoder may apply a loop filter to prediction reference 224 to reduce or remove distortion (e.g., blocking artifacts) introduced during encoding of prediction reference 224. The encoder may apply various loop filter techniques in loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture may be stored in buffer 234 (or a “decoded picture buffer”) for later use (e.g., for use as an inter-prediction reference picture for future pictures in video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in the binary encoding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.
[0072]
[0100] FIG. 3A shows a schematic diagram of an example of a decoding process according to some embodiments of the present disclosure. As shown in FIG. 3A, process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the basic processing unit (BPU) level for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A iteratively, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for each picture region (e.g., regions 114-118) encoded in video bitstream 228.
[0073]
[0101] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a “coded BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can provide the predicted reference 224 to the prediction stage 204 for performing the prediction operation in the next iteration of the process 300A.
[0074]
[0102] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a predicted reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0075]
[0103] In binary decoding stage 302, the decoder may perform the inverse transform of the binary encoding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a type of transform, parameters of the quantization process (e.g., quantization scale factor), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted over a network in packets, the decoder may depacketize video bitstream 228 before providing it to binary decoding stage 302.
[0076]
[0104] 3B shows a schematic diagram of another example of a decoding process according to some embodiments of the present disclosure. As shown in FIG. 3B, the process 300B can be modified from the process 300A. For example, the process 300B can be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0077]
[0105] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a coded picture (referred to as a "current picture") being decoded, prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data, depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. Parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0078]
[0106] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.
[0079]
[0107] In process 300B, the decoder may provide the predicted reference 224 to a spatial prediction stage 2042 or a temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the decoder may provide the prediction reference 224 to a loop filter stage 232 to reduce or remove distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-prediction reference picture for a future coded picture in the video bitstream 228). The decoder stores one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data includes loop filter parameters.
[0080]
[0108] There may be four types of loop filters. For example, the loop filter may include a deblocking filter, a sample adaptive offset ("SAO") filter, a luma mapping with chroma scaling ("LMCS") filter, and an adaptive loop filter ("ALF"). The order of applying the four types of loop filters may be the LMCS filter, the deblocking filter, the SAO filter, and the ALF. The LMCS filter may include two main components. The first component may be an in-loop mapping of the luma component based on an adaptive piecewise linear model. The second component may be for the chroma component and may apply luma-dependent chroma residual scaling.
[0081]
[0109] FIG. 4 illustrates a block diagram of an example device for encoding or decoding video, according to some embodiments of the present disclosure. As illustrated in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may be a dedicated device for video encoding or decoding. The processor 402 may be any type of circuitry capable of handling or processing information. For example, the processor 402 may include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, processor 402 may also be a set of processors grouped as a single logical entity. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0082]
[0110] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped as a single logical entity.
[0083]
[0111] Bus 410 may be a communication device that transfers data between components within apparatus 400, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus (USB) port, a Peripheral Component Interconnect (PCI) Express port), etc.
[0084]
[0112] For the sake of clarity and avoidance of ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuitry may be a single, independent module, or may be fully or partially integrated into any other component of the device 400.
[0085]
[0113] Device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0086]
[0114] In some embodiments, apparatus 400 may further include a peripheral interface 408 to provide connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface communicatively coupled to a video archive), etc.
[0087]
[0115] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0088]
[0116] The quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 of FIG. 2A or 2B, inverse quantization 218 of FIG. 3A or 3B) use a quantization parameter (QP) to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for coding a picture or slice can be signaled at a high level, for example, using the syntax element init_qp_minus 26 in the Picture Parameter Set (PPS) and the syntax element slice_qp_delta in the slice header. Furthermore, the QP value can be adapted at a local level for each CU using a delta QP value signaled at the granularity of the quantization group.
[0089]
[0117] In VVC transform skip mode, residual blocks (e.g., the difference between the original block and the predicted block) can be directly quantized and entropy coded. The transform process can be bypassed in transform skip ("TS") mode. For example, a variable transform_skip_flag can be signaled at the transform block level to indicate whether TS mode is selected for processing. TS mode can be efficient for lossless compression. For example, TS mode can be efficient for camera capture or screen content sequences. In the case of lossy compression, TS mode can also improve the compression process for certain types of video content, such as computer-generated images or graphics mixed with camera image content (e.g., scrolling text). A transform block is a block of samples resulting from a transformation during the decoding process, where transformation is the process by which a block of transform coefficients is converted into a block of spatial domain values.
[0090]
[0118] In addition to the TS mode, VVC also employs a Block Differential Pulse-Code Modulation (BDPCM) mode. In the BDPCM mode, the residual block can be directly quantized, and the delta between the quantized residual and its predicted quantization value can be entropy coded. The predicted quantization value can be in the horizontal or vertical direction. The variable bdpcm_flag can be transmitted at the CU level to indicate whether BDPCM is applied. If BDPCM is applied, another flag can be sent to signal the direction of the BDPCM mode (e.g., horizontal or vertical). In some examples, when the BDPCM mode is selected, the value of transform_skip_flag can be inferred to be 1, signaling that the transform process is bypassed for the current block.
[0091]
[0119] In addition to scalar quantization, VVC (e.g., VVC draft 8) also allows the use of state-dependent scalar quantization, in which the set of allowable reconstruction values for a transform coefficient depends on the value of the transform coefficient level that precedes the current transform coefficient level in the reconstruction order. The sequence parameter set ("SPS")-level variable sps_dep_quant_enabled_flag can be used to enable state-dependent quantization ("DQ:Dependent Quantization") at the sequence level. When the variable sps_dep_quant_enabled_flag is equal to 1, another picture-level variable ph_dep_quant_enabled_flag can be sent to indicate that scalar quantization is applied to the picture.
[0092]
[0120] Sign Data Hiding (SDH) is a mechanism in HEVC or VVC (e.g., VVC Draft 8) to reduce the number of coded signs. For each coefficient group (CG), coding the sign of the last non-zero coefficient (e.g., in reverse scan order) can simply be omitted when SDH is enabled. Instead, the sign value can be embedded in the parity of the sum of the non-zero coefficient levels in the CG using a predefined convention. For example, an even sum may correspond to positive parity (e.g., "+"), and an odd sum may correspond to negative parity (e.g., "-"). One measure for using SDH is the distance between the first and last non-zero coefficients of the CG in scan order. For example, if this distance is 4 or greater, SDH is used for that CG. In VVC (e.g., VVC Draft 8), there is an SPS level gating variable sps_sign_data_hiding_enabled_flag that determines whether SDH is enabled for the current video sequence. If the variable sps_sign_data_hiding_enabled_flag is equal to 1, another picture level variable pic_sign_data_hiding_enabled_flag can be signaled in the picture header to indicate whether SDH is enabled for that picture.
[0093]
[0121] DQ and SDH may be mutually exclusive. Therefore, the VVC specification (e.g., VVC Draft 8) does not allow both DQ and SDH to be enabled for the same video sequence (e.g., sps_dep_quant_enabled_flag equals 1 and sps_sign_data_hiding_enabled_flag equals 1). For example, sps_sign_data_hiding_enabled_flag can be signaled only if sps_dep_quant_enabled_flag equals 0. If sps_dep_quant_enabled_flag equals 1, sps_sign_data_hiding_enabled_flag is inferred to be 0.
[0094]
[0122] In VVC coding (e.g., VVC Draft 8), there are two residual coding methods: normal residual coding (e.g., residual_coding) and transform skip residual coding (residual_ts_coding). In normal residual coding, the sign of each non-zero coefficient is coded in bypass mode in the third scan path. The last sign in a CG can be coded or hidden depending on whether SDH is enabled for the CG. Both TS and BDPCM blocks can choose between normal residual coding or TS residual coding. If the slice-level flag or variable slice_ts_residual_coding_disabled_flag has a value equal to 0, blocks coded in TS mode and BDPCM mode in that slice select residual_ts_coding as the residual coding process for the block. If the value of the slice-level flag slice_ts_residual_coding_disabled_flag is equal to 1, the TS-coded blocks and BDPCM-coded blocks of that slice select the normal residual coding (eg, residual_coding) method as the residual coding process for the blocks.
[0095]
[0123] If both of the following conditions are met: 1) the variable slice_ts_residual_coding_disabled_flag is equal to 1, and 2) the variable pic_sign_data_hiding_enabled_flag is equal to 1, TS blocks using non-BDPCM and TS blocks using BDPCM are enabled to use SDH. When SDH is enabled, each CG of an entropy coded or decoded block shall satisfy at least one of the following two conditions: 1) the sum of the absolute values of the coefficients is even and the sign of the top-left coefficient is positive, or 2) the sum of the absolute values of the coefficients is odd and the sign of the top-left coefficient is negative. If none of the CGs meets either of the above conditions, the encoder may adjust the absolute value of one of the coefficients in the CG to ensure that one of the above conditions is met.
[0096]
[0124] 5 shows an example table including support conditions for enabling or disabling SDH for TS and BDPCM blocks according to some embodiments of the present disclosure. As shown in FIG. 5, SDH is disabled when pic_sign_data_hiding_enabled_flag has a value of 1 and slice_ts_residual_coding_disabled_flag has a value of 0. SDH is enabled when pic_sign_data_hiding_enabled_flag has a value of 1 and slice_ts_residual_coding_disabled_flag has a value of 1.
[0097]
[0125] There are many challenges with the current design (e.g., VVC Draft 8). First, the VVC design allows BDPCM blocks to use SDH and guarantee the above-mentioned SDH conditions without an efficient encoding algorithm for adjusting the coefficient values of the BDPCM blocks. FIG. 6A shows an example encoder adjustment for a BDPCM block before adjustment, according to some embodiments of the present disclosure. FIG. 6B shows an example encoder adjustment for a BDPCM block after adjustment, according to some embodiments of the present disclosure. The coefficients before adjustment are shown in FIG. 6A. As shown in FIG. 6A, the sum of the coefficients of the horizontal BDPCM block is odd (e.g., 211), and the sign of the top-left coefficient (e.g., the sign of the number 14) is positive. This does not meet the requirements for SDH. As a result, an adjustment is necessary during encoding. The coefficients after adjustment are shown in FIG. 6B. As shown in FIG. 6B, an encoder adjustment was made to change the value −21 (shown in bold) to −22 to make the sum of the absolute values even (e.g., 212). As a result, CG can meet the requirements for SDH. However, changing one coefficient value can affect many more coefficients. As shown in Figure 6B, the values -12, -2, 4, -1, -1, and -1 shown in Figure 6A have all been changed (shown in bold), resulting in error propagation. This error propagation makes BDPCM with SDH less efficient in terms of compression performance.
[0098]
[0126] Another issue with the current design of VVC (e.g., VVC Draft 8) is the ability to perform lossless compression. With lossless compression, regular residual coding (e.g., the variable slice_ts_residual_coding_disabled_flag equals 1) can achieve higher compression gains than TS residual coding (e.g., the variable slice_ts_residual_coding_disabled_flag equals 0). As a result, one important use of the condition slice_ts_residual_coding_disabled_flag equals 1 is lossless compression. Because SDH is a lossy coding tool, it cannot always produce lossless results. To achieve lossless compression, it may be necessary to disable sign data hiding. As an alternative approach, the combination of slice_ts_residual_coding_disabled_flag == 1 and sign data hiding may be disabled.
[0099]
[0127] Syntax redundancy is another issue. The VVC (e.g., VVC Draft 8) specification supports the combination of the condition slice_ts_residual_coding_disabled_flag equal to 1 and the condition pic_sign_data_hiding_enabled_flag equal to 1. However, based on the shortcomings mentioned above, the combination of regular residual coding (e.g., slice_ts_residual_coding_disabled_flag equal to 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag equal to 1) is not a useful configuration. As a result, syntactic redundancy can be reduced by prohibiting this combination.
[0100]
[0128] Embodiments of the present disclosure provide methods for addressing the above-mentioned problems. In some embodiments, SDH is disabled for both non-BDPCM TS blocks and BDPCM TS blocks, regardless of the value of slice_ts_residual_coding_disabled_flag. Figure 7 shows an example table including conditions for disabling sign data hiding according to some embodiments of the present disclosure. As shown in Figure 7, the variable transform_skip_flag is set to 1, regardless of the value of the variable slice_ts_residual_coding_disabled_flag. If the block is coded in BDPCM mode, it is asserted that the variable transform_skip_flag can be inferred to be 1.
[0101]
[0129] 8 illustrates an example syntax, including a portion of the syntax for residual coding, according to some embodiments of the present disclosure. As shown in FIG. 8, changes from the previous VVC are indicated in bold italics, and syntax to be removed is further indicated in strikethrough. For example, if the variable transform_skip_flag is equal to 1, sign data hiding is disabled. Furthermore, the redundant condition check for the variable ph_dep_quant_enabled_flag can be eliminated. According to the VVC specification, there is no valid case where the flags pic_sign_data_hiding_enabled_flag and ph_dep_quant_enabled_flag both have the value 1.
[0102]
[0130] In some embodiments, SDH is disabled if the block is coded in BDPCM mode (e.g., the variable BdpcmFlag is equal to 1), regardless of the value of the variable slice_ts_residual_coding_disabled_flag. However, if the variable slice_ts_residual_coding_disabled_flag is equal to 1, SDH of the TS block using a non-BDPCM mode may be allowed. If the variable slice_ts_residual_coding_disabled_flag is equal to 0, SDH of the TS block (with or without BDPCM) is disabled. Figure 9 shows an example table including conditions for allowing sign data hiding for transform skip mode and block differential pulse code modulation mode, according to some embodiments of the present disclosure. As shown in Figure 9, in a non-BDPCM mode (e.g., the variable BdpcmFlag is equal to 0), SDH may be enabled if the variable slice_ts_residual_coding_disabled_flag is equal to 1.
[0103]
[0131] 10 illustrates an example syntax, including a portion of the syntax for residual coding for the conditions shown in FIG. 9, according to some embodiments of the present disclosure. As shown in FIG. 10, changes from the previous VVC are shown in bold italics, and syntax to be removed is further indicated by strikethrough. For example, if BdpcmFlag is equal to 0, SDH is disabled (e.g., signHidden is equal to 0). Furthermore, the redundant condition check of the variable ph_dep_quant_enabled_flag can be eliminated from the syntax.
[0104]
[0132] In some embodiments, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, sign data hiding is disabled for all coded blocks (e.g., both TS and non-TS blocks). This is because when the variable slice_ts_residual_coding_disabled_flag is equal to 1, it is likely that the slice will be coded in a lossless mode, in which case SDH may not be suitable. Figure 11 shows example syntax, including a portion of the syntax for residual coding for disabling sign data hiding, according to some embodiments of the present disclosure. As shown in Figure 11, changes from the previous VVC are shown in bold italics, and syntax to be removed is further shown in strikethrough. For example, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, SDH is disabled.
[0105]
[0133] In some embodiments, a slice-level sign data hiding flag (e.g., variable slice_sign_data_hiding_enabled_flag) can be introduced to control sign data hiding for a slice. For example, if the variable slice_sign_data_hiding_enabled_flag is equal to 0, sign bit hiding is disabled for the current slice. If the variable slice_sign_data_hiding_enabled_flag is equal to 1, sign bit hiding is enabled for the current slice. In some embodiments, when the variable slice_sign_data_hiding_enabled_flag is not present, it is inferred to be equal to 0.
[0106]
[0134] In some embodiments, for a given slice, both sign data hiding and slice_ts_residual_coding_disabled_flag == 1 are disabled. More specifically, the following combinations are considered: a.slice_sign_data_hiding_enabled_flag is equal to 1, and The combination with b.slice_ts_residual_coding_disabled_flag equal to 1 is not allowed.
[0107]
[0135] In some embodiments, the variable slice_sign_data_hiding_enabled_flag is signaled if both of the following conditions are met: 1) the variable pic_sign_data_hiding_enabled_flag is equal to 1, i.e., the current picture allows SDH, and 2) the variable slice_ts_residual_coding_disabled_flag is equal to 0, i.e., the slice is likely coded in a lossy mode, for which SDH can be a useful tool.
[0108]
[0136] 12 shows an example syntax, including a portion of the slice header syntax, for slice-level sign data hiding flag control, according to some embodiments of the present disclosure. As shown in FIG. 10, changes from the previous VVC are shown in bold italics. For example, as shown in FIG. 10, a new variable slice_sign_data_hiding_enabled_flag has been added.
[0109]
[0137] In some embodiments, a slice level slice_ts_residual_coding_disabled_flag is signaled if slice_sign_data_hiding_enabled_flag is not equal to 1. This is one way of disallowing the combination of both slice_ts_residual_coding_disabled_flag and slice_sign_data_hiding_enabled_flag from being enabled on the same slice.
[0110]
[0138] 13 illustrates example syntax, including a portion of the residual coding syntax, for slice-level sign data hiding flag control, according to some embodiments of the present disclosure. As shown in FIG. 13, changes from the previous VVC are indicated in bold, and syntax to be removed is further indicated in strikethrough. For example, when the variable slice_sign_data_hiding_enabled_flag is equal to 0, SDH is disabled regardless of the values of the variables ph_dep_quant_enabled_flag or pic_sign_data_hiding_enabled_flag.
[0111]
[0139] In some embodiments, the variable slice_sign_data_hiding_enabled_flag may be signaled before the signaling of the variable slice_ts_residual_coding_disabled_flag, and the variable slice_ts_residual_coding_disabled_flag may be conditionally signaled if the variable slice_sign_data_hiding_enabled_flag is equal to 0. Figure 14 shows example syntax, including a portion of the slice header syntax, for a slice-level sign data hiding flag, in accordance with some embodiments of the present disclosure. As shown in Figure 14, changes from the previous VVC are indicated in bold. For example, slice_sign_data_hiding_enabled_flag may be set if pic_sign_data_hidnig_enabled_flag is equal to 1. As shown in Figure 14, the signaling of the variable slice_sign_data_hiding_enabled_flag at the slice level is determined by the variable pic_sign_data_hiding_enabled_flag at the picture level, and furthermore the signaling of the variable slice_ts_residual_coding_disabled_flag is determined by the variable slice_sign_data_hiding_enabled_flag.
[0112]
[0140] In some embodiments, the signaling of the slice_sign_data_hiding_enabled_flag variable and the signaling of the slice_ts_residual_coding_disabled_flag variable can be processed independently of each other. In other words, one signaling does not depend on the other signaling. In some embodiments, this may require the encoder to send the slice_sign_data_hiding_enabled_flag variable with the value 0 even if the encoder is already in lossless mode. This may result in syntax redundancy. Nevertheless, this may create the flexibility in the future to use regular residual coding and SDH together for TS and BDPCM blocks to improve the efficiency of lossy coding.
[0113]
[0141] In some embodiments, the slice-level variable slice_ts_residual_coding_disabled_flag can be replaced with a slice-level lossless variable, namely slice_lossless_flag. In some embodiments, a value of 1 for the variable slice_lossless_flag indicates that the current slice is losslessly coded and that all residual blocks of that slice use the residual_coding() syntax to parse residual samples. A value of 0 for the variable slice_lossless_flag indicates that the current slice is not losslessly coded. In some embodiments, when the variable slice_lossless_flag is not present, it is inferred to be equal to 0.
[0114]
[0142] Figure 15 shows example syntax, including a portion of the slice header syntax, for a slice-level lossless flag, according to some embodiments of the present disclosure. As shown in Figure 15, changes from the previous VVC are shown in bold italics, and syntax to be removed is further shown in strikethrough. For example, a variable slice_lossless_flag has been added and incorporated into the syntax. The variable slice_lossless_flag is signaled after the slice type signaling. When the variable slice_lossless_flag is equal to 1, some or all of the loop filters (e.g., adaptive loop filter, sample adaptive offset, deblocking filter, and luma mapping with chroma scaling) can be disabled for a lossless slice. For example, if the variable slice_lossless_flag is equal to 1, the variables slice_alf_enabled_flag, slice_sao_luma_flag, slice_deblocking_filter_override_flag, and slice_lmcs_enabled_flag are not signaled.
[0115]
[0143] 16 illustrates an example syntax including a portion of the residual coding syntax for a slice-level lossless flag, according to some embodiments of the present disclosure. As shown in FIG. 16, changes from the previous VVC are indicated in bold italics, and syntax to be removed is further indicated in strikethrough. For example, a variable slice_lossless_flag has been added and incorporated into the syntax. In some embodiments, as shown in FIG. 16, the variable slice_lossless_flag can be replaced with the variable slice_ts_residual_coding_disabled_flag when determining the condition for residual coding.
[0116]
[0144] In some embodiments, the values of the variables pic_sign_data_hiding_enabled_flag and slice_ts_residual_coding_disabled_flag cannot both be equal to 1. When pic_sign_data_hiding_enabled_flag is equal to 1, the codec is almost certainly operating in lossy mode, and the likelihood of using the variable slice_ts_residual_coding_disabled_flag equal to 0 is very high for lossy compression. Consequently, to reduce syntax redundancy, it can be inferred that when the variable pic_sign_data_hiding_enabled_flag is equal to 1, the variable slice_ts_residual_coding_disabled_flag is 0. The variable slice_ts_residual_coding_disabled_flag is signaled only if the variable pic_sign_data_hiding_enabled_flag is equal to 0. 17 illustrates an example syntax including a portion of the syntax for slice headings with reduced syntactic redundancy, according to some embodiments of the present disclosure. As shown in FIG. 17, changes from the previous VVC are indicated in bold italics. For example, when the variable pic_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is equal to 1.
[0117]
[0145] In some embodiments, sign data hiding and state-dependent quantization cannot be supported simultaneously for a given picture. As a result, the variables ph_dep_quant_enabled_flag and pic_sign_data_hiding_enabled_flag may not be equal to 1 at the same time. To avoid this combination, the variable slice_ts_residual_coding_disabled_flag is signaled when the variable ph_dep_quant_enabled_flag is equal to 1. Figure 18 shows example syntax, including a portion of the slice header syntax, for conditions on sign data hiding and state-dependent quantization, according to some embodiments of the present disclosure. As shown in Figure 18, changes from previous VVC are indicated in bold italics. For example, as shown in Figure 18, the variable slice_ts_residual_coding_disabled_flag is signaled when the variable ph_dep_quant_enabled_flag is equal to 1.
[0118]
[0146] In some embodiments, for a given slice, it is not possible to simultaneously support state-dependent quantization and slice_ts_residual_coding_disabled_flag == 1. To avoid this combination, if the variable ph_dep_quant_enabled_flag is equal to 1, the variable slice_ts_residual_coding_disabled_flag is not signaled.
[0119]
[0147] In some embodiments, slice_ts_residual_coding_disabled_flag is signaled if slice_sign_data_hiding_enabled_flag and ph_dep_quant_enabled_flag are both non-zero values, as described more specifically below. if (!slice_sign_data_hiding_enabled_flag && ! ph_dep_quant_enabled_flag) signal _residual_coding_disabled_flag
[0120]
[0148] Embodiments of the present disclosure further provide a method for performing video encoding. Figure 19 shows a flowchart of an example of a video encoding method using transform skip mode and sign data hiding according to some embodiments of the present disclosure. In some embodiments, the method 19000 shown in Figure 19 may be performed by the device 400 shown in Figure 4. In some embodiments, the method 19000 shown in Figure 19 may be performed according to the syntax shown in Figure 8. In some embodiments, the method 19000 shown in Figure 19 is performed according to the VVC standard.
[0121]
[0149] In step S19010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0122]
[0150] In step S19020, it is determined whether the video frame is coded according to the transform skip mode at the transform block level. For example, as shown in Figure 8, a variable transform_skip_flag may be used to determine whether the video frame is coded according to the transform skip mode at the transform block level.
[0123]
[0151] In step S19030, in response to determining that the video frame is coded according to transform skip mode at the transform block level, sign data hiding is turned off for residual coding. For example, as shown in FIG. 8, when the variable transform_skip_flag is equal to 0, the variable signHidden is set to 0. As a result, sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. For example, as shown in FIG. 8, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden is independent of the variable ph_dep_quant_enabled_flag. In some embodiments, turning off sign data hiding for both TS blocks using non-BDPCM and TS blocks using BDPCM can improve compression performance efficiency. For example, the error propagation shown in FIG. 6 may be eliminated by turning off sign data hiding.
[0124]
[0152] Figure 20 shows a flowchart of an example of a video encoding method using block differential pulse code modulation mode and sign data hiding, according to some embodiments of the present disclosure. In some embodiments, the method 20000 shown in Figure 20 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 20000 shown in Figure 20 may be performed according to the syntax shown in Figure 10. In some embodiments, the method 20000 shown in Figure 20 is performed in accordance with the VVC standard.
[0125]
[0153] In step S20010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0126]
[0154] In step S20020, it is determined whether the video frame is encoded according to a block differential pulse code modulation mode. In some embodiments, it is determined whether the video frame is encoded according to a block differential pulse code modulation mode at the block level. For example, as shown in Figure 10, a variable BdpcmFlag can be used to determine whether the video frame is encoded according to a block differential pulse code modulation mode at the block level.
[0127]
[0155] In step S20030, in response to determining that the video frame is coded according to a block differential pulse code modulation mode, sign data hiding is turned off for residual coding. For example, as shown in FIG. 10 , when the variable BdpcmFlag is equal to 0, the variable signHidden is set to 0. As a result, sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. For example, as shown in FIG. 10 , the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden is independent of the variable ph_dep_quant_enabled_flag. In some embodiments, turning off sign data hiding for TS blocks that use BDPCM can improve compression performance efficiency. For example, the error propagation shown in FIG. 6 may be eliminated by turning off sign data hiding.
[0128]
[0156] Figure 21 shows a flowchart of an example of a video encoding method using transform skip residual coding and sign data hiding according to some embodiments of this disclosure. In some embodiments, the method 21000 shown in Figure 21 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 21000 shown in Figure 21 may be performed according to the syntax shown in Figure 11. In some embodiments, the method 21000 shown in Figure 21 is performed according to the VVC standard.
[0129]
[0157] In step S21010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0130]
[0158] In step S21020, it is determined whether the video frame is coded according to the transform skip residual coding mode at the slice level. For example, as shown in Figure 11, a variable slice_ts_residual_coding_disabled_flag may be used to determine whether the video frame is coded according to the transform skip residual coding mode at the slice level.
[0131]
[0159] In step S21030, in response to determining that the video frame is not coded according to a transform skip residual coding mode at the slice level, sign data hiding is turned off for residual coding. For example, as shown in FIG. 11 , it is determined that the video frame is not coded according to a transform skip residual coding mode at the slice level when the variable slice_ts_residual_coding_disabled_flag is equal to 1. As a result, the variable signHidden is set to 0, and sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. For example, as shown in FIG. 11 , the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden is independent of the variable ph_dep_quant_enabled_flag. In some embodiments, sign data hiding is a tool of lossy coding. Regular residual coding (e.g., slice_ts_residual_coding_disabled_flag equals 1) can achieve higher compression gain than TS residual coding. Therefore, when the variable slice_ts_residual_coding_disabled_flag equals 1, the compression being performed can be lossless. As a result, sign data hiding can be turned off.
[0132]
[0160] Figure 22 shows a flowchart of an example of a video encoding method using transform skip residual coding and picture-level sign data hiding according to some embodiments of this disclosure. In some embodiments, the method 22000 shown in Figure 22 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 22000 shown in Figure 22 may be performed according to the syntax shown in Figure 12 or the syntax shown in Figure 13. In some embodiments, the method 22000 shown in Figure 22 is performed according to the VVC standard.
[0133]
[0161] In step S22010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0134]
[0162] In step S22020, it is determined whether sign data hiding is enabled at the picture level of the video frame and whether transform skip residual coding is disabled at the slice level of the video frame. For example, as shown in Figure 12, a variable pic_sign_data_hiding_enabled_flag may be used to determine whether sign data hiding is enabled at the picture level of the video frame, and a variable slice_ts_residual_coding_disabled_flag may be used to determine whether the video frame is coded according to a transform skip residual coding mode at the slice level of the video frame.
[0135]
[0163] In step S22030, in response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual coding is enabled at the slice level of the video frame, sign data hiding is turned on at the slice level of the video frame. For example, as shown in FIG. 12, when the variable pic_sign_data_hiding_enabled_flag is equal to 1, it is determined that sign data hiding is enabled at the picture level of the video frame. Furthermore, when the variable slice_ts_residual_coding_disabled_flag is equal to 1, it is determined that the video frame is not coded according to the transform skip residual coding mode at the slice level. As a result, the variable slice_sign_data_hiding_enabled_flag is set to 1, and sign data hiding is turned on at the slice level. In some embodiments, sign data hiding is a tool of lossy coding. Regular residual coding (e.g., slice_ts_residual_coding_disabled_flag is equal to 1) can achieve higher compression gain than transform skip residual coding. Therefore, transform skip residual coding can be lossy, and as a result, sign data hiding can be turned on.
[0136]
[0164] In some embodiments, method 22000 further includes steps S22040 and S22050. In step S22040, it is determined whether sign data hiding is turned off at the slice level of the video frame. For example, as shown in Figure 13, a variable slice_sign_data_hiding_enabled_flag may be checked to determine whether sign data hiding is turned off at the slice level of the video frame.
[0137]
[0165] In step S22050, in response to determining that sign data hiding is turned off at the slice level of the video frame, sign data hiding is turned off for residual coding. For example, as shown in FIG. 13, when the variable slice_sign_data_hiding_enabled_flag is equal to 0, the variable signHidden is set to 0 and sign data hiding is turned off for residual coding. In some embodiments, turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. For example, as shown in FIG. 13, the variable ph_dep_quant_enabled_flag is excluded from the condition of the variable signHidden. In other words, the value of signHidden is independent of the variable ph_dep_quant_enabled_flag.
[0138]
[0166] Figure 23 shows a flowchart of an example of a video coding method using sign data hiding at the picture level, sign data hiding at the slice level, and transform skip residual coding at the slice level, according to some embodiments of this disclosure. In some embodiments, the method 23000 shown in Figure 23 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 23000 shown in Figure 23 may be performed according to the syntax shown in Figure 14. In some embodiments, the method 23000 shown in Figure 23 is performed in accordance with the VVC standard.
[0139]
[0167] In step S23010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0140]
[0168] In step S23020, it is determined whether sign data hiding is enabled at the picture level of the video frame. For example, as shown in Figure 14, a variable pic_sign_data_hiding_enabled_flag may be used to determine whether sign data hiding is enabled at the picture level of the video frame.
[0141]
[0169] In step S23030, in response to determining that sign data hiding is enabled at the picture level of the video frame, sign data hiding is turned on at the slice level of the video frame. For example, as shown in Figure 14, it is determined that sign data hiding is enabled at the picture level of the video frame when the variable pic_sign_data_hiding_enabled_flag is equal to 1. As a result, the variable slice_sign_data_hiding_enabled_flag is set to 1, and sign data hiding is turned on at the slice level.
[0142]
[0170] In step S23040, it is determined whether sign data hiding is turned off at the slice level of the video frame. For example, as shown in Figure 14, the variable slice_sign_data_hiding_enabled_flag is checked to determine whether sign data hiding is turned off at the slice level of the video frame.
[0143]
[0171] In step S23050, in response to determining that sign data hiding is turned off at the slice level of the video frame, transform skip residual coding is turned off at the slice level of the video frame. For example, as shown in FIG. 14, if the variable slice_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is set to 1, and transform skip residual coding is turned off at the slice level of the video frame. In some embodiments, sign data hiding is a lossy coding tool. Regular residual coding (e.g., slice_ts_residual_coding_disabled_flag is equal to 1) can achieve higher compression gain than transform skip residual coding. Thus, transform skip residual coding can be lossy compression. When sign data hiding is turned off, transform skip residual coding at the slice level can also be turned off.
[0144]
[0172] Figure 24 shows a flowchart of an example of a video encoding method using a lossless encoding mode and sign data hiding, according to some embodiments of the present disclosure. In some embodiments, the method 24000 shown in Figure 24 may be performed by the device 400 shown in Figure 4. In some embodiments, the method 24000 shown in Figure 24 may be performed according to the syntax shown in Figure 15 or the syntax shown in Figure 16. In some embodiments, the method 24000 shown in Figure 24 is performed according to the VVC standard.
[0145]
[0173] In step S24010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0146]
[0174] In step S24020, it is determined whether the video frame is encoded in a lossless mode. In some embodiments, it is determined whether the video frame is encoded in a lossless mode at the slice level. For example, as shown in Figure 15, a variable slice_lossless_flag can be used to determine whether the video frame is encoded in a lossless mode at the slice level.
[0147]
[0175] In step S24030, in response to determining that the video frame is encoded in a lossless mode at the slice level, one or more loop filters are turned off at the slice level. For example, as shown in Figure 15, when the variable slice_lossless_flag is equal to 1, the variables slice_alf_enabled_flag, slice_sao_luma_flag, slice_deblocking_filter_override_flag, and slice_lmcs_enabled_flag are not signaled. In some embodiments, turning off one or more loop filters can make video encoding more efficient.
[0148]
[0176] Figure 25 shows a flowchart of an example of a video encoding method using sign data hiding at the picture level and transform skip residual coding at the slice level, according to some embodiments of this disclosure. In some embodiments, the method 25000 shown in Figure 25 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 25000 shown in Figure 25 may be performed according to the syntax shown in Figure 17. In some embodiments, the method 25000 shown in Figure 25 is performed in accordance with the VVC standard.
[0149]
[0177] In step S25010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0150]
[0178] In step S25020, it is determined whether sign data hiding is turned off at the picture level of the video frame. For example, as shown in Figure 17, the variable pic_sign_data_hiding_enabled_flag can be used to determine whether sign data hiding is turned off at the picture level of the video frame.
[0151]
[0179] In step S25030, in response to determining that sign data hiding is turned off at the picture level of the video frame, transform skip residual coding is turned off at the slice level of the video frame. For example, as shown in FIG. 17, the variable pic_sign_data_hiding_enabled_flag is examined. If the variable pic_sign_data_hiding_enabled_flag is equal to 0, the variable slice_ts_residual_coding_disabled_flag is signaled. As a result, transform skip residual coding is turned off at the slice level of the video frame. In some embodiments, the combination of normal residual coding (e.g., slice_ts_residual_coding_disabled_flag == 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag == 1) is not a useful configuration. Therefore, disallowing this combination can reduce syntax redundancy.
[0152]
[0180] Figure 26 shows a flowchart of an example video coding method using state-dependent quantization and transform skip residual coding at the slice level, according to some embodiments of this disclosure. In some embodiments, the method 26000 shown in Figure 26 may be performed by the apparatus 400 shown in Figure 4. In some embodiments, the method 26000 shown in Figure 26 may be performed according to the syntax shown in Figure 18. In some embodiments, the method 26000 shown in Figure 26 is performed according to the VVC standard.
[0153]
[0181] In step S26010, a video frame is received for encoding. In some embodiments, the video frame is in a bitstream. In some embodiments, the video frame is received for residual encoding.
[0154]
[0182] In step S26020, it is determined whether state-dependent quantization is enabled for the video frame. For example, as shown in Figure 18, a variable ph_dep_quant_enabled_flag may be used to determine whether state-dependent quantization is enabled for the video frame.
[0155]
[0183] In step S26030, in response to determining that state-dependent quantization is enabled for the video frame, transform skip residual coding is turned off at the slice level of the video frame. For example, as shown in FIG. 18, the variable ph_dep_quant_enabled_flag is examined. If the variable ph_dep_quant_enabled_flag is equal to 1, the variable slice_ts_residual_coding_disabled_flag is signaled. As a result, transform skip residual coding is turned off at the slice level of the video frame. In some embodiments, the combination of normal residual coding (e.g., slice_ts_residual_coding_disabled_flag == 1) and sign data hiding (e.g., pic_sign_data_hiding_enabled_flag == 1) is not a useful configuration. Therefore, disallowing this combination can reduce syntax redundancy.
[0156]
[0184] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by a device (such as the disclosed encoders and decoders) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs, and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0157]
[0185] It should be noted that, as used herein, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent in meaning and are open-ended in that the item or items following any one of these words are not intended to be an exhaustive list of such items or items, nor are they intended to be limited only to the item or items listed.
[0158]
[0186] As used herein, unless otherwise specified, the term "or" includes all possible combinations unless impracticable. For example, if it is said that a database may include A or B, the database may include A, B, A and B, unless otherwise specified or impracticable. As a second example, if it is said that a database may include A, B, or C, the database may include A, B, C, A and B, A and C, B and C, or A, B and C, unless otherwise specified or impracticable.
[0159]
[0187] It will be understood that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and that each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0160]
[0188] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications can be made to the described embodiments. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims. It is also intended that the order of steps depicted in the figures is for illustrative purposes only, and is not intended to be limited to any particular order of steps. As such, one skilled in the art will recognize that these steps can be performed in different orders while implementing the same method.
[0161]
[0189] The following clauses may be used to further describe the embodiments. 1. A video encoding method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip mode at the transform block level; in response to determining that the video frame is coded according to transform skip mode, turning off sign data hiding for residual coding; A video encoding method comprising: 2. The video encoding method of clause 1, wherein turning off sign data hiding for residual encoding is independent of whether state-dependent quantization is enabled for the video frame. 3. A video encoding method as recited in clause 1, wherein the video frames are in a bitstream. 4. A video coding method according to clause 1, performed in accordance with the versatile video coding (VVC) standard. 5. A video encoding method comprising: receiving a video frame for residual coding; determining whether the video frame is encoded according to a block differential pulse code modulation mode; in response to determining that the video frame is encoded according to a block differential pulse code modulation mode, turning off sign data hiding for residual encoding; A video encoding method comprising: 6. Determining whether the video frame is coded according to a block differential pulse code modulation mode 6. The video encoding method of clause 5, further comprising determining whether the video frame is encoded according to a block differential pulse code modulation mode at a block level. 7. The video encoding method of clause 5, wherein turning off sign data hiding for residual encoding is independent of whether state-dependent quantization is enabled for the video frame. 8. A video encoding method as recited in clause 5, wherein the video frames are in a bitstream. 9. A video coding method according to clause 5, performed in accordance with the VVC (versatile video coding) standard. 10. A video encoding method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip residual coding mode at a slice level; in response to determining that the video frame is not coded according to a transform skip residual coding mode at the slice level, turning off sign data hiding for residual coding; A video encoding method comprising: 11. The video encoding method of clause 10, wherein turning off sign data hiding for residual encoding is independent of whether state-dependent quantization is enabled for the video frame. 12. A video encoding method as recited in clause 10, wherein the video frames are in a bitstream. 13. A video coding method according to clause 10, performed in accordance with the versatile video coding (VVC) standard. 14. A video encoding method comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level of the video frame and whether transform skip residual coding is disabled at a slice level of the video frame; In response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual coding is enabled at the slice level of the video frame, turning on sign data hiding at the slice level of the video frame; A video encoding method comprising: 15. Determining whether sign data hiding is turned off at the slice level of the video frame; in response to determining that sign data hiding is turned off at the slice level of the video frame, turning off sign data hiding for residual coding; 15. The video encoding method of clause 14, further comprising: 16. A video encoding method according to clause 14, wherein the video frames are in a bitstream. 17. A video coding method according to clause 14, performed in accordance with the versatile video coding (VVC) standard. 18. A video encoding method comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level of the video frame; responsive to determining that sign data hiding is enabled at the picture level of the video frame, turning on sign data hiding at the slice level of the video frame; determining whether sign data hiding is turned off at the slice level of the video frame; in response to determining that sign data hiding is turned off at the slice level of the video frame, turning off transform skip residual coding at the slice level of the video frame; A video encoding method comprising: 19. A video encoding method according to clause 18, wherein the video frames are in a bitstream. 20. A video coding method according to clause 18, performed in accordance with the versatile video coding (VVC) standard. 21. A video encoding method comprising: receiving a video frame for residual coding; determining whether the video frame is encoded in a lossless mode at a slice level; in response to determining that the video frame is encoded in a lossless mode at the slice level, turning off one or more loop filters at the slice level; A video encoding method comprising: 22. A video encoding method according to clause 21, wherein the video frames are in a bitstream. 23. A video coding method according to clause 21, performed in accordance with the versatile video coding (VVC) standard. 24. A video encoding method comprising: receiving a video frame for residual coding; determining whether sign data hiding is turned off at the picture level of the video frame; in response to determining that sign data hiding is turned off at the picture level of the video frame, turning off transform skip residual coding at the slice level of the video frame; A video encoding method comprising: 25. A video encoding method as recited in clause 24, wherein the video frames are in a bitstream. 26. A video coding method according to clause 24, performed in accordance with the versatile video coding (VVC) standard. 27. A video encoding method comprising: receiving a video frame for residual coding; determining whether state-dependent quantization is enabled for the video frame; in response to determining that state-dependent quantization is enabled for the video frame, turning off transform skip residual coding at a slice level for the video frame; A video encoding method comprising: 28. A video encoding method as recited in clause 27, wherein the video frames are in a bitstream. 29. A video coding method according to clause 27, performed in accordance with the versatile video coding (VVC) standard. 30. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip mode at the transform block level; and turning off sign data hiding for residual coding in response to determining that the video frame is coded according to transform skip mode; A system configured to cause 31. The system of clause 30, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 32. The system of clause 30, wherein the video frames are in a bitstream. 33. The system of clause 30, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 34. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether the video frame is encoded according to a block differential pulse code modulation mode; and turning off sign data hiding for residual encoding in response to determining that the video frame is encoded according to a block differential pulse code modulation mode; A system configured to cause 35. A processor executes a set of instructions to provide a system 35. The system of clause 34, further configured to perform determining whether the video frame is coded according to a block differential pulse code modulation mode at a block level. 36. The system of clause 34, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 37. The system of clause 34, wherein the video frames are in a bitstream. 38. The system of clause 34, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 39. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip residual coding mode at the slice level; and turning off sign data hiding for residual coding in response to determining that the video frame is not coded according to a transform skip residual coding mode at the slice level; A system configured to cause 40. The system of clause 39, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 41. The system of clause 39, wherein the video frames are in a bitstream. 42. The system of clause 39, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 43. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level of the video frame and whether transform skip residual coding is disabled at a slice level of the video frame; and turning on sign data hiding at a slice level of the video frame in response to determining that sign data hiding is enabled at a picture level of the video frame and that transform skip residual coding is enabled at a slice level of the video frame; A system configured to cause 44. A processor executes a set of instructions to provide a system determining whether sign data hiding is turned off at the slice level of the video frame; and turning off sign data hiding for residual coding in response to determining that sign data hiding is turned off at the slice level of the video frame. 44. The system of clause 43, further configured to: 45. The system of clause 43, wherein the video frames are in a bitstream. 46. The system of clause 43, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 47. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level for a video frame; responsive to determining that sign data hiding is enabled at the picture level of the video frame, turning on sign data hiding at the slice level of the video frame; determining whether sign data hiding is turned off at the slice level of the video frame; and and turning off transform skip residual coding at the slice level of the video frame in response to determining that sign data hiding is turned off at the slice level of the video frame. A system configured to cause 48. The system of clause 47, wherein the video frames are in a bitstream. 49. The system of clause 47, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 50. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether a video frame is encoded in a lossless mode at a slice level; and turning off one or more loop filters at the slice level in response to determining that the video frame is encoded in a lossless mode at the slice level A system configured to cause 51. The system of clause 50, wherein the video frames are in a bitstream. 52. The system of clause 50, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 53. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether sign data hiding is turned off at the picture level of the video frame; and turning off transform skip residual coding at a slice level of the video frame in response to determining that sign data hiding is turned off at a picture level of the video frame. A system configured to cause 54. The system of clause 53, wherein the video frames are in a bitstream. 55. The system of clause 54, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 56. A system for performing video data processing, comprising: a memory for storing a set of instructions; a processor that executes a set of instructions to provide the system with receiving a video frame for residual coding; determining whether state-dependent quantization is enabled for the video frame; and turning off transform skip residual coding at the slice level of the video frame in response to determining that state-dependent quantization is enabled for the video frame. A system configured to cause 57. The system of clause 56, wherein the video frames are in a bitstream. 58. The system of clause 56, wherein the residual coding is performed in accordance with the versatile video coding (VVC) standard. 59. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip mode at the transform block level; in response to determining that the video frame is coded according to transform skip mode, turning off sign data hiding for residual coding; 1. A non-transitory computer-readable medium comprising: 60. The non-transitory computer-readable medium of clause 59, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 61. The non-transitory computer-readable medium of clause 59, wherein the video frames are in a bitstream. 62. The non-transitory computer-readable medium of clause 59, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 63. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is encoded according to a block differential pulse code modulation mode; in response to determining that the video frame is encoded according to a block differential pulse code modulation mode, turning off sign data hiding for residual encoding; 1. A non-transitory computer-readable medium comprising: 64. A set of instructions is used in a computer system to and further determining whether the video frame is coded according to a block differential pulse code modulation mode at the block level. 64. A non-transitory computer-readable medium as set forth in clause 63, executable by at least one processor of a computer system. 65. The non-transitory computer-readable medium of clause 63, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 66. The non-transitory computer-readable medium of clause 63, wherein the video frames are in a bitstream. 67. The non-transitory computer-readable medium of clause 63, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 68. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is coded according to a transform skip residual coding mode at a slice level; in response to determining that the video frame is not coded according to a transform skip residual coding mode at the slice level, turning off sign data hiding for residual coding; 1. A non-transitory computer-readable medium comprising: 69. The non-transitory computer-readable medium of clause 68, wherein turning off sign data hiding for residual coding is independent of whether state-dependent quantization is enabled for the video frame. 70. The non-transitory computer-readable medium of clause 68, wherein the video frames are in a bitstream. 71. The non-transitory computer-readable medium of clause 68, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 72. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level of the video frame and whether transform skip residual coding is disabled at a slice level of the video frame; In response to determining that sign data hiding is enabled at the picture level of the video frame and that transform skip residual coding is enabled at the slice level of the video frame, turning on sign data hiding at the slice level of the video frame; 1. A non-transitory computer-readable medium comprising: 73. A set of instructions is used in a computer system to determining whether sign data hiding is turned off at the slice level of the video frame; and and further causing, in response to determining that sign data hiding is turned off at the slice level of the video frame, turning off sign data hiding for residual coding. 73. A non-transitory computer-readable medium as set forth in clause 72, executable by at least one processor of a computer system. 74. The non-transitory computer-readable medium of clause 72, wherein the video frames are in a bitstream. 75. The non-transitory computer-readable medium of clause 72, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 76. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether sign data hiding is enabled at a picture level of the video frame; responsive to determining that sign data hiding is enabled at the picture level of the video frame, turning on sign data hiding at the slice level of the video frame; determining whether sign data hiding is turned off at the slice level of the video frame; in response to determining that sign data hiding is turned off at the slice level of the video frame, turning off transform skip residual coding at the slice level of the video frame; 1. A non-transitory computer-readable medium comprising: 77. The non-transitory computer-readable medium of clause 76, wherein the video frames are in a bitstream. 78. The non-transitory computer-readable medium of clause 76, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 79. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether the video frame is encoded in a lossless mode at a slice level; in response to determining that the video frame is encoded in a lossless mode at the slice level, turning off one or more loop filters at the slice level; 1. A non-transitory computer-readable medium comprising: 80. The non-transitory computer-readable medium of clause 79, wherein the video frames are in a bitstream. 81. The non-transitory computer-readable medium of clause 79, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 82. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether sign data hiding is turned off at the picture level of the video frame; in response to determining that sign data hiding is turned off at the picture level of the video frame, turning off transform skip residual coding at the slice level of the video frame; 1. A non-transitory computer-readable medium comprising: 83. The non-transitory computer-readable medium of clause 82, wherein the video frames are in a bitstream. 84. The non-transitory computer-readable medium of clause 82, wherein the method is performed in accordance with the versatile video coding (VVC) standard. 85. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a video frame for residual coding; determining whether state-dependent quantization is enabled for the video frame; in response to determining that state-dependent quantization is enabled for the video frame, turning off transform skip residual coding at a slice level for the video frame; 1. A non-transitory computer-readable medium comprising: 86. The non-transitory computer-readable medium of clause 85, wherein the video frames are in a bitstream. 87. The non-transitory computer-readable medium of clause 85, wherein the method is performed in accordance with the versatile video coding (VVC) standard.
[0162]
[0190] The drawings and specification disclose illustrative embodiments. However, these embodiments are susceptible to many variations and modifications. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A video encoding method, comprising: receiving a video frame for residual coding; determining a value of a first flag indicating whether sign data hiding is turned off at a slice level for the video frame; determining whether to signal a second flag indicating whether transform skip residual coding is turned off at the slice level of the video frame based on the value of the first flag; Including, 10. A video encoding method, wherein if the first flag has a value indicating that the sign data hiding is turned on at the slice level of the video frame, the second flag is not signaled.
2. The video encoding method of claim 1 , wherein the video frames are in a bitstream.
3. 10. The video coding method of claim 1, performed in accordance with the versatile video coding (VVC) standard.
4. The video encoding method of claim 1 , wherein the first flag comprises a slice-level flag slice_sign_data_hiding_enabled_flag.
5. The video encoding method of claim 1 , wherein the second flag comprises a slice-level flag slice_ts_residual_coding_disabled_flag.
6. determining whether to signal the second flag indicating whether the transform skip residual coding is turned off at the slice level of the video frame based on the value of the first flag; 2. The video encoding method of claim 1, comprising signaling the second flag in response to the value of the first flag indicating that the sign data hiding is turned off at the slice level of the video frame.
7. determining whether to signal the second flag indicating whether the transform skip residual coding is turned off at the slice level of the video frame based on the value of the first flag; 2. The video encoding method of claim 1, comprising signaling the second flag in response to the value of the first flag being equal to zero.
8. 2. The video encoding method of claim 1, wherein if the value of the first flag is equal to 1, the second flag is not signaled.
9. 1. A video decoding method, comprising: receiving a bitstream coded based on residual coding, the bitstream including a first flag indicating whether sign data hiding is turned off at a slice level; decoding the first flag; determining whether to decode a second flag indicating whether transform skip residual coding is turned off at the slice level based on a value of the first flag; Including, 2. The video decoding method of claim 1, wherein the second flag is not decoded if the first flag has a value indicating that the sign data hiding is turned on at the slice level of a video frame.
10. The video decoding method of claim 9 , wherein the first flag comprises a slice-level flag slice_sign_data_hiding_enabled_flag.
11. The video decoding method of claim 9 , wherein the second flag comprises a slice-level flag slice_ts_residual_coding_disabled_flag.
12. determining whether to decode the second flag indicating whether the transform skip residual coding is turned off at the slice level based on the value of the first flag; 10. The video decoding method of claim 9, comprising: decoding the second flag in response to the value of the first flag indicating that the sign data hiding is turned off at the slice level.
13. determining whether to decode the second flag indicating whether the transform skip residual coding is turned off at the slice level based on the value of the first flag; 10. The video decoding method of claim 9, comprising decoding the second flag in response to the value of the first flag being equal to zero.
14. 10. The video decoding method of claim 9, further comprising: inferring that the transform skip residual coding is turned on at the slice level in response to the value of the first flag being equal to one.
15. 10. The video decoding method of claim 9, further comprising inferring a value of the second flag to be equal to zero in response to the value of the first flag being equal to one.
16. 10. The video decoding method of claim 9, performed in accordance with the versatile video coding (VVC) standard.
17. 1. A method of storing a video bitstream, said method comprising: determining a value of a first flag indicating whether sign data hiding is turned off at the slice level; generating a bitstream including encoded information based on the value of the first flag; storing the bitstream in a non-transitory computer-readable storage medium; Including, When the first flag has a first value, the bitstream further includes a second flag indicating whether transform skip residual coding is turned off at the slice level; A method, wherein when the first flag has a second value, the bitstream does not include the second flag.
18. The method of claim 17 , wherein the first flag comprises a slice-level flag slice_sign_data_hiding_enabled_flag.
19. The method of claim 17 , wherein the second flag comprises a slice-level flag slice_ts_residual_coding_disabled_flag.
20. 18. The method of claim 17, wherein the first value of the first flag indicates that the sign data hiding is turned off at the slice level.
21. 18. The method of claim 17, wherein the first value of the first flag is equal to 0.
22. 18. The method of claim 17, wherein the second value of the first flag indicates that the sign data hiding is turned on at the slice level.
23. 18. The method of claim 17, wherein the second value of the first flag is equal to one.
24. 18. The method of claim 17, wherein the bitstream is coded according to the versatile video coding (VVC) standard.
Citation Information
Patent Citations
Image decoding device and image encoding device
JP2021150703A
Image processing device and method
WO2018173798A1
Method and apparatus for decoding imaging related to sign data hiding
WO2021172912A1
Image decoding method for residual coding and device for same
WO2021172914A1
Method and apparatus for video encoding and decoding
WO2021180710A1