Method for decoding an image, method for encoding an image into a bit stream, decoding apparatus, encoding apparatus, computer program

By determining the use of sign bit hiding for sub-blocks of residual coefficients based on specific flags and parameters, the method improves video encoding efficiency and flexibility, particularly in reversible coding scenarios.

JP7695422B2Active Publication Date: 2025-06-18CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024020597
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2024-02-14
Publication Date
2025-06-18
Estimated Expiration
2040-11-23

Smart Images

  • Figure 0007695422000004
    Figure 0007695422000004
  • Figure 0007695422000005
    Figure 0007695422000005
  • Figure 0007695422000006
    Figure 0007695422000006
Patent Text Reader

Abstract

To provide a method, an apparatus, and a system for encoding and decoding a block of a video sample.SOLUTION: A method includes: determining whether sign bit hiding is used for a subblock, wherein the determination is based on a value of a conversion skip flag determined for the subblock and a value of a sign bit hiding flag related to the subblock; when the sign bit hiding is not used, decoding sign bits in a number equal to the number of significant coefficients in the subblock; and decoding the subblock by reconstructing a residual coefficient of the subblock using the decoded sign bit.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of 35 U.S.C § 119 based on the filing date of Australian Patent Application No. 2020201753, filed on March 10, 2020, and is hereby incorporated by reference herein in its entirety as if fully set forth herein.

[0002] The present invention generally relates to digital video signal processing, and more particularly, to methods, apparatuses, and systems for encoding and decoding blocks of video samples. The present invention also relates to a computer program product including a computer - readable medium having recorded thereon a computer program for encoding and decoding blocks of video samples.

Background Art

[0003] There are currently many applications for video encoding, including applications for the transmission and storage of video data. Many video encoding standards have also been developed, and other standards are currently under development. Recent developments in video encoding standardization have led to the formation of a group called the "Joint Video Experts Team" (JVET). The Joint Video Experts Team (JVET) includes members of Study Group 16, Question 6 (SG16 / Q6) of the Telecommunication Standardization Sector (ITU - T) of the International Telecommunication Union (ITU), also known as the "Video Coding Experts Group" (VCEG), and members of Sub - committee 29, Working Group 11 (ISO / IEC JTC1 / SC29 / WG11) of the Joint Technical Committee 1 of the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), also known as the "Moving Picture Experts group" (MPEG).

[0004] The Joint Video Experts Team (JVET) analyzed the responses at its 10th meeting held in San Diego, USA, and issued a Call for Proposals (CfP). The submitted responses demonstrated video compression capabilities significantly exceeding those of the current state-of-the-art video compression standard, namely "High Efficiency Video Coding" (HEVC). Based on this outperformance, it was decided to initiate a project to develop a new video compression standard named "versatile video coding" (VVC). VVC is expected to continuously address the increasing demand for higher compression performance, especially as video formats increase in capabilities (e.g., at higher resolutions and higher frame rates) and the market demand for service delivery over WANs with relatively high bandwidth costs increases. At the same time, VVC must be implementable with modern silicon processes and provide an acceptable trade-off between the achieved performance and the implementation costs (e.g., with respect to silicon area, CPU processor load, memory usage, and bandwidth).

[0005] Video data comprises a sequence of frames of image data, each frame including one or more color channels. Generally, one primary colour channel and two secondary colour channels are required. The primary colour channel is generally referred to as the "luma" channel and the secondary colour channels are generally referred to as the "chroma" channels. Video data is typically displayed in the RGB (Red - Green - Blue) colour space, but this colour space has a high degree of correlation between each of the three respective elements. The video data representation seen by an encoder or decoder often uses a colour space such as YCbCr. YCbCr concentrates the luminance mapped to "luma" according to a transfer function in the Y (primary) channel and concentrates the chroma in the Cb and Cr (secondary) channels. Further, the Cb and Cr channels may be spatially sampled (subsampled) at a lower rate, for example, half horizontally and half vertically, compared to the luma channel, known as the "4:2:0 chroma format". The 4:2:0 chroma format is used in Internet video streaming, broadcast television, Blu-Ray TMIt is commonly used in "consumer" applications such as saving to a disk. Subsampling the Cb and Cr channels at half the rate horizontally and not subsampling vertically is known as the "4:2:2 chroma format". The 4:2:2 chroma format is typically used in professional applications, including the capture of video for movie production and the like. The higher sampling rate of the 4:2:2 chroma format makes the resulting video more flexible for editing operations such as color grading. Before distribution to consumers, 4:2:2 chroma format material is often converted to the 4:2:0 chroma format and then encoded for distribution to consumers. In addition to the chroma format, video is also characterized by resolution and frame rate. Example resolutions are ultra-high definition (UHD) at a resolution of 3840x2160, or "8K" at a resolution of 7680x4320, and example frame rates are 60 or 120 Hz. The luma sample rate may range from about 500 megasamples / second to several gigasamples / second. In the case of the 4:2:0 chroma format, the sample rate of each chroma channel is one-fourth of the luma sample rate, and in the case of the 4:2:2 chroma format, the sample rate of each chroma channel is half of the luma sample rate.

[0006] The VVC standard is a "block-based" codec, and a frame is first divided into a square array of regions known as "Coding Tree Units" (CTUs). A CTU generally occupies a relatively large area, such as 128×128 luma samples. However, the CTUs at the right end and bottom end of each frame may have a smaller area. Each CTU is associated with a "coding tree" for the luma channel and an additional coding tree for the chroma channel. A coding tree is defined to decompose the area of the CTU into a series of blocks also called "Coding Blocks" (CBs). It is also possible for a single coding tree to specify blocks for both the luma and chroma channels. In that case, the set of collocated coding blocks is called a "Coding Unit" (CU), that is, each CU has coding blocks for each color channel. The CBs are processed for encoding or decoding in a specific order. As a result of using the 4:2:0 chroma format, a CTU having a luma coding tree for a 128×128 luma sample area has a corresponding chroma coding tree for a 64×64 chroma sample area arranged together with the 128×128 luma sample area. When a single coding tree is used for both the luma and chroma channels, the set of collocated blocks for a given area is generally called a "unit", such as the above-mentioned CU, as well as a "Prediction Unit" (PU) and a "Transformation Unit" (TU). When separate coding trees are used for a given area, the above-mentioned CBs, as well as "Prediction Blocks" (PBs) and "Transformation Blocks" (TBs) are used.

[0007] Despite the above distinction between "units" and "blocks", the term "block" may be used as a general term for an area or region of a frame to which operations apply to all color channels.

[0008] For each CU, a prediction unit (PU) of the content (sample values) of the corresponding region of the frame data is generated ("prediction unit"). When the PU is generated from sample values within a frame that was previously signaled, the prediction is referred to as inter prediction. When the PU is generated from previous samples within the same frame, the prediction is called intra prediction. Further, an expression of the difference between the prediction and the content of the region seen in the input to the encoder (or "residual" in the spatial domain) is formed. The difference for each color channel is transformed as a block of residual coefficients, encoded, and can form one or more transform units (TUs) for a given CU. The residual coefficients can be transformed by a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or other transform to produce a final block of transform coefficients that substantially decorrelates the residual samples. A substantial coding gain can be achieved by quantizing the transform coefficients. The quantized transform coefficients are then traversed in an order such as a back diagonal scan, and each coefficient is encoded by an entropy encoder. Entropy coding consists of representing each coefficient by a syntax element, and each of the syntax elements is binarized. The binarized syntax elements are then further encoded by a context adaptive binary arithmetic coder (CABAC) or passed to the bitstream ("bypass coding").

[0009] In some classes of video content, such as screen content, it may be advantageous to avoid performing a transform. If the transform is to be avoided, the residual coefficients are quantized, traversed, and encoded. Since the statistics of the residual coefficients are not the same as those of the transform coefficients, it is generally advantageous for the residual coefficients encoded using a different process than the encoding process for the transform coefficients. Typical methods used to encode the residual coefficients include the "regular residual coding" (RRC) process and the "transform skip residual coding" (TSRC) process, and a particular one of the processes is selected for a block depending on whether a transform was performed.

[0010] Depending on the use case, it may be desirable to compress video data losslessly (i.e., without encoding loss). The CU can be reversibly encoded by skipping both the transformation step and the quantization step. In the TSRC process, quantization can be avoided by setting the "quantization parameter" to a value that does not indicate quantization. However, as described above, the TSRC process is only suitable for classes of video content such as screen content. Therefore, forcing the reversible encoding of video data to use the TSRC process is not optimal. For reversible encoding, it is desirable to have more flexible options according to the statistics of the video data to be encoded, while minimizing the amount of additional logic required to support the additional flexibility.

SUMMARY OF THE INVENTION

[0011] An object of the present invention is to substantially overcome or at least improve one or more drawbacks of existing configurations.

[0012] One aspect of the present invention provides a method for decoding a sub-block of residual coefficients of a transform block from a video bitstream, the method comprising determining whether sign bit hiding is used for the sub-block, the determination being based on the value of a transform skip flag determined for the sub-block and the value of a sign bit hiding flag associated with the sub-block, and if sign bit hiding is not used, decoding a number of sign bits equal to the number of significant coefficients in the sub-block, and decoding the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0013] According to another aspect, sign bit hiding is used if the sign bit hiding flag has a value of TRUE, the transform skip flag has a value of FALSE, and the difference between the first significant position and the last significant position of the sub-block is greater than 3.

[0014] According to another aspect, when the sign bit hiding flag has a value of TRUE and the transform skip flag has a value of TRUE, sign bit hiding is not used.

[0015] According to another aspect, the method further comprises, when it is determined that sign bit hiding is used, decoding a number of sign bits equal to one less than the number of valid coefficients in the sub-block, and determining additional sign bits from the sum of the parities of the valid coefficients of the sub-block.

[0016] Another aspect of the present invention provides a method for decoding a sub-block of residual coefficients of a transform block from a video bit stream, the method comprising determining whether sign bit hiding is used for the sub-block, the determination being based on the value of a sign bit hiding flag and the value of a quantization parameter associated with the sub-block, decoding a number of sign bits equal to the number of significant coefficients in the sub-block when sign bit hiding is not used, and decoding the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0017] According to another aspect, when the sign bit hiding flag has a value of TRUE and the quantization parameter is equal to 4, sign bit hiding is not used.

[0018] According to another aspect, when the sign bit hiding flag has a value of TRUE, the quantization parameter is not equal to 4, and the difference between the first significant position and the last significant position of the sub-block is greater than 3, sign bit hiding is used.

[0019] Another aspect of the present invention provides a method for decoding a sub-block of residual coefficients of a transform block from a video bitstream. The method includes determining whether sign-bit hiding is used for the sub-block, the determination being based on the value of a sign-bit hiding flag and the value of a TSRC invalid flag. When sign-bit hiding is not used, decoding a number of sign bits equal to the number of significant coefficients in the sub-block, and decoding the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0020] According to another aspect, sign-bit hiding is used when the sign-bit hiding flag has a value of TRUE, the TSRC invalid flag has a value of FALSE, and the difference between the first significant position and the last significant position of the sub-block is greater than 3.

[0021] According to another aspect, the sign-bit hiding flag has a value of TRUE and the TSRC invalid flag has a value of TRUE.

[0022] Another aspect of the present invention provides a non-transitory computer-readable medium storing a computer program for implementing a method for decoding a sub-block of residual coefficients of a transform block from a video bitstream. The method includes determining whether sign-bit hiding is used for the sub-block, the determination being based on the value of a transform skip flag determined for the sub-block and the value of a sign-bit hiding flag associated with the sub-block. When sign-bit hiding is not used, decoding a number of sign bits equal to the number of significant coefficients in the sub-block, and decoding the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0023] Another aspect of the present invention provides a system having a memory and a processor, the processor being configured to execute code stored in the memory to perform a method of decoding a sub-block of residual coefficients of a transform block from a video bitstream, the method comprising determining whether sign bit hiding is used for the sub-block, the determination being based on a value of a transform skip flag determined for the sub-block and a value of a sign bit hiding flag associated with the sub-block, and if sign bit hiding is not used, decoding a number of sign bits equal to the number of valid coefficients in the sub-block and decoding the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0024] Another aspect of the present invention provides a video decoder configured to receive a sub-block of residual coefficients of a transform block from a video bitstream, determine whether sign bit hiding is used for the sub-block, the determination being based on a value of a transform skip flag determined for the sub-block and a value of a sign bit hiding flag associated with the sub-block, and if sign bit hiding is not used, decode a number of sign bits equal to the number of significant coefficients in the sub-block and decode the sub-block by reconstructing the residual coefficients of the sub-block using the decoded sign bits.

[0025] Other aspects are also described.

Brief Description of the Drawings

[0026] At least one embodiment of the present invention will be described with reference to the following drawings and appendices.

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8A

Figure 8B

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0027] When referring to steps and / or features having the same reference numerals in one or more of the accompanying drawings, those steps and / or features have the same function or operation for the purposes of this specification, unless the contrary intention appears.

[0028] As described above, reversible coding may desirably be supported by existing building blocks of a codec. However, since various classes of video data encoded in a reversible manner cannot guarantee that the TSRC process exhibits the statistical characteristics for which it was designed, using the TSRC process exclusively for reversible coding can produce sub-optimal coding performance. Therefore, greater flexibility in the selection of high-level building blocks where reversible coding can be used enables excellent coding performance with a minimal additional complexity to the overall design.

[0029] FIG. 1 is a schematic block diagram showing the functional modules of a video encoding and decoding system 100. The system 100 includes a source device 110 and a destination device 130. A communication channel 120 is used to communicate the encoded video information from the source device 110 to the destination device 130. In some configurations, the source device 110 and the destination device 130 can each or both include a respective cellular phone handset or “smartphone,” in which case the communication channel 120 is a wireless channel. In other configurations, the source device 110 and the destination device 130 can include video conferencing equipment, in which case the communication channel 120 is typically a wired channel such as an Internet connection. Further, the source device 110 and the destination device 130 can include any of a wide range of devices including those that support applications in which encoded video data is captured onto some computer-readable storage medium such as a wireless television broadcast, a cable television application, an Internet video application (including streaming), and a hard disk drive in a file server.

[0030] As shown in FIG. 1, the source device 110 includes a video source 112, a video encoder 114, and a transmitter 116. The video source 112 has a source of captured video frame data (shown as 113), typically an imaging sensor or the like, a previously captured video sequence stored on a non-transitory recording medium, or video supplied from a remote imaging sensor. The video source 112 can also be the output of a computer graphics card, for example, displaying the video output of an operating system and various applications running on a computing device such as a tablet computer. Examples of source devices 110 that can include an imaging sensor as the video source 112 include smartphones, video cameras, professional video cameras, and network video cameras.

[0031] As will be further described with reference to FIG. 3, video encoder 114 converts (or "encodes") the captured frame data from video source 112 (indicated by arrow 113) into a bitstream (indicated by arrow 115). Bitstream 115 is transmitted by transmitter 116 via communication channel 120 as encoded video data (or "encoded video information"). Bitstream 115 may also be stored in a non-transitory storage device 122, such as a "flash" memory or a hard disk drive, until it is later transmitted via communication channel 120 or instead of transmission via communication channel 120.

[0032] Destination device 130 includes a receiver 132, a video decoder 134, and a display device 136. Receiver 132 receives the encoded video data from communication channel 120 and passes the received video data to video decoder 134 as a bitstream (indicated by arrow 133). Video decoder 134 then outputs the decoded frame data to display device 136 (indicated by arrow 135) to play the video data. The decoded frame data 135 has the same chroma format as frame data 113. Examples of display device 136 include liquid crystal displays such as cathode ray tubes, smartphones, tablet computers, computer monitors, or stand-alone television sets. Also, the functionality of each of source device 110 and destination device 130 may be implemented in a single device, examples of which include mobile phone handsets and tablet computers.

[0033] Notwithstanding the exemplary devices described above, each of source device 110 and destination device 130 can typically be configured within a general-purpose computing system via a combination of hardware and software components. FIG. 2A shows such a computer system 200, including a computer module 201, input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280, and output devices including a printer 215, a display device 214 that can be configured as a display 136, and a speaker 217. An external modem (modem) transceiver device 216 can be used by computer module 201 to communicate with communication network 220 via connection 221. Communication network 220, which can represent communication channel 120, may be a wide area network (WAN) such as the Internet, a cellular telecommunications network, or a private WAN. If connection 221 is a telephone line, modem 216 may be a conventional “dial-up” modem. Alternatively, if connection 221 is a high-capacity (e.g., cable or optical) connection, modem 216 may be a broadband modem. A wireless modem may also be used for a wireless connection to communication network 220. Transceiver device 216 can provide the functionality of transmitter 116 and receiver 132, and communication channel 120 can be embodied within connection 221.

[0034] Computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 can have a semiconductor random access memory (RAM) and a semiconductor read only memory (ROM). The computer module 201 also includes several input / output (I / O) interfaces, including an audio / video interface 207 that couples to a video display 214, a speaker 217, and a microphone 280, a keyboard 202, a mouse 203, a scanner 226, a camera 227, and an I / O interface 213 that couples to an optional joystick or other human interface device (not shown), as well as an interface 208 for an external modem 216 and a printer 215. The signal from the audio / video interface 207 to the computer monitor 214 is generally the output of a computer graphics card. In some implementations, the modem 216 may be incorporated within the computer module 201, for example, within the interface 208. The computer module 201 also has a local network interface 211, which enables the connection of the computer system 200 to a local area communication network 222 known as a local area network (LAN) via a connection 223. As shown in Figure 2A, the local communication network 222 can also be coupled to a wide network 220 via a connection 224, typically including a so-called "firewall" device or a device with similar functionality. The local network interface 211 TM circuit card, Bluetooth TM can include a wireless configuration or an IEEE 802.11 wireless configuration, although many other types of interfaces may be implemented for the interface 211. The local network interface 211 can also provide the functionality of a transmitter 116, and the receiver 132 and the communication channel 120 can also be embodied in the local communication network 222.

[0035] The I / O interfaces 208 and 213 can provide either or both of serial connectivity and parallel connectivity. The former is typically implemented in accordance with the Universal Serial Bus (USB) standard and has a corresponding USB connector (not shown). A memory device 209 is provided, typically including a hard disk drive (HDD) 210. Other memory devices such as a floppy disk drive and a magnetic tape drive (not shown) can also be used. The optical disk drive 212 is typically provided to function as a non-volatile source of data. Optical disks (e.g., CD-ROM, DVD, Blu ray Disc TM ), USB-RAM, portable, external hard drives, and floppy disks, etc., portable memory devices can be used, for example, as a suitable source of data for the computer system 200. Typically, any of the HDD 210, optical drive 212, networks 220 and 222 may be configured to operate as a video source 112 or as a destination for decoded video data to be stored for playback via the display 214. The source device 110 and the destination device 130 of the system 100 may be embodied in the computer system 200.

[0036] The components 205 - 213 of the computer module 201 typically communicate via an interconnect bus 204 in a manner that provides a conventional mode of operation for the computer system 200 known to those skilled in the art. For example, the processor 205 is coupled to the system bus 204 using connection 218. Similarly, the memory 206 and the optical disk drive 212 are coupled to the system bus 204 by connection 219. Examples of computers for which the above configuration is executable include IBM-PC and compatibles, Sun SPARC stations, Apple Mac TM or similar computer systems.

[0037] When appropriate or necessary, video encoder 114 and video decoder 134, and the methods described below can be implemented using computer system 200. Specifically, video encoder 114, video decoder 134, and the methods described can be implemented as one or more software application programs 233 executable within computer system 200. Specifically, the steps of video encoder 114, video decoder 134, and the methods described are executed by instructions 231 (see FIG. 2B) within software 233 executed within computer system 200. The software instructions 231 may each be formed as one or more code modules for performing one or more specific tasks. The software may also be divided into two separate parts, in which case the code modules corresponding to the first part execute the methods described, and the code modules corresponding to the second part manage the user interface between the first part and the user.

[0038] The software can be stored, for example, on a computer-readable medium including the storage devices described below. The software is loaded from the computer-readable medium into computer system 200 and then executed by computer system 200. A computer-readable medium having such software or a computer program recorded thereon is a computer program product. The use of the computer program product in computer system 200 preferably provides an advantageous apparatus for implementing video encoder 114, video decoder 134, and the methods described.

[0039] Software 233 is typically stored in HDD 210 or memory 206. The software is loaded from a computer-readable medium into computer system 200 and executed by computer system 200. Thus, for example, software 233 can be stored on an optically readable disk storage medium (e.g., CD-ROM) 225 that is read by optical disk drive 212.

[0040] In some cases, application program 233 may be encoded on one or more CD-ROMs 225 and supplied to the user, read via the corresponding drive 212, or alternatively, read by the user from network 220 or 222. Further, the software can also be loaded into computer system 200 from other computer-readable media. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides instructions and / or data recorded for execution and / or processing to computer system 200. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs TM , hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, etc., regardless of whether such devices are internal or external to computer module 201. Examples of transitory or non-tangible computer-readable transmission media that can also participate in providing software, application programs, instructions and / or video data or encoded video data to computer module 401 include wireless or infrared transmission channels, as well as network connections to another computer or network-connected device, and the Internet or intranet including electronic mail transmissions and information recorded on websites such as.

[0041] The second part of the application program 233 and the corresponding code module described above may be executed to implement one or more graphical user interfaces (GUIs) that are rendered on the display 214 or otherwise presented. Typically, through the operation of the keyboard 202 and the mouse 203, a user of the application and the computer system 200 can operate the interface in a functionally adaptable manner and provide control commands and / or input to the application associated with the GUI. Other forms of functionally adaptable user interfaces can also be implemented, such as an audio interface that utilizes speech prompts output via the speaker 217 and user voice commands input via the microphone 280.

[0042] FIG. 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents the logical aggregate of all memory modules (including the HDD 209 and the semiconductor memory 206) accessible by the computer module 201 of FIG. 2A.

[0043] When the computer module 201 is first powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in the ROM 249 of the semiconductor memory 206 in FIG. 2A. Hardware devices such as the ROM 249 that store software are sometimes called firmware. The POST program 250 inspects the hardware within the computer module 201 to confirm that it functions properly and usually checks the processor 205, the memory 234 (209, 206), and the basic input / output system software (BIOS) module 251, which is also typically stored in the ROM 249, for correct operation. When the POST program 250 is executed successfully, the BIOS 251 activates the hard disk drive 210 in FIG. 2A. When the hard disk drive 210 is activated, a bootstrap loader program 252 resident on the hard disk drive 210 is executed via the processor 205. As a result, the operating system 253 is loaded into the RAM memory 206, and the operating system 253 begins to operate thereon. The operating system 253 is a system-level application executable by the processor 205 and satisfies various high-level functions including processor management, memory management, device management, storage management, software application interface, and general-purpose user interface.

[0044] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without colliding with the memory allocated to another process. Further, the different types of memory available in the computer system 200 of FIG. 2A must be used appropriately so that each process can execute effectively. Thus, the aggregated memory 234 is not intended to indicate how a particular segment of memory is allocated (unless otherwise specified), but rather to provide a general view of the memory accessible by the computer system 200 and how such segments are used.

[0045] As shown in FIG. 2B, the processor 205 includes a number of functional modules including a control unit 239, an arithmetic logic unit (ALU) 240, and sometimes a local or internal memory 248, sometimes called a cache memory. The cache memory 248 typically includes a number of storage registers 244-246 within a register section. One or more internal buses 241 functionally interconnect these functional modules. The processor 205 also typically has one or more interfaces 242 for communicating with external devices via the system bus 204 using connection 218. The memory 234 is coupled to the bus 204 using connection 219.

[0046] The application program 233 includes a sequence 231 of instructions that may include conditional branch and loop instructions. The program 233 may also include data 232 used for the execution of the program 233. The instructions 231 and the data 232 are stored at memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending on the relative sizes of the instructions 231 and the memory locations 228 to 230, a particular instruction can be stored in a single memory location as indicated by the instruction shown at memory location 230. Alternatively, the instructions may be segmented into several parts, each stored in a separate memory location, as indicated by the instruction segments shown at memory locations 228 and 229.

[0047] Generally, the processor 205 is provided with a set of instructions to be executed therein. The processor 205 waits for subsequent input, and in response to this input, the processor 205 reacts by executing another set of instructions. Each input can be provided from one or more of several sources, including data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into the corresponding reader 212, all shown in FIG. 2A. When a set of instructions is executed, data may be output. Execution may also include storing data or variables in the memory 234.

[0048] The video encoder 114, the video decoder 134, and the methods described can use input variables 254 stored at corresponding memory locations 255, 256, 257 within the memory 234. The video encoder 114, the video decoder 134, and the methods described generate output variables 261, which are stored at corresponding memory locations 262, 263, 264 within the memory 234. Intermediate variables 258 may be stored at memory locations 259, 260, 266, and 267.

[0049] Referring to the processor 205 of FIG. 2B, the registers 244, 245, 246, the arithmetic logic unit (ALU) 240, and the control unit 239 cooperate to execute a sequence of micro-operations necessary to perform "fetch, decode, and execute" cycles for all instructions within the instruction set that constitutes the program 233. Each fetch, decode, and execute cycle includes a fetch operation of fetching or reading the instruction 231 from the memory locations 228, 229, 230 a decode operation in which the control unit 239 determines which instruction has been fetched and an operation of executing the instruction by the control unit 239 and / or the ALU 240 .

[0050] Thereafter, further fetch, decode, and execute cycles for the next instruction can be executed. Similarly, the control unit 239 can execute a store cycle of storing or writing a value to the memory location 232.

[0051] Each step or sub-process in the methods of FIGS. 19 to 14 described below is associated with one or more segments of the program 233, and typically the register sections 244, 245, 247, the ALU 240, and the control unit 239 within the processor 205 cooperate to perform fetch, decode, and execute cycles for all instructions within the instruction set for the noted segments of the program 233.

[0052] FIG. 3 is a schematic block diagram showing the functional modules of the video encoder 114. FIG. 4 is a schematic block diagram showing the functional modules of the video decoder 134. Generally, data passes between the functional modules of the video decoder 134 and the video encoder 114 in groups of samples or coefficients, such as by splitting the blocks into sub-blocks of a fixed size, or as an array. The video encoder 114 and the video decoder 134 can be implemented using the general-purpose computer system 200 as shown in FIGS. 2A and 2B, and the various functional modules reside on the hard disk drive 205 and are controlled by the processor 205 during its execution, such as one or more software code modules of the software application program 233, and can be realized by software executable within the computer system 200, or by dedicated hardware within the computer system 200. Alternatively, the video encoder 114 and the video decoder 134 may be implemented by a combination of software and dedicated hardware executable within the computer system 200. The video encoder 114, the video decoder 134, and the methods described can alternatively be implemented in dedicated hardware, such as one or more integrated circuits that perform the functions or sub-functions of the methods described. Such dedicated hardware can include a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific standard product (ASSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or one or more microprocessors and associated memories. In particular, the video encoder 114 includes modules 310 to 386, and the video decoder 134 includes modules 420 to 496, which can each be implemented as one or more software code modules of the software application program 233.

[0053] The video encoder 114 of FIG. 3 is an example of a general-purpose video coding (VVC) video coding pipeline, but other video codecs can also be used to execute the processing stages described herein. The video encoder 114 receives imaged frame data 113, such as a series of frames, and each frame includes one or more color channels. The frame data 113 may be in any chroma format, for example, 4:0:0, 4:2:0, 4:2:2, or 4:4:4 chroma format. The block partitioner 310 first divides the frame data 113 into CTUs, which are generally square-shaped and configured such that a specific size is used for the CTUs. The size of the CTUs can be, for example, 64×64, 128×128, or 256×256 luma samples. The block partitioner 310 further divides each CTU into one or more CUs according to a luma coding tree and a chroma coding tree. The CUs have various sizes and may include both square and non-square aspect ratios. In the VVC standard, the CUs, PUs, and TUs always have side lengths that are powers of two. Thus, the current CU, represented as 312, is output from the block partitioner 310 and proceeds according to iterations over one or more blocks of the CTU according to the chroma coding tree and the luma coding tree of the CTU. Options for dividing the CTU into CUs are further described below with reference to FIGS. 5 and 6.

[0054] The CTUs obtained from the first division of the frame data 113 are scanned in raster scan order and can be grouped into one or more "slices". The slice may be an "intra" (or "I") slice. An intra slice (I slice) indicates that all CUs within the slice are intra-predicted. Alternatively, the slice may be a uni- or bi-predicted (respectively, "P" or "B" slice), indicating the further availability of uni- and bi-prediction in the slice, respectively.

[0055] For each CTU, the video encoder 114 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 310 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides low distortion and high compression efficiency. This test generally involves Lagrangian optimization, whereby candidate CBs are evaluated based on a weighted combination of rate (encoding cost) and distortion (error with respect to the input frame data 113). The "best" candidate CB (the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding into the bitstream 115. The evaluation of candidate CBs includes the option of using a CB for a given area, or further dividing the area according to various partitioning options and encoding each of the resulting smaller areas with a further CB, or further further dividing the area. As a result, both the CB and the coding tree itself are selected in the search stage.

[0056] The video encoder 114 generates a predicted block (PB) indicated by the arrow 320 for each CB, e.g., CB312. The PB320 is a prediction of the content of the associated CB312. The subtractor module 322 generates a difference, indicated as 324 (or "residual", referring to the difference within the spatial region), between the PB320 and the CB312. The residual 324 is the block-size difference between corresponding samples in the PB320 and the CB312. The residual 324 is transformed, quantized, and represented as a transformed block (TB) indicated by the arrow 336. The PB320 and the associated TB336 are typically selected from among many possible candidate CBs, for example, based on the evaluated cost or distortion.

[0057] A candidate coding block (CB) is a CB resulting from one of the prediction modes available to video encoder 114 for the associated PB and the resulting residual. Each candidate CB results in one or more corresponding TBs. TB 336 is the quantized and transformed representation of residual 324. When combined with the PB predicted at video decoder 114, TB 336 reduces the difference between the decoded CB and the original CB 312 at the expense of additional signals in the bitstream.

[0058] Accordingly, each candidate coding block (CB), i.e., a prediction block (PB) combined with a transform block (TB), has an associated coding cost (or "rate") and an associated difference (or "distortion"). The rate is typically measured in bits. The distortion of a CB is typically estimated as the difference of sample values such as the sum of absolute differences (SAD) or the sum of squared differences (SSD). The estimation obtained from each candidate PB is determined by mode selector 386 using residual 324 to determine the prediction mode (represented by arrow 388). The estimation of the coding cost associated with the residual coding corresponding to each candidate prediction mode can be performed at a much lower cost than the entropy coding of the residual. Thus, a number of candidate modes can be evaluated to determine the optimal mode in rate-distortion detection.

[0059] Determining the optimal mode from the perspective of rate distortion is typically achieved using a variation of Lagrangian optimization. The selection of the prediction mode 388 typically involves determining the encoding cost for the residual data resulting from the application of a particular prediction mode. The encoding cost can be approximated by using the "sum of absolute transform differences" (SATD), thereby obtaining an estimated transform residual cost using a relatively simple transform such as the Hadamard transform. In some embodiments that use relatively simple transforms, the cost obtained from the simplified estimation method is monotonically related to the actual cost that would otherwise be determined from a full evaluation. In embodiments having a monotonically related estimated cost, the same decision (i.e., prediction mode) can be made using the simplified estimation method while reducing the complexity of the video encoder 114. To allow for possible non-monotonicity in the relationship between the estimated cost and the actual cost, a simplified estimation method can be used to generate a list of the best candidates. Non-monotonicity can result, for example, from additional mode decisions available for encoding the residual data. The list of the best candidates can be any number. Using the best candidates, a more complete search can be performed to establish the optimal mode selection for encoding the residual data for each of the candidates, enabling the final selection of the prediction mode 388 along with other mode decisions.

[0060] Prediction modes are broadly classified into two categories. The first category is "intra-frame prediction" (also called "intra prediction"). In intra-frame prediction, a prediction for a block is generated, and the generation method may use other samples obtained from the current frame. Types of intra prediction include intra planner, intra DC, intra angle, and matrix weighted intra prediction (MIP). In the case of intra-predicted PB, different intra prediction modes can be used for luma and chroma, and thus intra prediction is mainly described with respect to operation on PB. Further, the chroma CB may be predicted from luma samples located at the same location by cross-component linear model prediction.

[0061] The second category of prediction mode is "inter-frame prediction" (also called "inter prediction"). In inter-frame prediction, the prediction of a block is generated using samples from one or two frames preceding the current frame in the order in which the frames in the bitstream are encoded. Further, for inter-frame prediction, a single coding tree is typically used for both the luma channel and the chroma channel. The encoding order of the frames in the bitstream may be different from the order of the frames at capture or display. When one frame is used for prediction, the block is said to be "single prediction" and has one associated motion vector. When two frames are used for prediction, the block is said to be "dual predicted" and has two associated motion vectors. In the case of a P slice, each CU can be intra predicted or single predicted. In the case of a B slice, each CU can be intra predicted, single predicted, or dual predicted. Frames are typically encoded using a "group of pictures" structure that allows for a temporal hierarchy of frames. The temporal hierarchy of frames allows a frame to reference preceding and subsequent pictures in the order in which the frames are to be displayed. Images are encoded in an order necessary to ensure that the dependencies for decoding each frame are satisfied.

[0062] The subcategory of inter prediction is called "skip mode". The inter prediction mode and the skip mode are described as two separate modes. However, both the inter prediction mode and the skip mode include motion vectors that reference a block of samples from a preceding frame. Inter prediction includes an encoded motion vector delta that specifies the motion vector relative to a motion vector predictor. The motion vector predictor is obtained from a list of one or more candidate motion vectors selected by a "merge index". The encoded motion vector delta provides a spatial offset to the selected motion vector prediction. Also, inter prediction uses the encoded residuals within the bitstream 133. The skip mode uses only an index (also called the "merge index") to select one of several motion vector candidates. The selected candidate is used without further signaling. Also, the skip mode does not support encoding of residual coefficients. When the skip mode is used, the absence of encoded residual coefficients means that there is no need to perform a transform for the skip mode. Thus, the skip mode typically does not cause pipeline processing problems, which can occur for intra-predicted CUs and inter-predicted CUs. Due to the limited signaling of the skip mode, the skip mode is useful for achieving very high compression performance when a relatively high-quality reference frame is available. Bi-predicted CUs in higher temporal layers of a random access picture group structure typically have high-quality reference pictures and motion vector candidates that accurately reflect the underlying motion.

[0063] Samples are selected according to motion vectors and reference picture indices. The motion vectors and reference picture indices are applied to all color channels, and thus, inter prediction is mainly described with respect to operation on PUs rather than PBs. Different techniques can be applied to generate PUs within each category (i.e., intra and inter frame prediction). For example, intra prediction can use values from adjacent rows and columns of previously reconstructed samples in combination with a direction to generate a PU according to a predetermined filtering and generation process. Alternatively, a PU may be described using a small number of parameters. The inter prediction method may vary in the number and accuracy of motion parameters. Motion parameters typically include a reference frame index indicating which reference frame should be used from a list of reference frames and a spatial transformation for each of the reference frames, but can include more frames, special frames, or complex affine parameters such as scaling and rotation. Further, a predetermined motion refinement process can be applied to generate high density motion estimation based on the referenced sample block.

[0064] Lagrange processing or similar optimization processing can be adopted to select both the optimal partitioning of the CTU into CUs (by the block partitioner 310) and the selection of the best prediction mode from multiple possibilities. Through the application of the Lagrange optimization process for candidate modes in the mode selector module 386, the prediction mode with the lowest cost measurement is selected as the "best" mode. The lowest cost mode is the selected prediction mode 388 and is also encoded into the bitstream 115 by the entropy encoder 338. The selection of the prediction mode 388 by the operation of the mode selector module 386 extends to the operation of the block partitioner 310. For example, the candidates for the selection of the prediction mode 388 can include the modes applicable to a given block and, further, the modes applicable to a plurality of smaller blocks that are collectively arranged with the given block. When including the modes applicable to a given block and the smaller collocated blocks, the process of implicitly selecting candidates is also the process of determining the best hierarchical decomposition of the CTU into CUs.

[0065] In the second stage of the operation of the video encoder 114 (referred to as the "encoding" stage), the selected luma coding tree and the selected chroma coding tree, and thus the iterations over each selected CU, are performed in the video encoder 114. In the iterations, the CUs are encoded into the bitstream 115 as further described herein.

[0066] Entropy encoder 338 supports both variable-length coding of syntax elements and arithmetic coding of syntax elements. Arithmetic coding is supported using a context-adaptive binary arithmetic coding (CABAC) process. The arithmetically coded syntax elements consist of a sequence of one or more "bins". A bin has a value of "0" or "1", similar to a bit. A bin is not encoded as discrete bits into bitstream 115. A bin has an associated prediction (or "likelihood" or "most likely") value and an associated probability, known as a "context". When the actual bin to be encoded matches the predicted value, the "most probable symbol" (MPS) is encoded. Encoding the most probable symbol is relatively inexpensive in terms of bits consumed. If the actual bin to be encoded does not match the likely value, the "least probable symbol" (LPS) is encoded. Encoding the least probable symbol has a relatively high cost in terms of bits consumed. Bin coding techniques enable efficient coding of bins where the probability of "0" versus "1" is skewed. For syntax elements with two possible values (i.e., "flags"), a single bin is sufficient. For syntax elements with a large number of possible values, a sequence of bins is required.

[0067] The presence of subsequent bins in the sequence may be determined based on the value of previous bins in the sequence. Additionally, each bin can be associated with more than one context. The choice of a particular context can depend on previous bins of the syntax element, bin values of adjacent syntax elements (i.e., from adjacent blocks), etc. Each time a context-coded bin is encoded, the context (if any) selected for that bin is updated in a way that reflects the new bin value. Thus, the binary arithmetic coding scheme is said to be adaptive.

[0068] Also, what the video encoder 114 supports is a bin without context (a "bypass bin"). The bypass bin is encoded assuming an equiprobable distribution between "0" and "1". Therefore, each bin occupies 1 bit within the bitstream 115. Without context, memory is saved, complexity is reduced, and thus, when the distribution of the values of a particular bin is not skewed, the bypass bin is used.

[0069] The entropy encoder 338 encodes the prediction mode 388 using a combination of context-encoded bins and bypass-encoded bins. For example, when the prediction mode 388 is an intra prediction mode, a list of "most probable modes" is generated in the video encoder 114. The list of most probable modes is typically of a fixed length such as 3 or 6 modes and can include the modes encountered in previous blocks. The context-encoded bin encodes a flag indicating whether the prediction mode is one of the most probable modes. If the intra prediction mode 388 is one of the most probable modes, additional signaling using bypass-encoded bins is encoded. The encoded additional signaling indicates, for example, using a truncated unary bitstring, which of the most probable modes corresponds to the intra prediction mode 388. Otherwise, the intra prediction mode 388 is encoded as a "remaining mode". Encoding as a remaining mode uses an alternative syntax such as a fixed-length code that is also encoded using bypass-encoded bins to represent intra prediction modes other than those present in the most probable mode list.

[0070] The multiplexer module 384 outputs the PB320 according to the determined best prediction mode 388 and selects from the tested prediction modes of each candidate CB. The candidate prediction modes do not necessarily have to include all possible prediction modes supported by the video encoder 114.

[0071] When PB320 is determined and selected, and PB320 is subtracted from the original sample block by subtractor 322, a residual represented by 324 with the lowest encoding cost is obtained and undergoes irreversible compression. The irreversible compression process includes steps of transformation, quantization, and entropy encoding. The forward primary transformation module 326 applies a forward transformation to the residual 324, transforms the residual 324 from the spatial domain to the frequency domain, and generates primary transformation coefficients represented by arrow 328. The primary transformation coefficients 328 are passed to the forward secondary transformation module 330, which generates transformation coefficients represented by arrow 332 by performing a non-separable secondary transformation (NSST) operation. The forward primary transformation is typically separable and typically uses a type II discrete cosine transform (DCT-2) to transform a set of rows and then a set of columns of each block, although type VII discrete sine transform (DST-7) and type VIII discrete cosine transform (DCT-8) are also available, for example, horizontally for block widths not exceeding 16 samples and vertically for block heights not exceeding 16 samples. The transformation of each set of rows and columns is performed by first applying a one-dimensional transformation to each row of the block to generate an intermediate result and then applying a one-dimensional transformation to each column of the intermediate result to generate the final result. The forward secondary transformation is generally a non-separable transformation, which is applied only to the residual of the intra-predicted CU and may nevertheless be bypassed. The forward secondary transformation operates on either 16 samples (arranged as the upper left 4x4 sub-block of the primary transformation coefficients 328) or 64 samples (arranged as 4 4x4 sub-blocks of the primary transformation coefficients 328, arranged as the upper left 8x8 coefficients). Furthermore, the matrix coefficients for the forward secondary transform are selected from multiple sets according to the intra prediction mode of the CU so that two sets of coefficients are available for use. Using one of the sets of matrix coefficients, i.e., bypassing the forward secondary transform, is signaled by the syntax element of "nsst_index" and is encoded using truncated unary binary to represent value zero (no secondary transform is applied), one (the first set of selected matrix coefficients), or two (the second set of selected matrix coefficients).

[0072] Video encoder 114 may also select to skip both the primary transform and the secondary transform, known as the "transform skip" mode. Skipping the transform is suitable for residual data that lacks the appropriate correlation to reduce the encoding cost through the representation as transform basis functions. Certain types of content, such as relatively simple computer-generated graphics, may exhibit similar behavior. When the transform skip mode is used, the transform coefficients 332 are the same as the residual coefficients 324.

[0073] The transformation coefficient 332 is passed to the quantizer module 334. In module 334, quantization is performed by the "quantization parameter", generating the quantized coefficient represented by arrow 336. The quantization parameter is constant for a given TB, thus providing uniform scaling for the generation of the residual coefficients for the TB. Non-uniform scaling is also possible by applying a "quantization matrix", whereby the scaling coefficient applied to each residual coefficient is derived from a combination of the quantization parameter and the corresponding entry in a scaling matrix typically having a size equal to the size of the TB. The scaling matrix can have a size smaller than the size of the TB, and when applied to the TB, the nearest neighbor approach is used to provide the scaling value for each residual coefficient from a scaling matrix smaller in size than the TB size. The quantized coefficient 336 is supplied to the entropy encoder 338 for encoding in the bitstream 115. Typically, the quantized coefficients of each TB having at least one significant quantized coefficient are scanned according to a scan pattern to generate a list of values in sorted order. The scan pattern generally scans the TB as a sequence of 4x4 "sub-blocks", providing a regular scan operation at the granularity of 4x4 sets of residual coefficients, and the arrangement of the sub-blocks depends on the size of the TB. Additionally, the prediction mode 388 and the corresponding block partitioning are also encoded in the bitstream 115.

[0074] As described above, the video encoder 114 requires access to a frame representation corresponding to the frame representation seen by the video decoder 134. Thus, the quantization coefficients 336 are also inverse quantized by the inverse quantizer module 340 to generate the reconstructed transform coefficients represented by arrow 342. The reconstructed transform coefficients 342 pass through the inverse secondary transform module 344 to generate the reconstructed primary transform coefficients represented by arrow 346. The reconstructed primary transform coefficients 346 are passed to the inverse primary transform module 348 to generate the reconstructed residual samples of the TU represented by arrow 350. The type of inverse transform performed by the inverse secondary transform module 344 corresponds to the type of forward transform performed by the forward secondary transform module 330. The type of inverse transform performed by the inverse primary transform module 348 corresponds to the type of primary transform performed by the primary transform module 326. The addition module 352 adds the reconstructed residual samples 350 and the PU 320 to generate the reconstructed samples of the CU (indicated by arrow 354).

[0075] The reconstructed sample 354 is passed to a reference sample cache 356 and an in-loop filter module 368. The reference sample cache 356 is typically implemented using static RAM on an ASIC (thus avoiding costly off-chip memory accesses) and provides the minimum sample storage necessary to satisfy the dependencies for generating intra-frame PBs for subsequent CUs within a frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a CTU row, which are used by the next row and column buffering of the CTU, and the extent is set by the height of the CTU. The reference sample cache 356 supplies reference samples (indicated by arrow 358) to a reference sample filter 360. The sample filter 360 applies a smoothing operation to generate filtered reference samples (indicated by arrow 362). The filtered reference samples 362 are used by an intra-frame prediction module 364 to generate an intra prediction block of samples represented by arrow 366. For each candidate intra prediction mode, the intra-frame prediction module 364 generates a block of samples, i.e., 366.

[0076] The in-loop filter module 368 applies several filtering stages to the reconstructed sample 354. The filtering stages include a "deblocking filter" (DBF) that applies smoothing aligned at CU boundaries to reduce artifacts resulting from discontinuities. Another filtering stage present in the in-loop filter module 368 is the "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. A further available filtering stage in the in-loop filter module 368 is the "sample adaptive offset" (SAO) filter. The SAO filter first classifies the reconstructed samples into one or more categories and operates by applying an offset at the sample level according to the assigned category.

[0077] The filtered samples represented by arrow 370 are output from the in-loop filter module 368. The filtered samples 370 are stored in the frame buffer 372. The frame buffer 372 typically has a capacity for storing several (e.g., up to 16) pictures and is thus stored in the memory 206. Since the frame buffer 372 requires a large memory consumption, it is typically not stored using on-chip memory. Therefore, access to the frame buffer 372 is costly in terms of memory bandwidth. The frame buffer 372 provides a reference frame (represented by arrow 374) to the motion estimation module 376 and the motion compensation module 380.

[0078] The motion estimation module 376 estimates several "motion vectors" (shown as 378), each of which is a Cartesian space offset from the position of the current CB and refers to a block within one of the reference frames in the frame buffer 372. A filtered block of reference samples (represented as 382) is generated for each motion vector. The filtered reference samples 382 form additional candidate modes available for potential selection by the mode selector 386. Further, for a given CU, the PU320 may be formed using one reference block ("single prediction") or two reference blocks ("dual prediction"). For the selected motion vector, the motion compensation module 380 generates the PB320 according to a filtering process that supports sub-pixel accuracy within the motion vector. Thus, the motion estimation module 376 (operating on many candidate motion vectors) can perform a simplified filtering process compared to that of the motion compensation module 380 (operating only on the selected candidates) to achieve a reduced computational complexity. When the video encoder 114 selects inter-prediction for a CU, the motion vector 378 is encoded into the bitstream 115.

[0079] The video encoder 114 of FIG. 3 is described with reference to Versatile Video Coding (VVC), but other video coding standards or implementations can also use the processing stages of modules 310 to 386. The frame data 113 (and bitstream 115) can be read from (or written to) the memory 206, hard disk drive 210, CD-ROM, Blu-ray disc TM , or other computer-readable storage media. Further, the frame data 113 (and bitstream 115) may be received (or transmitted) from an external source such as a communication network 220 or a server connected to a radio frequency receiver.

[0080] The video decoder 134 is shown in FIG. 4. The video decoder 134 of FIG. 4 is an example of a Versatile Video Coding (VVC) video decoding pipeline, but other video codecs can also be used to perform the processing stages described herein. As shown in FIG. 4, the bitstream 133 is input to the video decoder 134. The bitstream 133 can be read from the memory 206, hard disk drive 210, CD-ROM, Blu-ray disc TM , or other non-transitory computer-readable storage media. Alternatively, the bitstream 133 may be received from an external source such as a communication network 220 or a server connected to a radio frequency receiver. The bitstream 133 contains encoded syntax elements representing the decoded captured frame data.

[0081] The bitstream 133 is input to the entropy decoder module 420. The entropy decoder module 420 extracts syntax elements from the bitstream 133 by decoding a sequence of "bins" and passes the values of those syntax elements to other modules within the video decoder 134. An example of a syntax element extracted from the bitstream 133 is the quantized coefficient 424. The entropy decoder module 420 uses an arithmetic decoding engine to decode each syntax element as a sequence of one or more bins. Each bin can use one or more "contexts" along with a context that describes the probability level used to encode the "1" and "0" values of the bin. When multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the contexts available for decoding the bin. The process of decoding the bins forms a sequential feedback loop. The number of operations in the feedback loop is preferably minimized to enable the entropy decoder 420 to achieve a high throughput in bins per second. Context modeling depends on other properties of the bitstream known to the video decoder 134 when selecting a context, i.e., the properties of the bins preceding the current bin. For example, the context can be selected based on the quadtree depth of the current CU within the coding tree. The dependencies are preferably based on properties known prior to decoding the bins or are determined without requiring long sequential processing.

[0082] The quantized coefficient 424 is input to the inverse quantizer module 428. The inverse quantizer module 428 performs inverse quantization (or "scaling") on the quantized coefficient 424 according to the quantization parameter to generate the reconstructed intermediate transform coefficient represented by arrow 432. If the use of a non-uniform inverse quantization matrix is indicated in the bitstream 133, the video decoder 134 reads the quantization matrix from the bitstream 133 as a sequence of scaling factors and arranges the scaling factors in the matrix. Inverse scaling uses the quantization matrix in combination with the quantization parameter to generate the reconstructed intermediate transform coefficient 432. The reconstructed intermediate transform coefficient 432 is passed to the inverse secondary transform module 436, where the secondary transform can be applied according to the decoded "nsst_index" syntax element. "nsst_index" is decoded from the bitstream 133 by the entropy decoder 420 under the execution of the processor 205. The inverse secondary transform module 436 generates the reconstructed transform coefficient 440.

[0083] The reconstructed transform coefficient 440 is passed to the inverse primary transform module 444. Module 444 converts the coefficient back from the frequency domain to the spatial domain. The result of the operation of module 444 is a block of residual samples represented by arrow 448. The block of residual samples 448 is equal in size to the corresponding CU. The type of inverse primary transform can be type II discrete cosine transform (DCT-2), type VII discrete sine transform (DST-7), type VIII discrete cosine transform (DCT-8), or "transform skip" mode. The use of the transform skip mode is signaled by a transform skip flag decoded from the bitstream 133 or inferred in some other way. When the transform skip mode is used, the residual sample 448 is the same as the reconstructed transform coefficient 440.

[0084] The residual sample 448 is supplied to an addition module 450. In the addition module 450, the residual sample 448 is added to the decoded PB (represented as 452) to generate a block of reconstructed samples represented by an arrow 456. The reconstructed samples 456 are supplied to a reconstructed sample cache 460 and an in-loop filtering module 488. The in-loop filtering module 488 generates a reconstructed block of frame samples represented as 492. The frame samples 492 are written to a frame buffer 496.

[0085] The reconstructed sample cache 460 operates in the same manner as the reconstructed sample cache 356 of the video encoder 114. The reconstructed sample cache 460 provides storage for the reconstructed samples necessary for intra prediction of subsequent CBs without memory 206 (e.g., typically by using data 232 which is on-chip memory instead). The reference samples represented by an arrow 464 are obtained from the reconstructed sample cache 460 and supplied to a reference sample filter 468 to generate filtered reference samples indicated by an arrow 472. The filtered reference samples 472 are supplied to an intra-frame prediction module 476. The module 476 generates a block of intra prediction samples represented by an arrow 480 according to an intra prediction mode parameter 458 signaled in the bitstream 133 and decoded by the entropy decoder 420.

[0086] If it is shown that the prediction mode of CB is an intra prediction in the bitstream 133, the intra prediction sample 480 forms the decoded PB452 via the multiplexer module 484. Intra prediction generates a block within one color component that is derived using the predicted block (PB) of samples, i.e., the "adjacent samples" within the same color component. The adjacent samples are samples adjacent to the current block and have already been reconstructed because they precede in the block decoding order. When luma and chroma blocks are juxtaposed, the luma and chroma blocks can use different intra prediction modes. However, the two chroma channels each share the same intra prediction mode.

[0087] Intra prediction of luma blocks consists of four types. "DC intra prediction" involves populating the PB with a single value representing the average of adjacent samples. "Planar intra prediction" involves populating the PB with samples according to a plane using vertical and horizontal gradients and a DC offset derived from adjacent samples. "Angular intra prediction" involves populating the PB with adjacent samples that are filtered and propagated in a specific direction (or "angle") across the PB. In VVC, the PB can be selected from up to 65 angles, and rectangular blocks can utilize different angles that are not available for square blocks. "Matrix intra prediction" involves populating the PB by multiplying a reduced set of adjacent samples by one of a number of available matrices that are available to the video decoder 134. The reduced set of adjacent samples is generated by filtering and subsampling the adjacent samples. Next, a set of reduced prediction samples is generated by multiplying the set of reduced samples by a matrix and adding an offset vector. The matrix and associated offset vector are selected from a number of possible matrices according to the size of the PB, and a specific selection of the matrix and offset vector is indicated by the "MIP mode" syntax element. For example, there are 11 MTP modes for PBs with a size larger than 8×8, while there are 19 MIP modes for 8×8 sized PBs. Finally, the PB generated by matrix intra prediction is populated from a reduced set of predicted samples by interpolation.

[0088] A fifth type of intra prediction is available for chroma PBs, whereby the PB is generated from collocated luma reconstruction samples according to the "Cross Component Linear Model" (CCLM) mode. Three different CCLM modes are available, each of which uses a different model derived from adjacent luma and chroma samples. The derived model is then used to generate a block of samples for the chroma PB from the collocated luma samples.

[0089] When it is shown that the prediction mode of CB is inter prediction in the bitstream 133, the motion compensation module 434 uses the motion vector and the reference frame index to select and filter a block of samples 498 from the frame buffer 496 to generate a block of inter prediction samples represented as 438. The block of samples 498 is obtained from a previously decoded frame stored in the frame buffer 496. In the case of bi-prediction, two sample blocks are generated and blended together to generate samples for the decoded PB 452. Filtered block data 492 from the in-loop filtering module 488 is input to the frame buffer 496. Similar to the in-loop filtering module 368 of the video encoder 114, the in-loop filtering module 488 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vector is applied to both the luma channel and the chroma channel, but the filtering processes for the sub-sampled interpolation luma channel and the chroma channel are different. The frame buffer 496 outputs the decoded video samples 135.

[0090] Figure 5 is a schematic block diagram showing a set 500 of available divisions or splits of a region into one or more sub-regions within the tree structure of general video coding. The divisions shown in the set 500 are available to the block partitioner 310 of the encoder 114 to divide each CTU into one or more CUs or CBs according to the coding tree, as determined by Lagrange optimization, as described with reference to FIG. 3.

[0091] The set 500 indicates that only the square region is divided into other, perhaps non-square, sub-regions. Although FIG. 500 shows a potential division, it should be understood that the enclosing region need not be square. If the enclosing region is non-square, the dimensions of the blocks resulting from the division are scaled according to the aspect ratio of the enclosing block. When the region is no longer divisible, i.e., at the leaf nodes of the coding tree, the CU occupies that region. The specific sub-division of a CTU into one or more CUs by the block partitioner 310 is called the "coding tree" of the CTU.

[0092] The process of sub-dividing a region into sub-regions must end when the resulting sub-regions reach the minimum CU size. In addition to constraining the CUs such that a given minimum size, e.g., a block region smaller than 16 samples, is prohibited, the CUs are constrained to have a minimum width or height of 4. Other minimum values are possible, both with respect to both width and height, or with respect to width or height. The process of sub-division can also end before the deepest level of decomposition, such that the CU is larger than the minimum CU size. It is possible that no division occurs, and as a result, a single CU occupies the entire CTU. A single CU that occupies the entire CTU is the largest available coding unit size. The use of a sub-sampled chroma format such as 4:2:0 allows the configuration of the video encoder 114 and the video decoder 134 to end the division of regions in the chroma channel earlier than in the luma channel.

[0093] At the leaf nodes of the coding tree, there are CUs without further sub - divisions. For example, leaf node 510 contains one CU. At the non - leaf nodes of the coding tree, there are divisions into two or more further nodes, each of which can be a leaf node forming one CU or a non - leaf node containing further divisions into smaller regions. At each leaf node of the coding tree, there is one coding block for each color channel. Divisions that terminate at the same depth for both luma and chroma result in three juxtaposed CBs. Divisions that terminate at a luma depth deeper than chroma will result in multiple luma CBs being juxtaposed with the CBs of the chroma channels.

[0094] As shown in FIG. 5, the quadtree split 512 divides the encompassing region into four regions of equal size. Compared with HEVC, Versatile Video Coding (VVC) achieves further flexibility by adding horizontal binary splits 514 and vertical binary splits 516. Each of the splits 514 and 516 divides the encompassing region into two regions of equal size. The splits are along the horizontal boundary (514) or vertical boundary (516) within the encompassing block.

[0095] Further flexibility is achieved in Versatile Video Coding by adding horizontal ternary splits 518 and vertical ternary splits 520. The ternary splits 518 and 520 divide the block into three regions bounded either horizontally (518) or vertically (520) along the 1 / 4 and 3 / 4 of the width or height of the encompassing region. The combination of quadtree, binary tree, and ternary tree is called "QTBTTT". At the root of the tree, there are zero or more quadtree splits (the "QT" section of the tree). When the QT section ends, zero or more binary or ternary splits (the "multi - tree" or "MT" section of the tree) occur, finally ending at the CBs or CUs of the leaf nodes of the tree. When the tree describes all color channels, the tree leaf nodes are CUs. When the tree describes the luma channel or chroma channel, the tree leaf nodes are CBs.

[0096] Supporting only quad-trees and thus only square blocks, QTBTTT results in more possible CU sizes compared to HEVC, especially considering the possible recursive application of binary and / or ternary tree partitions. The possibility of non-regular (non-square) block sizes can be reduced by constraining the partitioning options to exclude partitions where either the block width or height is less than 4 samples or not a multiple of 4 samples. Generally, this constraint is applied when considering luma samples. However, in the described configuration, the constraint can be applied separately to blocks for the chroma channels. The application of the constraint to the partitioning options for the chroma channels can result in different minimum block sizes for luma and chroma, for example, when the frame data is in 4:2:0 chroma format or 4:2:2 chroma format. For each partition, sub-regions are generated that are of the same side dimension, halved, or quartered with respect to this encompassing region. And since the CTU size is a power of 2, all CU side dimensions are also powers of 2.

[0097] Figure 6 is a schematic flow diagram showing the data flow 600 of a QTBTTT (or "coding tree") structure used in general video coding. The QTBTTT structure is used for each CTU to define the partitioning of the CTU into one or more CUs. The QTBTTT structure for each CTU is determined by the block partitioner 310 within the video encoder 114 and is either encoded into the bitstream 115 or decoded from the bitstream 133 by the entropy decoder 420 within the video decoder 134. The data flow 600 further characterizes the acceptable combinations available to the block partitioner 310 for partitioning the CTU into one or more CUs according to the partitions shown in Figure 5.

[0098] Starting from the top level of the hierarchy, i.e., the CTU, zero or more quadtree partitions are first performed. Specifically, the quadtree (QT) partition decision 610 is made by the block partitioner 310. The decision at 610 to return a "1" symbol indicates a decision to split the current node into four sub-nodes according to the quadtree partition 512. As a result, four new nodes, such as 620, are generated, and for each new node, the QT partition decision 610 is revisited. Each new node is considered in raster (or Z-scan) order. Alternatively, if the QT partition decision 610 indicates that no further split should be performed (returns a "0" symbol), the quadtree partition stops and the multi-tree (MT) partition is then considered.

[0099] First, the MT partition decision 612 is made by the block partitioner 310. At 612, a decision to perform the MT partition is indicated. Returning a "0" symbol at decision 612 indicates that no further split to the sub-nodes of the node is to be performed. If no further split of the node is performed, the node is a leaf node of the coding tree and corresponds to a CU. The leaf node is output at 622. Alternatively, if the MT partition 612 indicates a decision to perform the MT partition (returns a "1" symbol), the block partitioner 310 proceeds to the direction decision 614.

[0100] The direction decision 614 indicates the direction of the MT partition as either horizontal ("H" or "0") or vertical ("V" or "1"). The block partitioner 310 proceeds to decision 616 if decision 614 returns a "0" indicating the horizontal direction. The block partitioner 310 proceeds to decision 618 if decision 614 returns a "1" indicating the vertical direction.

[0101] In each of decisions 616 and 618, the number of partitions for MT splitting is shown as either two (2-way split or "BT" node) or three (3-way split or "TT") for the BT / TT split. That is, the BT / TT split decision 616 is made by the block partitioner 310 when the indicated direction from 614 is horizontal, and the BT / TT split decision 618 is made by the block partitioner 310 when the indicated direction from 614 is vertical.

[0102] The BT / TT split decision 616 indicates whether it is a 2-way split 514 indicated by the horizontal split returning "0" or a 3-way split 518 indicated by returning "1". When the BT / TT split decision 616 indicates a 2-way split, in the HBT CTU node generation step 625, two nodes are generated by the block partitioner 310 according to the horizontal 2-way split 514. When the BT / TT split 616 indicates a 3-way split, in the HTT CTU node generation step 626, three nodes are generated by the block partitioner 310 according to the horizontal 3-way split 518.

[0103] The BT / TT split decision 618 indicates whether it is a 2-way split 516 indicated by the vertical split returning "0" or a 3-way split 520 indicated by returning "1". When the BT / TT split 618 indicates a 2-way split, in the VBT CTU node generation step 627, two nodes are generated by the block partitioner 310 according to the vertical 2-way split 516. When the BT / TT split 618 indicates a 3-way split, in the VTT CTU node generation step 628, three nodes are generated by the block partitioner 310 according to the vertical 3-way split 520. For each node resulting from steps 625 - 628, the recursion of the data flow 600 back to the MT split decision 612 is applied in the order from left to right or from top to bottom according to the direction 614. As a result, binary and ternary tree splits can be applied to generate CUs of various sizes.

[0104] Figures 7A and 7B provide a splitting example 700 for some CUs or CBs of CTU 710. An example of CU 712 is shown in Figure 7A. Figure 7A shows the spatial arrangement of the CUs in CTU 710. The splitting example 700 is also shown as a coding tree 720 in Figure 7B.

[0105] In each non-leaf node within CTU 710 of Figure 7A, such as nodes 714, 716, and 718, the contained nodes (which may be further split or may be CUs) are scanned or traversed in "Z-order" to create a list of nodes, and are represented as columns within coding tree 720. In the case of quadtree splitting, the Z-order scan is in the order from top left to right followed by bottom left to right. In the case of horizontal and vertical splitting, the Z-order scan (traversal) simplifies to a scan from top to bottom and a scan from left to right, respectively. The coding tree 720 in Figure 7B lists all the nodes and CUs according to the applied scan order. Each split generates a list of 2, 3, or 4 new nodes at the next level of the tree until a leaf node (CU) is reached.

[0106] When an image is decomposed into CTUs and further into CUs by block partitioner 310, and each residual block (324) is generated using the CUs as described with reference to Figure 3, the residual blocks are forward-transformed by video encoder 114. An equivalent inverse transformation process is performed in video decoder 134 to obtain the TBs from bitstream 133.

[0107] In video encoder 114, the quantized coefficients 336 can be rearranged into a one-dimensional list by performing a 2-level back diagonal scan. Similarly, in video decoder 134, the quantized coefficients 424 can be rearranged from a one-dimensional list into a two-dimensional collection of sub-blocks by the same 2-level back diagonal scan.

[0108] FIG. 8A shows an exemplary two-level backward diagonal scan 810 of an 8×8 TB800. Scan 810 is shown to proceed from the bottom-right residual coefficient position of TB800 back to the top-left (DC) residual coefficient position of TB800. The path of scan 810 proceeds from one 4×4 region, known as a sub-block, to the next sub-block. For a TB with a width or height of 2, sub-block sizes of 2×2, 2×8, or 8×2 are available. The scan within a particular sub-block is performed according to an “encoded sub-block flag” or the sub-block is skipped. When the scan of a sub-block is skipped, all residual coefficients within the sub-block are assumed to have a value of zero. Although scan 810 is shown to start from the bottom-right residual coefficient position of TB800, for a given set of residual coefficients, the scan starts from the position of the “last significant coefficient” and is considered the “last” coefficient when the order of the coefficients proceeds from the DC coefficient instead of the scan order.

[0109] FIG. 8B shows an exemplary alternative two-level forward diagonal scan 860 of an 8×8 TB850 that is used when the TSRC process is selected. When the TSRC process is used in video encoder 114, the quantized coefficients 336 are rearranged into a one-dimensional list by scan 860. Similarly, when the TSRC process is used for the current TB within video decoder 134, the quantized coefficients 424 are rearranged from a one-dimensional list into a two-dimensional collection of sub-blocks by scan 860. Scan 860 is shown to proceed from the top-left (DC) residual coefficient position of TB850 to the bottom-right residual coefficient position of TB850. Unlike scan 810, scan 860 does not end at the “last significant coefficient”.

[0110] FIGS. 8A and 8B show scan patterns typically used in VVC. The examples described herein use scan pattern 810 to encode the residual coefficients transformed by module 326, and scan pattern 860 is used for the transformed blocks with transform skipped. However, in some implementations, other scan patterns can be used.

[0111] As described above, the transform coefficient 332 is the same as the residual coefficient 324 when the transform skip mode is used. Thus, regardless of whether the transform skip mode is selected, the transform coefficient 332 may be referred to in the same way as the residual coefficient. When reversible coding is desired, the video encoder 114 selects a transform skip for the current TB and signals a transform skip flag having a value of "TRUE" to the bitstream 133. The residual coefficient 332 associated with the current TB is encoded into the bitstream 133. Two residual coding processes, a "regular residual coding" (RRC) process and a "transform skip residual coding" (TSRC) process are available. In the normal operation of the video encoder 114, when a transform skip is selected (the transform skip flag has a value of "TRUE"), the TSRC process is selected, and when not (the transform skip flag has a value of "FALSE"), the RRC process is selected. However, it is typically not desirable for the encoding of the residual coefficient 332 to be exclusively processed by the TSRC in the case of reversible coding.

[0112] In one configuration of the video encoder 114, the TSRC disable flag is signaled within the bitstream 133. The TSRC disable flag can be signaled at a relatively high level, such as once per sequence or once per picture, so that the relative cost of signaling the TSRC disable flag is low. High-level syntax elements are typically grouped into parameter sets such as "sequence parameter set" (SPS) for sequence-level flags and "picture parameter set" (PPS) for parameter-level flags. The TSRC disable flag may be set to "TRUE" if the video data 113 is thought to be unsuitable for encoding by TSRC processing (in terms of encoding loss and function reproduction). An example of video data unsuitable for encoding by TSRC is natural scene content. The TSRC disable flag may be set to "FALSE" when the video data 113 is thought to belong to a class that can be successfully encoded by TSRC processing. Video data suitable for encoding by TSRC processing includes artificial screen content.

[0113] If the video encoder 114 selects a transform skip for the current TB and the TSRC disable flag is set to "TRUE", the residual coefficients 332 are encoded into the bitstream 133 using the RRC process. Similarly, if the video decoder 134 determines that a transform skip is used for the current TB and the TSRC disable flag is set to "TRUE", the residual coefficients 432 are decoded from the bitstream 133 using the RRC process.

[0114] FIG. 9 shows a method 900 for encoding a transform block of residual coefficients 332 using the RRC process. The method 900 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, the method 900 may be executed by the video encoder 114 under the execution of the processor 205. Thus, the method 900 can be implemented as a module of software 233 stored in a computer-readable storage medium and / or memory 206.

[0115] Method 900 is executed in several configurations by video encoder 114, upon receiving residual coefficient 332, first by quantizer 334 and then by entropy encoder 338. Method 900 begins with coefficient quantization step 910.

[0116] In coefficient quantization step 910, step 910 invokes method 1100, which is described below in connection with FIG. 11. Method 1100 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1100 may be executed by video encoder 114 under the execution of processor 205. Thus, method 1100 can be implemented as a module of software 233 stored in computer-readable storage medium and / or memory 206. Method 1100 quantizes residual coefficient 332 and generates quantized coefficient 336. Method 900 proceeds from step 910 to end position encoding step 920 under the control of processor 205.

[0117] In end position encoding step 920, video encoder 114 finds the position of the last significant coefficient within quantized coefficient 336 for the transform block of residual coefficient 332). The last significant coefficient is determined in relation to the forward direction of an appropriate scan pattern, e.g., the direction of two-level forward diagonal scan 860. A quantized coefficient is significant if it has any value other than zero. The position of the last valid coefficient is written to bitstream 133. Method 900 proceeds from step 920 to state initialization step 930 under the control of processor 205.

[0118] In state initialization step 930, quantizer state Qstate is set to the value zero. Furthermore, in step 930, a sub-block containing the last valid coefficient is selected. Method 900 proceeds from step 930 to encoded sub-block flag determination step 940 under the control of processor 205.

[0119] The description herein refers to several flags being either "TRUE" or "FALSE". Setting to "TRUE" means that the flag value indicates that the requirement is met or the mode is selected. Setting to "FALSE" means that the flag value indicates that the requirement is not met or the mode is not selected.

[0120] In the encoded sub-block flag determination step 940, the video encoder 114 determines and sets the encoded sub-block flag. If the currently selected sub-block is the first sub-block selected in the state initialization step 930, the encoded sub-block flag is set to "TRUE", but is not encoded into the bitstream 133. If the currently selected sub-block is identified as the last sub-block, as will be described later in connection with the last sub-block test 970, the encoded sub-block flag is set to "TRUE", but is not encoded into the bitstream 133.

[0121] Otherwise, the video encoder 114 sets the encoded sub-block flag to (i) "TRUE" if there is at least one significant coefficient in the 4×4 quantization coefficients belonging to the selected sub-block, or (ii) "FALSE" if there are no significant coefficients, and encodes the encoded sub-block flag into the bitstream 133. The method 900 proceeds from step 940 to the encoded sub-block flag test step 950 under the control of the processor 205.

[0122] In the encoded sub-block flag test step 950, the method 900 determines the value of the encoded sub-block flag. If the encoded sub-block flag is set to "TRUE", the method 900 proceeds to the sub-block encoding step 960. Otherwise, if the encoded sub-block flag is set to "FALSE", the method 900 proceeds to the last sub-block test step 970.

[0123] In sub-block encoding step 960, entropy encoder 338 encodes the quantized coefficients within the selected sub-block into bitstream 133. Step 960 invokes method 1300 described below in connection with FIG. 13. Method 900 proceeds from step 960 to the last sub-block test 970 under the control of the processor.

[0124] In the last sub-block test 970, method 900 operates to determine whether the selected sub-block is the last sub-block in the current transform block. If the currently selected sub-block is the top-left sub-block of the transform block, step 900 returns "YES" and method 900 ends. Otherwise, if the currently selected sub-block is not the top-left sub-block of the transform block, step 970 returns "NO" and method 900 proceeds to step 980 where the next sub-block is selected.

[0125] In step 980 of selecting the next sub-block, the next sub-block within the transform block is selected. The next sub-block in the backward diagonal scan order 810 is selected. Method 900 proceeds from step 980 to step 940 of determining the encoded sub-block flag for the selected sub-block.

[0126] FIG. 10 shows a method 1000 for decoding a transform block of residual coefficients 432 by the RRC process. Method 1000 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, method 1000 may be executed by video decoder 134 under the execution of processor 205. Thus, method 1000 can be implemented as a module of software 233 stored in a computer-readable storage medium and / or memory 206.

[0127] Method 1000 is implemented in several configurations in entropy decoder 420 by video encoder 134 upon receipt of bitstream 133 and in inverse quantization module 428. Method 1000 begins with last position decoding step 1010.

[0128] In last position decoding step 1010, the last significant coefficient position of the transform block of residual coefficients 432 is decoded from bitstream 133. Method 1000 proceeds from step 1010 to state initialization step 1020 under the control of processor 205.

[0129] In state initialization step 1020, video decoder 134 initializes the quantizer state Qstate to the value 0. Further, in step 1020, a sub-block including the last significant coefficient position is selected. Method 1000 proceeds from step 1020 to encoded sub-block flag determination step 1030 under the control of processor 205.

[0130] In encoded sub-block flag determination step 1030, video decoder 134 determines the encoded sub-block flag. If the currently selected sub-block is the first sub-block selected in state initialization step 1020, the encoded sub-block flag is set to "TRUE" (i.e., the encoded sub-block flag is presumed to be "TRUE"). If the currently selected sub-block is identified as the last sub-block as described below in last sub-block test 1060, the encoded sub-block flag is inferred as "TRUE". Otherwise, video decoder 134 decodes the encoded sub-block flag from bitstream 133. Method 1000 proceeds from step 1030 to encoded sub-block flag test 1040 under the control of processor 205.

[0131] In symbolic sub-block flag test 1040, method 1000 tests the value of the symbolic sub-block flag determined in step 1030. If it is determined in step 1040 that the symbolic sub-block flag has a value of "TRUE", method 1000 proceeds to sub-block decoding step 1050. Otherwise, in step 1040, if it is determined that the symbolic sub-block flag has a value of "FALSE", a zero value is assigned to all of the quantized coefficients within the currently selected sub-block, and method 1000 proceeds to the last sub-block test 1060.

[0132] In sub-block decoding step 1050, entropy decoder 420 decodes the quantized coefficients of the sub-block selected from bitstream 133. Step 1050 invokes method 1400, which will be described below in connection with FIG. 14. Method 1000 proceeds to the last sub-block test 1060 under the control of processor 205.

[0133] In the last sub-block test 1060, if the currently selected sub-block is the top-left sub-block of the transform block, step 1060 returns "YES" and method 1000 proceeds to coefficient scale step 1080. Otherwise, step 1060 returns "NO" and method 1000 proceeds to step 1070 of selecting the next sub-block.

[0134] In step 1070 of selecting the next sub-block, the next sub-block in the reverse diagonal scan order 810 is selected. Method 1000 proceeds from step 1070 to symbolic sub-block flag determination step 1030 under the control of processor 205.

[0135] In coefficient scale step 1080, the inverse quantizer module 428 applies scaling to the quantized coefficients 424 to generate the reconstructed residual coefficients 432. The sub-block is decoded by reconstructing the residual coefficients of the sub-block using the decoded coded bits. Step 1080 calls method 1200 described below in connection with FIG. 12. Method 1000 ends with the execution of step 1080.

[0136] FIG. 11 shows a method 1100 for quantizing the residual coefficients 332 of the transform block to generate quantized coefficients 336. Method 1100 is performed on the TB in step 910 of method 900. Method 1100 begins with a DQ test 1110.

[0137] In DQ test 1110, the video encoder 114 determines whether dependent quantization is used to quantize the residual coefficients 332. The video encoder 114 checks the value of the valid dependent quantization flag, which is signaled as high-level syntax in the bitstream 133. The valid dependent quantization flag determines whether dependent quantization is permitted within the range of the flag. For example, the sequence-level dependent quantization flag determines whether dependent quantization is permitted when encoding the entire video sequence. The picture-level dependent quantization flag determines whether dependent quantization is permitted when encoding the current picture and takes precedence over the value of the sequence-level dependent quantization flag. If the valid dependent quantization flag is "FALSE", step 1110 returns "NO" and method 1100 proceeds to the scalar quantization step 1120.

[0138] In one configuration of DQ test 1110, when the valid dependent quantization flag is "TRUE", video encoder 114 also checks the value of the conversion skip flag for the current TB. If the valid dependent quantization flag is "TRUE" and the conversion skip flag is "TRUE", step 1110 returns "NO" and method 1100 proceeds to scalar quantization step 1120. Otherwise, if the valid dependent quantization flag is "TRUE" and the conversion skip flag is "FALSE", step 1110 returns "YES" and method 1100 proceeds to dependent quantization step 1130.

[0139] In another configuration of DQ test 1110, when the valid dependent quantization flag is "TRUE", video encoder 114 also checks the value of the TSRC invalid flag. If the valid dependent quantization flag is "TRUE" and the TSRC invalid flag is "TRUE", step 1110 returns "NO" and method 1100 proceeds to scalar quantization step 1120. Otherwise, if the valid dependent quantization flag is "TRUE" and the TSRC invalid flag is "FALSE", step 1110 returns "YES" and method 1100 proceeds to dependent quantization step 1130.

[0140] In yet another configuration of DQ test 1110, when the valid dependent quantization flag is "TRUE", video encoder 114 also checks the value of the quantization parameter (QP) of the current TB. QP indicates the degree of quantization applied to the residual coefficients 332. QP is determined as QP = QP i + 6*(BD - 8), from an initial QP i and an offset that depends on the bit depth BD of video encoder 114. For example, if QP i is 4 and the bit depth is 8, QP is determined to be 4. QP iWhen it is -8 and the bit depth is 10, the QP is determined to be 4. Typically, a QP of 4 indicates that the residual coefficients are not quantized, and thus, reversible operations are possible. However, higher values of QP can still achieve reversible operations. For example, if video data 113 was originally captured at a bit depth of 8 but is supplied to video encoder 114 at a higher bit depth, reversible operations are possible at a higher QP. For example, if video data 113 was captured at bit depth 8 but is supplied to video encoder 114 at bit depth 10, reversible operations are possible at QPs of 4, 10, or 16. The QP at which reversible operations are possible can be indicated by the minimum QP for transform skip blocks notified in the high-level syntax parameter set. If the valid dependent quantization flag is "TRUE" and the QP is 4 (or any value indicating reversible operations), step 1110 returns "NO" and method 1100 proceeds to scalar quantization step 1120. Otherwise, if the valid dependent quantization flag is "TRUE" and the QP is not 4 (or a similar value indicating reversible operations), step 1110 returns "YES" and method 1100 proceeds to dependent quantization step 1130. In scalar quantization step 1120, the residual coefficient 332 is represented as r[n]. Then, a quantized coefficient q[n] is generated by quantizing the residual coefficient r[n] according to the following equation (1). q[n]=(k*r[n]+offset)>>qbits (1) In equation (1), k is the scaling factor, qbits is the coarse quantization factor, and offset controls the placement of the quantization threshold. k, qbits, and offset are determined based on the values of the quantization parameters of the current TB. For example, when the QP is 4, k = 1, qbits = 0, and offset = 0. Then, when the QP is 4, q[n]=r[n], and no loss occurs in the scalar quantization step. Method 1100 proceeds from step 1120 to SBH test 1140 under the control of processor 205.

[0141] In the dependent quantization step 1130, each of the residual coefficients r[n] can be quantized by one of a plurality of scalar quantizers. For the same QP, the scalar quantizers have the same quantization division number, but the quantization thresholds are offset relative to each other. The scalar quantizer for a particular residual coefficient r[n] depends on the current quantizer state Qstate updated for each coefficient and depends on the parity (least significant bit) of the resulting q[n]. Due to the dependence on the previous state, the optimal quantization result is not determined for each coefficient. One efficient way to determine the optimal quantization result is by constructing a "trellis" of possible quantization states at each coefficient position. The optimal quantization result can be found by equivalently finding the best path through the trellis. The optimal trellis path can be determined by applying the Viterbi algorithm. Method 1100 ends with the execution of step 1130.

[0142] In the SBH test 1140, the video encoder 114 determines whether sign bit hiding is used to modify the quantized coefficient q[n] before encoding the coefficients of the TB. The video encoder 114 checks the value of the valid sign bit hiding flag. The valid sign bit hiding flag is signaled as high-level syntax in the bitstream 133. For example, the valid sign bit hiding flag can be signaled in the picture header. If the valid dependent quantization flag is "TRUE", the valid sign bit hiding flag is implicitly "FALSE". If the valid sign bit hiding flag is "FALSE", step 1140 returns "NO" and method 1100 ends.

[0143] In one configuration of the SBH test 1140, the determination depends on the value of the valid sign bit hiding flag and the value of the current TB's conversion skip flag. If the valid sign bit hiding flag is "TRUE", the video encoder 114 also checks the value of the conversion skip flag for the current TB. If the valid sign bit hiding flag is "TRUE" and the conversion skip flag is "TRUE", step 1140 returns "NO" and method 1100 ends. Otherwise, if the valid sign bit hiding flag is "TRUE" and the conversion skip flag is "FALSE", step 1140 returns "YES" and method 1100 proceeds to the parity adjustment step 1150.

[0144] In another configuration of the SBH test 1140, the determination depends on the value of the valid sign bit hiding flag and the value of the TSRC invalid flag. If the valid sign bit hiding flag is "TRUE", the video encoder 114 also checks the value of the TSRC invalid flag. If the valid sign bit hiding flag is "TRUE" and the TSRC invalid flag is "TRUE", step 1140 returns "NO" and method 1100 ends. If not, when the valid sign bit hiding flag is "TRUE" and the TSRC invalid flag is "FALSE", step 1140 returns "YES" and method 1100 proceeds to the parity adjustment step 1150.

[0145] In yet another configuration of the SBH test 1140, the determination depends on the value of the valid sign bit hiding flag and the value of the quantization parameter (QP) of the current TB. If the valid sign bit hiding flag is "TRUE", the video encoder 114 also checks the value of QP for the current TB. If the valid sign bit hiding flag is "TRUE" and the QP is 4 (or any value indicating a reversible operation), step 1140 returns "NO" and method 1100 ends. Otherwise, if the valid sign bit hiding flag is "TRUE" and the QP is not 4 (or a similar value indicating a reversible operation), step 1140 returns "YES" and method 1100 proceeds to the parity adjustment step 1150.

[0146] In the parity adjustment step 1150, the video encoder 114 checks the positions of the first and last significant coefficients for each sub-block within the current TB. If the difference between the first and last significant positions of a sub-block is greater than a threshold (typically 3), sign bit hiding is used for that sub-block. For each sub-block for which sign bit hiding is used, the video encoder 114 checks the sign of the first significant coefficient within the sub-block and adjusts the parity of the coefficients within the sub-block accordingly. The parity of a coefficient is zero if the coefficient is even and 1 if the coefficient is odd. The sum of the parities of multiple coefficients is zero if the number of odd coefficients is odd and 1 if the number of odd coefficients is even. If the sign of the first significant coefficient within the sub-block is positive, the coefficients within the sub-block are adjusted so that the sum of the parities is zero. If the sign of the first significant coefficient within the sub-block is negative, the coefficients within the sub-block are adjusted so that the sum of the parities is 1. Method 1100 ends after the execution of step 1150.

[0147] FIG. 12 illustrates a method 1200 for applying scaling to quantization coefficients 424 to generate reconstructed residual coefficients 432. Method 1200 may be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, method 1200 may be executed by video decoder 134 under the execution of processor 205. Thus, method 1200 can be implemented as a module of software 233 stored in computer-readable storage medium and / or memory 206. Method 1200 is implemented in step 1080 of method 1000. Method 1200 begins with DQ test 1210.

[0148] In DQ test 1210, video decoder 134 determines whether dependent quantization is used to inverse-quantize quantization coefficients 424. Video decoder 134 checks the value of the valid dependent quantization flag, which may be decoded from bitstream 133 or inferred based on the value of other high-level syntax flags. If the valid dependent quantization flag is "FALSE", step 1210 returns "NO" and method 1200 proceeds to inverse scalar quantization step 1220.

[0149] In one configuration of DQ test 1210, if the valid dependent quantization flag is "TRUE", video decoder 134 also checks the value of the current TB's transform skip flag. If the valid dependent quantization flag is "TRUE" and the transform skip flag is "TRUE", step 1210 returns "NO" and method 1100 proceeds to inverse scalar quantization step 1220. Otherwise, if the valid dependent quantization flag is "TRUE" and the transform skip flag is "FALSE", step 1210 returns "YES" and method 1100 proceeds to inverse dependent quantization step 1230.

[0150] In another configuration of the DQ test 1210, when the valid dependent quantization flag is "TRUE", the video decoder 134 also checks the value of the TSRC invalid flag. When the valid dependent quantization flag is "TRUE" and the TSRC invalid flag is "TRUE", step 1210 returns "NO" and method 1100 proceeds to the inverse scalar quantization step 1220. Otherwise, when the valid dependent quantization flag is "TRUE" and the TSRC invalid flag is "FALSE", step 1210 returns "YES" and method 1100 proceeds to the inverse dependent quantization step 1230.

[0151] In yet another configuration of the DQ test 1210, when the valid dependent quantization flag is "TRUE", the video decoder 134 also checks the value of the quantization parameter (QP) of the current TB. When the valid dependent quantization flag is "TRUE" and the QP is 4 (or any value indicating a reversible operation), step 1210 returns "NO" and method 1100 proceeds to the inverse scalar quantization step 1220. Otherwise, when the valid dependent quantization flag is "TRUE" and the QP is not 4 (or a similar value indicating a reversible operation), step 1210 returns "YES" and method 1100 proceeds to the inverse dependent quantization step 1230.

[0152] In the inverse scalar quantization step 1220, the video decoder 134 scales the quantization coefficient 424 to generate the reconstructed residual coefficient 432. The quantization coefficient 424 is represented as q[n]. The reconstructed residual coefficient r[n] is generated by scaling the quantization coefficient q[n] according to the following equation (2).

[0153]

Equation

[0154] In equation (2), s is a scaling factor determined based on the value of the QP for the current TB. For example, when the QP is 4, s = 1 and r[n]=q[n]. Method 1200 ends with the execution of step 1220.

[0155] In the inverse-dependent quantization step 1230, the video decoder 134 applies inverse-dependent quantization to the quantization coefficients 424 to generate the reconstructed residual coefficients 432. The quantizer state Qstate is initially reset to zero. The quantization coefficients 424 are represented by q[n]. Each coefficient position n is visited in the reverse diagonal scan order 810, and each reconstructed residual coefficient r[n] is calculated according to Equation (3).

[0156]

Equation

[0157] In Equation (3), s is a scaling factor determined based on the value of QP for the current TB.

[0158] After each reconstructed residual coefficient r[n] is calculated, the quantizer state is updated based on the parity of q[n] according to Table 1.

[0159]

Table 1

[0160] The method 1200 ends with the execution of step 1230.

[0161] To utilize the statistical characteristics of the quantization coefficients 336, the quantization coefficients are binarized into several syntax elements by the video encoder 114 (typically by the entropy encoder 338) before encoding. For example, since the quantization coefficients 336 often have a value of zero, one syntax element is a significant flag, which is set to "FALSE" for quantization coefficients with a value of zero. If the significant flag is set to "FALSE", no further syntax elements of the related quantization coefficients are signaled. The significant flag can be encoded into the bitstream 133 by using a context-adaptive binary arithmetic coding (CABAC) entropy encoder.

[0162] The CABAC coder encodes context - coding syntax elements relatively efficiently, but generally it is desirable to limit the number of context - coding syntax elements to minimize the computational requirements and cost for hardware implementation. Thus, after quantization coefficient 336 is binarized into several syntax elements by entropy encoder 338, some of the syntax elements are context - coded into bitstream 133 and other syntax elements are bypass - coded into bitstream 133. The total number of context - coding syntax element bins is limited per transform block. In the VVC standard, the limit is set to 1.75 bins per sample. For example, for an 8×8 transform block consisting of 64 samples, the context - coding bin budget is set to 112 bins. During the process of coding the TB into bitstream 133, whenever a syntax element is context - coded, the remaining context - coding bin budget is tracked and decremented. When the remaining context - coding bin budget is depleted, the remaining quantization coefficients and related syntax elements must be bypass - coded.

[0163] FIG. 13 shows a method 1300 for encoding the quantization coefficients (336) of the currently selected sub - block into bitstream 133. Method 1300 is implemented at step 960 of method 900. Method 1300 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, method 1300 can be executed by video encoder 114 under the execution of processor 205. Thus, method 1300 can be implemented as a module of software 233 stored in a computer - readable storage medium and / or memory 206. Method 1300 begins with step 1310 of selecting an initial coefficient.

[0164] In step 1310 of selecting the initial coefficient, method 1300 selects the quantization coefficient of the current sub-block. If the current sub-block contains the last significant coefficient position, the currently selected coefficient is set as the last significant coefficient. Otherwise, if the current sub-block does not contain the last significant coefficient position, the currently selected coefficient is set as the bottom-right coefficient of the current sub-block. Method 1300 proceeds to context coding usage check 1320.

[0165] In context coding usage check 1320, video encoder 114 checks whether the remaining context coding bin budget is 4 or more. If the remaining context coding bin budget is 4 or more, step 1320 returns "YES" and method 1300 proceeds to context coding syntax element encoding step 1330. Otherwise, if the current context coding bin budget is less than 4, step 1320 returns "NO" and method 1300 proceeds to remainder pass encoding step 1370.

[0166] In context coding syntax element encoding step 1330, video encoder 114 can encode a number of syntax elements into bitstream 133 using a CABAC encoder that includes a potentially significant flag, a flag greater than 1, a parity flag, and a flag greater than 3. Each bin associated with a syntax element is encoded by the CABAC encoder using a "context model". The context model for each bin can be selected according to the current value of the quantizer state Qstate. Further, whenever a context coding bin is encoded into bitstream 133 by the CABAC encoder, the remaining context coding bin budget is reduced by 1 in step 1330.

[0167] In step 1330, if the current coefficient is the last significant coefficient, the significance flag is set to "TRUE", but it is not encoded in the bitstream 133. If the current selected sub-block is not the first or last sub-block in the backward scan order 810, and the current selected coefficient is the last coefficient as described below in the last coefficient check 1350, and all significance flags for the previous coefficients in the current selected sub-block are "FALSE", the significance flag is set to "TRUE". The significance flag is not encoded in the bitstream 133. If the current coefficient has a magnitude of zero, in step 1330, the significance flag is set to "FALSE" and context-encoded in the bitstream 133. Otherwise, the significance flag is set to "TRUE" and context-encoded in the bitstream 133 in step 1330.

[0168] If the current coefficient has a magnitude of 1, the flag greater than 1 is set to "FALSE" and context-encoded in the bitstream 133 in step 1330. Otherwise, the flag greater than 1 is set to "TRUE" and context-encoded in the bitstream 133.

[0169] If the current coefficient has a magnitude of at least 2, the parity flag is set to "FALSE" if the current coefficient is even, and set to "TRUE" if the current coefficient is odd. The parity flag is context-encoded in the bitstream 133 in step 1330. If the current coefficient has a magnitude greater than 3, the flag greater than 3 is set to "TRUE" and context-encoded in the bitstream 133 in step 1330. Otherwise, if the current coefficient has a magnitude of 2 or 3, the flag greater than 3 is set to "FALSE" and context-encoded in the bitstream 133.

[0170] Method 1300 proceeds from step 1330 to DQ test 1340 under the control of processor 205. Depending on the coefficient selected at 1310, method 1300 sets (or in some cases encodes) a significance flag and then proceeds to step 1340. Otherwise, method 1300 encodes the last appropriate one of the flags greater than 1, the parity flag, and the flags greater than 3, and then proceeds to step 1340.

[0171] In DQ test 1340, using the same conditions checked in DQ test 1110, it is determined whether step 1340 returns "YES" or "NO". If step 1340 returns "YES", method 1300 proceeds to Qstate update step 1345. Otherwise, if step 1340 returns "NO", method 1300 proceeds to last coefficient check 1350.

[0172] In Qstate update step 1345, the quantizer state Qstate is updated based on the parity of the current coefficient according to Table 1. Method 1300 proceeds from step 1345 to last coefficient check 1350.

[0173] In last coefficient check 1350, video encoder 114 checks whether the currently selected coefficient is the top-left coefficient of the currently selected sub-block. If the currently selected coefficient is the top-left coefficient of the currently selected sub-block, step 1350 returns "YES" and method 1300 proceeds to reminder path encoding step 1370. Otherwise, if the current coefficient is not the top-left coefficient, step 1350 returns "NO" and method 1300 proceeds to step 1360 of selecting the next coefficient.

[0174] In step 1360 of selecting the next coefficient, the next coefficient of the currently selected sub-block is selected in the backward diagonal scan order 810. Method 1300 proceeds from step 1360 to context encoding usage check 1320.

[0175] In step 1370 of encoding the reminder path, the remaining magnitudes of the quantization coefficients of the currently selected sub-block are binarized and bypass-encoded into the bitstream 133, for example, by the entropy encoder 338. The quantization coefficients are encoded, for example, in the reverse diagonal scan order 810. If the quantization coefficients are context-encoded by the CABAC coder (i.e., the context encoding usage check 1320 has passed (returned "YES")) and the flag greater than 3 is "TRUE", the quantization coefficient at the scan position n has the remaining magnitude r[n]. The remaining magnitude is determined using equation (4). r[n]=(x[n]-4)>>1, (4)

[0176] Here, in equation (4), x[n] is the absolute magnitude of the quantization coefficient at the scan position n. The magnitude r[n] is binarized and bypass-encoded into the bitstream 133. If the quantization coefficients are not context-encoded (the context encoding usage check 1320 has not passed / returned "NO"), the absolute magnitude x[n] is binarized and bypass-encoded into the bitstream 133. The method 1300 proceeds from step 1370 to the SBH test 1380.

[0177] In the SBH test 1380, the same conditions checked in the SBH test 1140 are used to determine whether step 1380 returns "YES" or "NO". If the SBH test 1140 returns "NO", step 1380 returns "NO" and the method 1300 proceeds to step 1390 of encoding N codes. Otherwise, the video encoder 114 checks the positions of the first and last significant coefficients of the current sub-block. If the difference between the first significant position and the last significant position is greater than 3, step 1380 returns "YES" and the method 1300 proceeds to step 1395 of encoding N - 1 codes. Otherwise, step 1380 returns "NO" and the method 1300 proceeds to step 1390 of encoding N codes.

[0178] As described in connection with step 1140, the sign bit hiding test can depend on a number of alternative flags or settings in different implementations. If the valid sign bit hiding flag is set (has a "TRUE" value), different implementations can make a determination based on whether the TB conversion skip flag, the TSRC invalid flag, or the TB's QP meets a threshold related to reversible coding. Thus, step 1380 determines whether sign bit hiding is enabled depending on values or flags related to the conversion block itself, or upper level values of the TSRC invalid flag. Step 1380 provides a certain degree of flexibility for performing reversible coding. Implementations that use values or flags related to the conversion block to determine whether to enable sign hiding are particularly suitable for enabling flexibility when implementing reversible coding using RRC.

[0179] In step 1390 of encoding N symbols, the sign bits of any significant coefficients of the currently selected sub-block are bypass encoded into the bit stream 133. The sign bits are bypass encoded into the bit stream 133, for example, based on the backward diagonal scan order 810. Method 1300 ends after the execution of step 1390.

[0180] In step 1395 of encoding N - 1 symbols, the sign bits of the significant coefficients of the currently selected sub-block are bypass encoded into the bit stream 133 based on the backward diagonal scan order 810. The sign bits associated with the first significant coefficient (visited last in the backward diagonal scan order 810) are not encoded into the bit stream 133. In other words, if there are N significant coefficients in the currently selected sub-block, N - 1 sign bits are bypass encoded into the bit stream 133. Method 1300 ends with the execution of step 1395.

[0181] Figure 14 shows a method 1400 for decoding the quantization coefficients (424) of the currently selected sub-block from the bitstream 133. Method 1400 is implemented in step 1050 of method 1000. Method 1400 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, method 1400 may be executed by the video decoder 134 under the execution of the processor 205. Thus, method 1400 can be implemented as a module of software 233 stored in a computer-readable storage medium and / or memory 206. Method 1400 begins with step 1410 of selecting the first coefficient.

[0182] In step 1410 of selecting the first coefficient, method 1400 selects the first quantization coefficient of the current sub-block. If the current sub-block includes the last significant coefficient position, the currently selected coefficient is set to the last significant coefficient. Otherwise, the currently selected coefficient is set to the bottom-right coefficient of the current sub-block. Method 1400 proceeds from step 1410 to the context coding usage check step 1420.

[0183] In the context coding usage check 1420, the video decoder 134 checks whether the remaining context coding bin budget meets a threshold, typically whether the remaining context coding bin budget for the transform block is 4 bins or more. If the remaining budget is 4 or more, step 1420 returns "YES" and method 1400 proceeds to the context coding syntax element determination step 1430. Otherwise, if the remaining CABAC budget is less than the threshold (4 bins), step 1420 returns "NO" and method 1400 proceeds to the reminder path decoding step 1470.

[0184] In context - encoded syntax element determination step 1430, video decoder 134 can decode a number of context - encoded syntax elements from bitstream 133 using a CABAC coder. Each bin related to a syntax element is decoded by the CABAC coder using a "context model". The context model for each bin can be selected according to the current value of quantizer state Qstate. Further, whenever a context - encoded bin is decoded from bitstream 133 by the CABAC coder, the remaining context - encoded bin budget is decreased by 1.

[0185] If the current coefficient is the last significant coefficient, the significant flag is inferred as "TRUE" rather than being decoded from bitstream 133. If the current selected sub - block is not the first or the last sub - block in the backward scan order 810, the current selected coefficient is the last coefficient as described below in last coefficient check 1450, and all significant flags for previous coefficients in the current selected sub - block are "FALSE", the significant flag is inferred as "TRUE". Otherwise, the significant flag is context - decoded from bitstream 133 in step 1430. If the significant flag is set to "FALSE", a zero value is assigned to the currently selected coefficient and method 1400 proceeds to DQ test 1440.

[0186] If the significant flag is set to "TRUE", in step 1430, a flag greater than 1 is context - decoded from bitstream 133. If the flag greater than 1 is set to "FALSE", a magnitude of 1 is assigned to the currently selected coefficient and method 1400 proceeds to DQ test 1440.

[0187] If a flag greater than 1 is set to "TRUE", the parity flag and the flag greater than 3 are context-decoded from the bit stream 133. Method 1400 proceeds to the DQ test 1440. The number of flags determined in step 1430 depends on the position and value of the coefficient selected in step 1410. The progression from step 1430 can occur after a significant flag has been inferred or decoded, or after the appropriate one of the flag greater than 1, the parity flag, or the flag greater than 3 has been decoded.

[0188] In the DQ test 1440, using the same conditions checked in the DQ test 1210, it is determined whether step 1440 returns "YES" or "NO". If step 1440 returns "YES", method 1400 proceeds to the Qstate update step 1445. Otherwise, if step 1440 returns "NO", method 1400 proceeds to the last coefficient check 1450.

[0189] In the Qstate update step 1445, the quantizer state Qstate is updated based on the parity of the currently selected coefficient according to Table 1. If the currently selected coefficient has a value of zero, the parity is zero. The parity is 1 if the currently selected coefficient has a magnitude of 1. Otherwise, the parity is zero if the parity flag is set to "FALSE", and the parity is 1 if the parity flag is set to "TRUE". Method 1400 proceeds from step 1445 to the last coefficient check 1450.

[0190] In the final coefficient check step 1450, the video decoder 134 checks whether the currently selected coefficient is the coefficient at the upper left of the currently selected sub-block. If the currently selected coefficient is the coefficient at the upper left of the currently selected sub-block, step 1450 returns "YES" and method 1400 proceeds to the reminder path decoding step 1470. Otherwise, if the currently selected coefficient is not the upper left coefficient, step 1450 returns "NO" and method 1400 proceeds to step 1460 to select the next coefficient.

[0191] In step 1460 to select the next coefficient, in the reverse diagonal scan order 810, the next coefficient of the currently selected sub-block is selected. Method 1400 proceeds from step 1460 to the context coding usage check 1420.

[0192] In the reminder path decoding step 1470, the remaining magnitude of the quantization coefficient of the currently selected sub-block is bypass decoded from the bitstream 133. The quantization coefficients are processed in the reverse diagonal scan order 810. If a flag greater than 3 for which the quantized coefficient has been context decoded (the context coding usage check 1420 has passed or returned "YES") is decoded with a value of "TRUE", the remaining magnitude r[n] is bypass decoded from the bitstream 133, where n is the scan position of the quantization coefficient. The absolute magnitude x[n] of the quantization coefficient is determined as x[n]=4 + p[n]+2*r[n], where p[n] has a value of zero if the parity flag is decoded as "FALSE" and p[n] has a value of 1 if the parity flag is decoded as "TRUE".

[0193] The quantization coefficient is context decoded, and if the flag greater than 1 is decoded as "TRUE" but the flag greater than 3 is not decoded or is decoded as "FALSE", the absolute magnitude is determined as x[n]=2 + p[n]. If the quantization coefficient is not context decoded (the context coding usage check 1420 fails and returns "NO"), the absolute magnitude x[n] is bypass decoded from the bitstream 133. The method 1400 proceeds from step 1470 to the SBH test 1480.

[0194] In the SBH test 1480, the video decoder 134 determines whether sign bit hiding is being used, i.e., whether one sign bit for the currently selected sub-block is inferred. The test used in step 1480 relates to the test used in step 1140 on the encoder side. The video decoder 134 checks the value of the valid sign bit hiding flag, which may be signaled as high level syntax within the bitstream 133. If the valid dependent quantization flag is "TRUE", the valid sign bit hiding flag is inferred as "FALSE". If the valid sign bit hiding flag is "FALSE", step 1480 returns "NO" and the method 1400 proceeds to the sign decoding step 1490.

[0195] The video decoder 134 checks the positions of the first and last significant coefficients of the currently selected sub-block. If the difference between the first significant position and the last significant position is 3 or less, step 1480 returns "NO" and the method 1400 proceeds to the sign decoding step 1490.

[0196] In one configuration of SBH test 1480, the determination depends on the value of the valid sign bit hiding flag and the value of the conversion skip flag for the current TB. If the valid sign bit hiding flag is "TRUE" and the difference between the first significant position and the last significant position is greater than 3, the video decoder 134 also checks the value of the conversion skip flag for the current TB. If the conversion skip flag is "TRUE", step 1480 returns "NO" and method 1400 proceeds to the sign decoding step 1490. Otherwise, if the conversion skip flag is "FALSE", step 1480 returns "YES" and method 1400 proceeds to the step 1495 of decoding and inferring the sign.

[0197] In another configuration of SBH test 1480, the determination depends on the value of the valid sign bit hiding flag and the value of the TSRC invalid flag. If the valid sign bit hiding flag is "TRUE" and the difference between the first significant position and the last significant position is greater than 3, the video decoder 134 also checks the value of the TSRC invalid flag. If the TSRC invalid flag is "TRUE", step 1480 returns "NO" and method 1400 proceeds to the sign decoding step 1490. Otherwise, if the TSRC invalid flag is "FALSE", step 1480 returns "YES" and method 1400 proceeds to the step 1495 of decoding and inferring the sign.

[0198] In yet another configuration of SBH test 1480, the determination depends on the value of the valid sign bit hiding flag, the value of the quantization parameter (QP) of the current TB, and the difference between the first significant position and the last significant position of the quantization parameter QP of the transform block. If the valid sign bit hiding flag is "TRUE" and the difference between the first significant position and the last significant position is greater than 3, the video decoder 134 also checks the value of the QP of the current TB. If the QP is 4 (or any value indicating a reversible operation), step 1480 returns "NO" and method 1400 proceeds to sign decoding step 1490. Otherwise, if the QP is not 4 (or a similar value indicating a reversible operation), step 1480 returns "YES" and method 1400 proceeds to step 1495 of decoding and inferring the sign.

[0199] In sign path decoding step 1490, the sign bits of any significant coefficients of the currently selected sub-block are bypass decoded from the bitstream 133. The sign bits are bypass decoded from the bitstream 133 in the reverse diagonal scan order 810. If the associated sign bit has a value of 1, the value of the quantized coefficient is set to -x[n]. The value of the quantized coefficient is set to x[n] if the associated sign bit has a value of zero. Method 1400 ends with the execution of step 1490.

[0200] In step 1495 of decoding and inferring symbols, the sign bits of the significant coefficients of the currently selected sub-block are bypass decoded from the bit stream 133 in the reverse diagonal scan order 810. The sign bit associated with the first significant coefficient (the last one visited in the reverse diagonal scan order 810) is not decoded from the bit stream 133. In other words, if there are N significant coefficients in the currently selected sub-block, N - 1 sign bits are bypass decoded from the bit stream 133. The sign bit associated with the first significant coefficient is inferred based on the sum of the parities of the significant coefficients. If the sum of the parities is zero, the sign bit associated with the first significant coefficient is inferred to be zero. If the sum of the parities is 1, the sign bit associated with the first significant coefficient is inferred to be 1. If the associated sign bit has a value of 1, the value of the quantized coefficient is set to -x[n]. The value of the quantized coefficient is set to x[n] if the associated sign bit has a value of zero. Thereafter, method 1400 ends.

[0201] The configurations described in methods 900 and 1000 enable reversible compression of video data while using a normal residual coding process. Dependent quantization and sign bit hiding are non-reversible coding tools that can be flexibly disabled when a reversible operation is desired, but can still be available to achieve improved coding performance in non-reversible coding blocks.

[0202] Industrial Applicability The described configurations are applicable to the computer and data processing industries, particularly for digital signal processing for decoding and encoding signals such as video and image signals, and achieve high compression efficiency.

[0203] The above describes only some embodiments of the present invention, and modifications and / or changes can be made to the present invention without departing from the scope and spirit of the present invention. The embodiments are illustrative and not limiting.

Claims

1. 1. A method for decoding a transform block from a bitstream, comprising: decoding a first flag used to determine whether to use dependent quantization in the transform block; determining whether the transform block uses sign bit hiding, in which data indicative of a sign of a significant coefficient at a position is not decoded from the bitstream; If the first flag is TRUE, the sign bit hiding is not used in the transform block; decoding the transform block using the dependent quantization when it is determined based on the first flag that the dependent quantization is used for the transform block; if it is determined that the sign bit hiding is to be used, decoding the transform block using the sign bit hiding; After checking the first flag, a disable flag for transform skip residual coding is checked, and if the disable flag is TRUE, the dependent quantization is not used; The signaling of the invalid flag depends on information indicated by the first flag; The invalid flag indicates whether the first residual coding is applied instead of the second residual coding even when the transform process is skipped, and the first residual coding is a process for a block for which the transform process is not skipped, and the second residual coding is a process for a block for which the transform process is skipped. A method comprising:

2. the first residual coding corresponds to a backward scan order starting at a bottom right position of a sub-block and ending at a top left position of the sub-block; 2. The method of claim 1, wherein the second residual coding corresponds to a forward scan order starting at a top-left location of a sub-block and ending at a bottom-right location of the sub-block.

3. 2. The method of claim 1, wherein if the sign bit hiding is not used, then as many pieces of sign data as there are significant coefficients in a subblock are decoded.

4. 2. The method of claim 1, wherein an enable flag for the sign bit hiding is at least TRUE if the sign bit hiding is to be used.

5. 2. The method of claim 1, wherein the first flag is not a flag related to a sequence parameter set.

6. 1. A method for encoding a transform block into a bitstream, comprising the steps of: encoding a first flag used to determine whether to use dependent quantization in the transform block; determining whether the transform block uses sign bit hiding, in which data indicating the sign of a significant coefficient at a position is not coded into the bitstream; If the first flag is TRUE, the sign bit hiding is not used in the transform block; encoding the transform block using the dependent quantization when it is determined based on the first flag that the dependent quantization is to be used for the transform block; if it is determined that the sign bit hiding is to be used, encoding the transform block using the sign bit hiding; After checking the first flag, a disable flag for transform skip residual coding is checked, and if the disable flag is TRUE, the dependent quantization is not used; The signaling of the invalid flag depends on information indicated by the first flag; The invalid flag indicates whether the first residual coding is applied instead of the second residual coding even when the transform process is skipped, and the first residual coding is a process for a block for which the transform process is not skipped, and the second residual coding is a process for a block for which the transform process is skipped. A method comprising:

7. the first residual coding corresponds to a backward scan order starting at a bottom right position of a sub-block and ending at a top left position of the sub-block; 7. The method of claim 6, wherein the second residual coding corresponds to a forward scan order starting at a top-left location of a sub-block and ending at a bottom-right location of the sub-block.

8. 7. The method of claim 6, wherein if the sign bit hiding is not used, then as many pieces of sign data as there are significant coefficients in a subblock are encoded.

9. 7. The method of claim 6, wherein an enable flag for the sign bit hiding is at least TRUE if the sign bit hiding is to be used.

10. 7. The method of claim 6, wherein the first flag is not a flag related to a sequence parameter set.

11. 1. A decoding device for decoding a transform block from a bitstream, comprising: means for decoding a first flag used to determine whether dependent quantization is used in the transform block; means for determining whether the transform block uses sign bit hiding, in which data indicative of the sign of a significant coefficient at a position is not decoded from the bitstream; If the first flag is TRUE, the sign bit hiding is not used in the transform block; means for decoding the transform block using the dependent quantization when it is determined based on the first flag that the dependent quantization is to be used for the transform block; means for decoding the transform block using the sign bit hiding if it is determined that the sign bit hiding is to be used; After checking the first flag, a disable flag for transform skip residual coding is checked, and if the disable flag is TRUE, the dependent quantization is not used; The signaling of the invalid flag depends on information indicated by the first flag; The invalid flag indicates whether the first residual coding is applied instead of the second residual coding even when the transform process is skipped, and the first residual coding is a process for a block for which the transform process is not skipped, and the second residual coding is a process for a block for which the transform process is skipped. A decoding device comprising:

12. 1. An encoding device for encoding a transform block into a bitstream, comprising: means for encoding a first flag used to determine whether to use dependent quantization in the transform block; means for determining whether the transform block uses sign bit hiding, in which data indicative of the sign of a significant coefficient at a position is not coded into the bitstream; If the first flag is TRUE, the sign bit hiding is not used in the transform block; means for encoding the transform block using the dependent quantization when it is determined based on the first flag that the dependent quantization is to be used for the transform block; means for encoding the transform block using the sign bit hiding if it is determined that the sign bit hiding is to be used; After checking the first flag, a disable flag for transform skip residual coding is checked, and if the disable flag is TRUE, the dependent quantization is not used; The signaling of the invalid flag depends on information indicated by the first flag; The invalid flag indicates whether the first residual coding is applied instead of the second residual coding even when the transform process is skipped, and the first residual coding is a process for a block for which the transform process is not skipped, and the second residual coding is a process for a block for which the transform process is skipped.

13. An encoding device comprising:

13. A computer program product for causing a computer to carry out the method according to any one of claims 1 to 5.

14. A computer program product for causing a computer to carry out the method according to any one of claims 6 to 10.

Citation Information

Patent Citations

  • Image decoding device and image coding device

    JP2021136460A