Method and apparatus for video coding using a palette mode
The palette mode in video encoding and decoding addresses the challenge of high-definition video data compression by optimizing block sizes and using CABAC for efficient entropy encoding, enhancing compression efficiency and image quality.
Patent Information
- Application Number
- JP2023145424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-11
- Filing Date
- 2023-09-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-01-11
AI Technical Summary
The exponential increase in video data with high-definition formats like 4K×2K or 8K×4K poses challenges in efficiently encoding and decoding digital video while maintaining image quality, necessitating improved compression techniques.
Implementing a palette mode for video encoding and decoding, which involves determining a minimum palette mode block size, activating a palette mode flag for units larger than this size, and using a palette table to encode video data, along with context adaptive binary arithmetic coding (CABAC) for entropy encoding.
Enhances encoding efficiency by reducing the bit requirement for palette indices, improving compression performance, and maintaining image quality during high-definition video processing.
Smart Images

Figure 0007704816000021 
Figure 0007704816000022 
Figure 0007704816000023
Abstract
Description
Technical Field
[0001] Related Applications This application claims priority to U.S. Provisional Application No. 62 / 959,913, filed on January 11, 2020, entitled "VIDEO CODING USING PALETTE MODE", the entire disclosure of which is incorporated herein by reference.
[0002] This application generally relates to the coding and compression of video data, and more particularly, to methods and systems for video coding using a palette mode.
Background Art
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Such electronic devices perform the transmission, reception, encoding, decoding, and / or storage of digital video data by implementing video compression extension standards defined by standards such as MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding). Generally, video compression includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in the video data. For block-based video coding, a video frame is divided into one or more slices, and each slice has a plurality of video blocks that can also be referred to as coding tree units (CTUs). Each CTU may contain one coding unit (CU), or may be recursively divided into smaller CUs until a predetermined minimum CU size is reached. Each CU (also named a leaf CU) contains one or more transform units (TUs) and also includes one or more prediction units (PUs). Each CU can be encoded in either an intra mode, an inter mode, or an IBC mode. Video blocks within an intra-coded (I) slice in a video frame are encoded using spatial prediction with respect to reference samples in neighboring blocks within the same video frame.Video blocks within an inter-coded (P or B) slice in a video frame may use spatial prediction with respect to reference samples in neighboring blocks within the same video frame, or may use temporal prediction with respect to reference samples in other previous and / or other future reference video frames.
[0004] For example, spatial or temporal prediction based on previously encoded reference blocks, such as neighboring blocks, results in a prediction block for the current video block to be encoded. The process of finding the reference block can be achieved by a block matching algorithm. Residual data representing the pixel difference between the current block to be encoded and the prediction block is referred to as a residual block or prediction error. An inter-coded block is encoded according to a motion vector indicating the reference block in the reference frame forming the prediction block and the residual block. The process of determining the motion vector is generally referred to as motion prediction. An intra-coded block is encoded according to an intra-prediction mode and the residual block. For further compression, the residual block may be transformed from the pixel domain to a transform domain, such as the frequency domain, to yield residual transform coefficients, which may then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, may be scanned to generate a one-dimensional vector of transform coefficients, which is then entropy encoded into the video bitstream to achieve further compression.
[0005] Next, the encoded video bitstream is stored in a computer-readable recording medium (such as a flash memory) that can be accessed by another electronic device with digital video capabilities, or directly transmitted to the electronic device via wire or wirelessly. Next, the electronic device, for example, analyzes the encoded video bitstream to obtain syntax elements from the bitstream, and based at least in part on the syntax elements obtained from the bitstream, performs video extension (a process opposite to the aforementioned video compression) by reconstructing the digital video data from the encoded video bitstream into its original format, and draws the reconstructed digital video data on the display of the electronic device.
[0006] As the quality of digital video migrates from high definition to 4K×2K or 8K×4K, the amount of video data to be encoded / decoded increases exponentially. This has led to continuous efforts in terms of how to more efficiently encode / decrypt video data while maintaining the image quality of the decoded video data. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] This application describes implementations related to the encoding and decoding of video data, and more particularly, describes a system and method for encoding and decoding video using a palette mode. MEANS FOR SOLVING THE PROBLEMS
[0008] According to a first aspect of the present application, a method for decrypting video data includes receiving, from a bitstream, a plurality of syntax elements associated with an encoded unit, wherein the plurality of syntax elements indicate the size of the encoded unit and the encoding tree type; determining a minimum palette mode block size of the encoded unit according to the encoding tree type of the encoded unit; receiving, from the bitstream, a palette mode activation flag associated with the encoded unit according to a determination that the size of the encoded unit is larger than the minimum palette mode block size; and decrypting the encoded unit from the bitstream according to the palette mode activation flag.
[0009] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When executed by the one or more processing units, these programs cause the electronic device to implement the method for decrypting video data as described above.
[0010] According to a third aspect of the present application, a non-transitory computer-readable recording medium stores a plurality of programs that are executed by an electronic device having one or more processing units. When executed by the one or more processing units, these programs cause the electronic device to implement the method for decrypting video data as described above.
[0011] The accompanying drawings, which are included to provide a further understanding of the implementation forms and constitute a part of this specification, illustrate the described implementation forms and are useful for explaining the basic principles together with the description. Similar reference numerals refer to corresponding parts.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7
[0013] Next, specific implementations will be referred to in detail, and examples thereof are shown in the accompanying drawings. In the following detailed description, many non-limiting and specific details are set forth to assist in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative forms may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented in many types of electronic devices with digital video capabilities.
[0014] FIG. 1 is a block diagram showing an exemplary system 10 for performing video block encoding and decoding in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that is later decoded by a destination device 14. Source device 12 and destination device 14 may comprise any of a variety of electronic devices including, for example, a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, and the like. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0015] In some implementations, destination device 14 receives the encoded Video data may be received. Link 16 may comprise any type of communication medium or communication device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, Link 16 may comprise a communication medium to enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may comprise any wireless or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission paths. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other devices useful for facilitating communication from the source device 12 to the destination device 14.
[0016] In some other implementations, the encoded video data can be transmitted from the output interface 22 to the recording device 32. Subsequently, the encoded video data in the recording device 32 can be accessed by the destination device 14 via the input interface 28. The recording device 32 can include any of a variety of distributed or locally accessible data recording media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or other digital recording media suitable for storing the encoded video data. In a further example, the recording device 32 may correspond to a file server or another intermediate recording device that can hold the encoded video data generated by the source device 12. The destination device 14 can access the stored video data by streaming or downloading from the recording device 32. The file server can be any type of computer that can store the encoded video data or transmit the encoded video data to the destination device 14. Exemplary file servers include a web server (for example, for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 can access the encoded video data through any standard data connection that includes a wireless channel (for example, a Wi-Fi connection), a wired connection (for example, DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the recording device 32 can be a streaming transmission, a download transmission, or a combination of both.
[0017] As shown in FIG. 1, the information source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as, for example, a video camera, a video archive including previously captured video, a video supply interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources, including a source such as a video capture device. As an example, when the video source 18 is a video camera of a security monitoring system, the information source device 12 and the destination device 14 may form a camera phone or a video phone. However, the implementations described in this application may generally be applicable to video coding and may be applied to wireless and / or wired applications.
[0018] Captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 through the output interface 22 of the information source device 12. The encoded video data may also (or instead) be stored in the recording device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.
[0019] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and receives the encoded video data through the link 16. The encoded video data communicated through the link 16 or supplied by the recording device 32 may include various syntax elements generated by the video encoder 20 that are used when the video decoder 30 decodes the video data. The encoded video data that may include such syntax elements is transmitted over a communication medium and stored in a recording medium or a file server.
[0020] In some implementations, the display device 34 that the destination device 14 may include can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 can display the decoded video data to the user and can include any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0021] The video encoder 20 and the video decoder 30 can operate based on intellectual property or industry standards such as VVC, HEVC, MPEG-4 Part 10 AVC (Advanced Video Coding), or an extended version of these standards. It should be understood that this application is not limited to a specific video encoding / decoding standard and can be applicable to other video encoding / decoding standards. In general, it is intended that the video encoder 20 of the source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally intended that the video decoder 30 of the destination device 14 can be configured to decode video data according to any of these current or future standards.
[0022] The video encoder 20 and the video decoder 30 each include one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), It may be implemented as any of various suitable encoded circuit configurations such as an integrated circuit, a field programmable gate array (FPGA), discrete logic, software, hardware, firmware, or any combination thereof. When the electronic device is implemented partially in software, instructions related to the software are stored in a suitable non-transitory computer-readable medium, and the instructions are executed in hardware using one or more processors to perform the video encoding / decoding processes disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of them may be integrated as part of an integrated encoder / decoder (CODEC) combined in their respective devices.
[0023] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some implementations described in the present application. The video encoder 20 may perform intra prediction encoding and inter prediction encoding of video blocks inside a video frame. Intra prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in the video data within adjacent video frames or pictures of a video sequence.
[0024] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB )It includes an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a division unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse conversion processing unit 60, and an adder 62 for reconstructing video blocks. From the reconstructed video, a deblocking filter (not shown) may be arranged between the adder 62 and the DPB 64 to filter the block boundaries to remove block distortion. In addition to the deblocking filter, a loop filter (not shown) may also be used to filter the output of the adder 62. The video encoder 20 may take the form of a non-changeable or programmable hardware unit, or may be divided among one or more non-changeable or programmable hardware units.
[0025] The video data memory 40 may store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18. The DPB 64 is a buffer that records reference video data used to encode video data by the video encoder 20 (e.g., in the intra prediction encoding mode or the inter prediction encoding mode). The video data memory 40 and the DPB 64 may be formed by any of various recording devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or may be off-chip with respect to those components.
[0026] As shown in FIG. 2, the splitting unit 45 inside the prediction processing unit 41 splits the received video data into video blocks. This splitting may include splitting the video frame into slices, tiles, or other larger coding units (CUs) according to a predetermined splitting structure such as a quadtree structure associated with the video data. The video frame may be split into a plurality of video blocks (or a set of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of possible prediction coding modes, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes, for the current video block based on error results (e.g., coding rate or level of distortion). The prediction processing unit 41 may supply the resulting intra prediction coded block or inter prediction coded block to the adder 50 to generate a residual block, and may also supply this coded block to the adder 62 to reconstruct it for later use as part of a reference frame. The prediction processing unit 41 also supplies syntax elements such as motion vectors, intra mode indicators, splitting information, and other such syntax information to the entropy coding unit 56.
[0027] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 inside the prediction processing unit 41 may perform intra prediction coding of the current video block with respect to one or more neighboring blocks in the same frame as the current block to be coded, resulting in spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 inside the prediction processing unit 41 perform inter prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames, resulting in temporal prediction. The video encoder 20 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.
[0028] In some implementations, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating a motion vector indicating the displacement of a prediction unit (PU) of a video block within the current video frame relative to a prediction block within a reference video frame according to a predetermined pattern within a series of video frames. The motion prediction executed by the motion estimation unit 42 is a process of generating a motion vector for estimating the motion of a video block. The motion vector may indicate, for example, the displacement of a PU of a video block within the current video frame or picture relative to a prediction block within a reference frame (or another coding unit) in relation to the current block to be coded within the current frame (or within another coding unit). The predetermined pattern may specify the video frame as a P-frame or a B-frame in the sequence. The intra BC unit 48 may determine a vector for intra BC coding, such as a block vector, in a manner similar to the determination of the motion vector by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vector. The prediction block is a block of the reference frame that is considered to closely correspond to the PU of the video block to be coded from the perspective of pixel difference, and may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference criteria. In some implementations, the video encoder 20 may calculate the values of the sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 may perform a motion search regarding the overall pixel position and fractional pixel position and output a motion vector having fractional pixel accuracy.
[0029] The prediction block is a block of the reference frame that is considered to closely correspond to the PU of the video block to be coded from the perspective of pixel difference, and may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference criteria. In some implementations, the video encoder 20 may calculate the values of the sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 may perform a motion search regarding the overall pixel position and fractional pixel position and output a motion vector having fractional pixel accuracy.
[0030] The motion estimation unit 42 calculates a motion vector by comparing the position of a prediction block in a reference frame selected from the first reference frame list (list 0) or the second reference frame list (list 1) with the position of the PU of the video block of the inter-predicted coded frame. Here, the first reference frame list or the second reference frame list identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.
[0031] Motion compensation performed by the motion compensation unit 44 may include fetching or generating a prediction block based on the motion vector determined by the motion estimation unit 42. When the motion compensation unit 44 receives a motion vector for the PU of the current video block, it searches for the prediction block indicated by the motion vector in one of the reference frame lists, extracts the prediction block from the DPB 64, and transfers the prediction block to the adder 50. Then, the adder 50 forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the current video block to be coded. The pixel difference values forming the residual video block may include a luminance difference component or a chroma difference component, or both. The motion compensation unit 44 may also generate syntax elements related to the video blocks of the video frame used when the video decoder 30 decodes the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the prediction block, any flag indicating the prediction mode, or other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 can be mostly integrated but are shown separately for conceptual purposes.
[0032] In some implementations, the intra BC unit 48 can generate vectors and capture prediction blocks in the same way as described above for the motion estimation unit 42 and the motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are referred to as block vectors as opposed to motion vectors. Specifically, the intra BC unit 48 may be determined to use an intra prediction mode to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during an individual encoding pass, and analyze the performance of those intra prediction modes by rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use to generate an intra mode indicator among the various intra prediction modes tested. For example, the intra BC unit 48 may use rate-distortion analysis to calculate rate-distortion values for the various intra prediction modes tested and select an intra prediction mode having the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original block before encoding used to generate the encoded block, along with the bit rate (i.e., number of bits) used to generate these encoded blocks. The intra BC unit 48 may calculate the ratio of distortion to rate for various encoded blocks and determine an intra prediction mode that indicates the best rate-distortion value for that block. and select an intra prediction mode having the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original block before encoding used to generate the encoded block, along with the bit rate (i.e., number of bits) used to generate these encoded blocks. The intra BC unit 48 may calculate the ratio of distortion to rate for various encoded blocks and determine an intra prediction mode that indicates the best rate-distortion value for that block.
[0033] In other examples, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction according to the implementation forms described in this specification. In either case, for intra block copy, the prediction block may be a block that is considered to closely correspond to the block to be coded from the perspective of pixel difference, and may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. The determination of the prediction block may include the calculation of the values of sub-integer pixel positions.
[0034] Whether the prediction block is from the same frame by intra prediction or from different frames by inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block to be coded, and form pixel difference values. The pixel difference values for forming the residual video block may include both the luminance difference component and the chrominance difference component.
[0035] As described above, the intra prediction processing unit 46 may intra predict the current video block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction performed by the intra BC unit 48. Specifically, the intra prediction processing unit 46 may determine to use an intra prediction mode to code the current block. To do so, the intra prediction processing unit 46 may, for example, code the current block using various intra prediction modes during an individual coding pass, and the intra prediction processing unit 46 (or in some examples, the mode selection unit) may select an appropriate intra prediction mode to use from the tested intra prediction modes. The intra prediction processing unit 46 may supply information representing the intra prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 may code the information indicating the selected intra prediction mode in the bitstream.
[0036] After the prediction processing unit 41 determines a prediction block for the current video block by either inter prediction or intra prediction, the adder 50 generates a residual video block by subtracting the prediction block from the current video block. Residual video data in the residual block may be included in one or more transform units (TUs) and is supplied to the transform processing unit 52. The transform processing unit 52 uses a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to transform the residual video data into residual transform coefficients. transform) or a transform such as a conceptually similar transform to transform the residual video data into residual transform coefficients.
[0037] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be changed by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0038] Following quantization, the entropy coding unit 56 performs, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), context-adaptive binary arithmetic coding), probability interval partitioning entropy coding (PIPE), or another entropy coding technique or technology to entropy code the quantized transform coefficients into a video bitstream. The encoded bitstream can then be transmitted to the video decoder 30, recorded in the recording device 32 for later transmission to or retrieval by the video decoder 30. The entropy coding unit 56 may also entropy code the motion vectors and other syntax elements for the current video frame to be encoded.
[0039] To generate a reference block for predicting other video blocks, the inverse quantization unit 58 applies inverse quantization and the inverse transform processing unit 60 applies inverse transform to reconstruct the residual video block in the pixel region. As described above, the motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion prediction.
[0040] The adder 62 generates a reference block for storing the reconstructed residual block in the DPB 64 in addition to the motion compensation prediction block generated by the motion compensation unit 44. The reference block can then be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to inter-predict another video block in a subsequent video frame.
[0041] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is overall inverse to the encoding process described with respect to the video encoder 20 in relation to FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra prediction processing unit 84 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.
[0042] In some examples, tasks may be assigned to the units of the video decoder 30 to execute the implementations of the present application. Also, in some examples, the implementations of the present disclosure may be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 may execute the implementations of the present application alone or in combination with other units such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80 of the video decoder 30. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functionality of the intra BC unit 85 may be executed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0043] The video data memory 79 can store video data such as an encoded video bit stream decoded by other components of the video decoder 30. The video data stored in the video data memory 79 can be acquired from the recording device 32, from a local video source such as a camera, by wired or wireless network communication of the video data, or by accessing a physical data recording medium such as a flash drive or a hard disk. The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bit stream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data used to encode video data by the video decoder 30 (e.g., in an intra prediction coding mode or an inter prediction coding mode). The video data memory 79 and the DPB 92 can be formed by any of various memory devices including synchronous dynamic random access memory (SDRAM), magneto-resistive RAM (MRAM), resistive random access memory (RRAM), or other types of memory devices. For purposes of illustration, the video data memory 79 and the DPB 92 are shown as two separate components of the video decoder 30 in FIG. 3. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or individual memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30 or off-chip with respect to those components.
[0044] During the decoding process, video decoder 30 receives an encoded video bitstream representing an encoded video frame and video blocks of associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or at the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Then, entropy decoding unit 80 transfers the motion vectors and other syntax elements to prediction processing unit 81.
[0045] When a video frame is encoded as an intra prediction encoded (I) frame or for intra encoded prediction blocks in other types of frames, intra prediction processing unit 84 of prediction processing unit 81 may generate prediction data for video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.
[0046] When a video frame is encoded as an inter prediction encoded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be generated from one of the internal reference frames of a reference frame list. Video decoder 30 may configure reference frame lists, list 0 and list 1, using a default configuration technique based on the reference frames stored in DPB 92.
[0047] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed area of the same picture as the current video block defined by the video encoder 20.
[0048] The motion compensation unit 82 and / or the intra BC unit 85 determines prediction information for the video block of the current video frame by analyzing the motion vector and other syntax elements, and then uses the prediction information to generate a prediction block for the current video block to be decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode the video block of the video frame, the inter prediction frame type (e.g., B or P), the configuration information of one or more of the reference frame lists for the frame, the motion vector of each inter prediction encoded video block in the frame, the inter prediction state of each inter prediction encoded video block in the frame, and other information for decoding the video block in the current video frame.
[0049] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as a flag, to determine that the current video block is predicted using the intra BC mode, the configuration information of the video block of the frame that should be stored in the DPB 92 within the reconstructed area, the block vector of each intra BC prediction video block in the frame, the intra BC prediction state of each intra BC prediction video block in the frame, and other information for decoding the video block in the current video frame.
[0050] The motion compensation unit 82 may also perform interpolation using an interpolation filter such as that used to calculate the interpolation values of sub-pixel of the reference block during the encoding of the video block by the video encoder 20. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements, and generate a prediction block using the interpolation filter.
[0051] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients given in the bitstream and entropy decoded by the entropy decoder 80 using the same quantization parameter as that calculated to determine the degree of quantization for each video block in the video frame by the video encoder 20. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.
[0052] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) may be arranged between the adder 90 and the DPB 92 to further process the decoded video block. Then, the decoded video block in a given frame is stored in the DPB 92 that stores the reference frames used for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for presentation to a display device such as the display device 34 of FIG. 1 later.
[0053] In a general video coding process, a video sequence generally includes an ordered set of frames or pictures. Each frame may include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array consisting of luma samples. SCb is a two-dimensional array consisting of Cb chroma samples. SCr is a two-dimensional array consisting of Cr chroma samples. In other cases, the frame may be black and white and thus include only one two-dimensional array of luma samples.
[0054] As shown in FIG. 4A, video encoder 20 (more specifically, splitting unit 45) generates an encoded representation of a frame by first splitting the frame into a set of coding tree units (CTUs). A video frame may include a whole number of CTUs that are sequentially ordered in a raster scan order from left to right and top to bottom. Each CTET is the largest logical coding unit, and the width and height of the CTU are signaled in the sequence parameter set by video encoder 20 such that all CTUs in the video sequence have the same size, which is one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU may include one coding tree block (CTB) consisting of luma samples, a coding tree block consisting of two corresponding chroma samples, and syntax elements used to encode the samples of the coding tree block. The syntax elements describe the characteristics of various types of units of the encoded block of pixels and how the video sequence can be reconstructed at video decoder 30, including inter prediction or intra prediction, intra prediction mode, motion vectors, and other parameters. For a black and white picture or a picture having three separate color planes, the CTU may include a single coding tree block and syntax elements used to encode the samples of the coding tree block. The coding tree block may be an N×N block of samples.
[0055] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary-tree partitioning, ternary-tree partitioning, quadtree partitioning, or combinations thereof, on the coding tree blocks of the CTU to divide the CTU into smaller coding units (CUs). As shown in FIG. 4C, a 64×64 CTU 400 is first divided into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. Two 16×16 CUs, 430 and 440, are further divided into four CUs each with a block size of 8×8. FIG. 4D represents a quadtree data structure showing the final result of the CTU 400 partitioning process as shown in FIG. 4C, where each leaf node of the quadtree corresponds to one CU of each size in the range of 32×32 to 8×8. Each CU may include, similar to the CTU shown in FIG. 4B, a coding block (CB) of luminance samples, two corresponding coding blocks of chrominance samples of a frame of the same size, and syntax elements used to code the samples of the coding block. In a monochrome picture or a picture having three separate color planes, the CU may include a single coding block and a syntax structure used to code the samples of the coding block. Note that the quadtree partitioning shown in FIGS. 4C and 4D is for illustrative purposes only, and it should be noted that one CTU can be divided into CUs based on quadtree / ternary-tree / binary-tree partitioning to adapt to various local characteristics. In a composite tree structure, one CTU can be divided by a quadtree structure, and each leaf CU of the quadtree can be further divided by binary-tree and ternary-tree structures. As shown in FIG. 4E, there are five types of partitioning, such as four-way partitioning, horizontal two-way partitioning, vertical two-way partitioning, horizontal three-way partitioning, and vertical three-way partitioning.
[0056] In some implementations, video coder 20 may further divide the coding block of a CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same (inter or intra) prediction is applied. The prediction unit (PU) of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements used to predict the prediction block. In a monochrome picture or a picture having three separate color planes, the PU may include a single prediction block and the syntax structure used to predict the prediction block. Video coder 20 may generate a predicted luma, Cb and Cr blocks for luma, and Cb and Cr prediction blocks, in each PU of the CU.
[0057] Video coder 20 may use intra prediction or inter prediction to generate the prediction block for the PU. When video coder 20 uses intra prediction to generate the prediction block of the PU, video coder 20 may generate the prediction block of the PU based on the decoded samples of the frame associated with the PU. When video coder 20 uses inter prediction to generate the prediction block of the PU, video coder 20 may generate the prediction block of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.
[0058] After the video encoder 20 generates a predicted luminance block, a predicted Cb block, and a predicted Cr block for one or more PUs in a CU, the video encoder 20 can generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from the original luminance encoded block of the CU such that each sample in the luminance residual block of the CU indicates a difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance encoded block of the CU. Similarly, the video encoder 20 can generate a Cb residual block and a Cr residual block for the CU such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb encoded block of the CU, and each sample in the Cr residual block of the CU can indicate a difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr encoded block of the CU.
[0059] Furthermore, as shown in FIG. 4C, the video encoder 20 uses quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of the CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of the CU can include a transform block of luminance samples, two corresponding transform blocks of chroma samples, and syntax elements used to predict the transform block samples. Thus, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with the TU can be a sub-block of the luminance residual block of the CU. The Cb transform block can be a sub-block of the Cb residual block of the CU. The Cr transform block can be a sub-block of the Cr residual block of the CU. In a black and white picture or a picture having three separate color planes, the TU can include a single transform block and a syntax structure used to transform the samples of the transform block.
[0060] Video encoder 20 may apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0061] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to the process by which transform coefficients are quantized to somehow reduce the amount of data used to represent the transform coefficients, resulting in further compression. After quantizing the coefficient block, video encoder 20 may entropy encode a syntax element indicating the quantized transform coefficients. For example, video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax element indicating the quantized transform coefficients. Finally, Video encoder 20 may output a bitstream including a series of bits forming a representation of the encoded frame and associated data, which is stored in recording device 32 or transmitted to destination device 14.
[0062] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can analyze the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 can also reconstruct the encoded block of the current CU by adding the samples of the prediction block for the current CU's PU to the samples of the transform block of the corresponding current CU's TU. After reconstructing the encoded block for each CU of the frame, the video decoder 30 can reconstruct the frame.
[0063] As described above, video coding mainly uses two modes, intra-frame prediction (i.e., intra prediction) and inter-frame prediction (i.e., inter prediction), to achieve video compression. Palette-based coding is another coding method adopted by many video coding standards. Palette-based coding is particularly suitable for coding the content generated on the screen. In this method, a video coder (e.g., the video encoder 20 or the video decoder 30) forms a color palette table that represents the video data of a given block. The palette table contains the most dominant (e.g., frequently used) pixel values in a given block. Pixel values that are not frequently represented in the video data of a given block are not included in the palette table or are included in the palette table as avoidable colors.
[0064] Each entry in the palette table contains an index for the corresponding pixel value in the palette table. The palette index for the samples in a block can be encoded to indicate the entry in the palette table used to predict or reconstruct the samples. This palette mode begins with a process that generates a palette predictor for the first block of such a grouping of a picture, slice, tile, or video block. As described below, the palette predictor for subsequent video blocks is generally generated by updating the previously used palette predictor. For purposes of illustration, it is assumed that the palette predictor is defined at the picture level. In other words, a picture may contain multiple coded blocks each having its own palette table, but there is one palette predictor for the entire picture.
[0065] To reduce the number of bits required to signal palette entries in the video bitstream, the video decoder may utilize the palette predictor to determine new palette entries for the palette table used to reconstruct the video block. For example, the palette predictor may include palette entries from a previously used palette table, or may be initialized using the most recently used palette table by including all entries of the most recently used palette table. In some implementations, the palette predictor includes fewer entries than all entries from the most recently used palette table, and may then incorporate some entries from other previously used palette tables. The size of the palette predictor may be the same as, larger than, or smaller than the size of the palette table used to encode different blocks. In one example, the palette predictor is implemented as a first-in first-out (FIFO) table containing 64 palette entries.
[0066] To generate a palette table for a block of video data from a palette predictor, the video decoder may receive a 1-bit flag from the encoded video bitstream for each input of the palette predictor. The 1-bit flag may have a first value (e.g., binary 1) indicating that the associated input of the palette predictor is included in the palette table or a second value (e.g., binary 0) indicating that the associated input of the palette predictor is not included in the palette table. If the size of the palette predictor is larger than the palette table used for the block of video data, the video decoder may stop receiving further flags once the maximum size of the palette table has been reached.
[0067] In some implementations, some entries of the palette table may not be determined using the palette predictor but may be signaled directly in the encoded video bitstream. For such entries, the video decoder may receive from the encoded video bitstream three separate m-bit values indicating pixel values for the luminance component and two chrominance components associated with the entry, where m represents the bit depth of the video data. While multiple m-bit values are required for directly signaled palette entries, only a 1-bit flag is required for palette entries derived from the palette predictor. Thus, signaling some or all of the palette inputs using the palette predictor can significantly reduce the number of bits required to signal the inputs of the new palette table, thereby improving the overall encoding efficiency of the palette mode encoding.
[0068] In many cases, the palette predictor for a block is determined based on the palette table used to encode one or more previously encoded blocks. However, when encoding the first coding tree unit in a picture, slice, or tile, it may not be possible to utilize the palette table of the previously encoded blocks. Therefore, it is not possible to generate the palette predictor using the entries of the previously used palette table. In such cases, a series of palette predictor initialization specifiers, which are the values used to generate the palette predictor when the previously used palette table is not available, may be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS). The SPS generally refers to the syntax structure of the syntax elements that conform to a series of consecutive encoded video pictures called the coded video sequence (CVS), as determined by the content of the syntax elements found in the PPS, which are referred to by the syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of the syntax elements that conform to one or more individual pictures within the CVS, as determined by the syntax elements found in each slice segment header. Therefore, the SPS is generally regarded as a higher-level syntax structure than the PPS, and the syntax elements included in the SPS generally do not change as frequently and conform to a larger portion of the video data compared to the syntax elements included in the PPS.
[0069] FIGS. 5A-5B are block diagrams showing examples of using a palette table to encode video data according to some implementations of the present disclosure.
[0070] For palette (PLT) mode signaling, the palette mode is coded as a prediction mode for the coding unit, i.e., the prediction mode for the coding unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. When the palette mode is utilized, the pixel values of the CU are represented by a small set of representative colors. This set is referred to as the palette. For pixels having values close to the palette colors, the palette index is signaled. Pixels having values outside the palette are represented with an escape symbol and the quantized pixel values are directly signaled. The syntax and related semantics of the palette mode in the current VVC draft specification are shown in Table 1 and Table 2 below, respectively.
[0071] To decode a block encoded in palette mode, the decoder needs to decode the palette colors and indices from the bitstream. The palette colors are defined by a palette table and encoded by a palette table encoding syntax (e.g., palette_predictor_run, num_signaled_palette_entries, new_palette_entries). For each CU, an escape flag palette_escape_val_present_flag is signaled to indicate whether there is an escape symbol in the current CU. If there is an escape symbol, the palette table is augmented by another entry and the last index is assigned to the escape mode. The palette indices of all pixels in the CU form a palette index map and are encoded by a palette index map encoding syntax (e.g., num_palette_indices_minus1, palette_idx_idc, copy_above_indices_for_final_run_flag, palette_transpose_flag, copy_above_palette_indices_flag, palette_run_prefix, palette_run_suffix). An example of a CU encoded in palette mode is shown in Figure 5A, where the palette size is 4. The first three samples in the CU use palette entries 2, 0, and 3, respectively, for reconstruction. The "x" sample in the CU represents an escape symbol. The CU-level flag, palette_escape_val_present_flag, indicates whether there is any escape symbol in the CU. If there is an escape symbol, the palette size is augmented by one and the last index is used to indicate the escape symbol. Thus, in Figure 5A, index 4 is assigned to the escape symbol.
[0072] If the palette index (e.g., index 4 in FIG. 5A) corresponds to an escape symbol, additional overhead is signaled to indicate the corresponding color of the sample.
[0073] In some embodiments, on the encoder side, it is necessary to derive an appropriate palette to be used with the CU. To derive the palette for lossy coding, a modified k - means clustering algorithm is used. The first sample of the block is added to the palette. Then, for each subsequent sample from the block, the sum of absolute differences (SAD) between the sample and each of the current palette colors is calculated. The sample is added to the cluster belonging to the palette entry if the distortion of each component of the sample is less than the threshold of the SAD corresponding to the palette entry with the minimum distortion. Otherwise, the sample is added as a new palette entry. When the number of samples mapped to a cluster exceeds the threshold, the centroid of that cluster is updated to become the palette entry for that cluster.
[0074] In the next step, the clusters are sorted in descending order of usage. Then, the palette entries corresponding to each entry are updated. Usually, the cluster centroid is used as the palette entry. However, when considering the encoding cost of the palette entry, a rate - distortion analysis is performed to analyze whether there is any entry from a more appropriate palette predictor that can be used as the updated palette entry instead of the centroid. This process continues until all clusters are processed or the maximum palette size is reached. Finally, if a cluster has only a single sample and the corresponding palette entry is not in the palette predictor, the sample is converted to an escape symbol. In addition, duplicate palette entries are removed and their clusters are combined.
[0075] After palette derivation, each sample in the block is assigned the index of the closest palette entry (in terms of SAD). The sample is then assigned to either the "INDEX" mode or the "COPY_ABOVE" mode. For each sample for which either the "INDEX" mode or the "COPY_ABOVE" mode is possible, the execution of each mode is determined. Then, the cost of encoding the mode is calculated. The mode with the lower cost is selected.
[0076] A palette predictor is maintained for encoding the palette table. Both the maximum size of the palette and the maximum size of the palette predictor can be signaled at the SPS (or other coding levels such as PPS, slice header). The palette predictor is initialized at the start of each slice where the palette predictor is reset to 0. For each entry of the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. As shown in Figure 5B, the reuse flag, palette_predictor_run is sent. After this, the number of new palette entries is signaled using an exponential Golomb code of order from 0 to the syntax num_signaled_palette_entries. Finally, the component values for the new palette entries, new_palette_entries[] are signaled. After encoding the current CU, the palette predictor is updated using the current palette, and at the end of the new palette predictor, entries from the previous palette predictor that are not reused in the current palette are added until the maximum possible size is reached.
[0077] To encode the palette index map, as shown in Figure 5C, the indexes are encoded using a horizontal scan or a vertical scan. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag.
[0078] The palette index is encoded using two main palette sample modes, "INDEX" and "COPY_ABOVE". In the "INDEX" mode, the palette index is explicitly signaled. In the "COPY_ABOVE" mode, the palette index of the sample in the above row is copied. For both the "INDEX" mode and the "COPY_ABOVE" mode, an execution value that defines the number of pixels encoded using the same mode is signaled. The mode is signaled using a flag, but the top row when horizontal scanning is used, or the first column when vertical scanning is used or when the previous mode was "COPY_ABOVE" is not signaled.
[0079] In some embodiments, the encoding order for the index map is as follows. First, the number of index values of the CU is signaled using the syntax num_palette_indices_minus1, and then the actual index values of the entire CU are signaled using the syntax palette_idx_idc. Both the number of indices and the index values are encoded in bypass mode. This groups the bypass encoding bins related to the indices. Then, using the syntax copy_above_palette_indices_flag, palette_run_prefix, and palette_run_suffix, the palette mode (INDEX or COPY_ABOVE) and the execution are signaled in an interleaved manner. The copy_above_palette_indices_flag is a context encoding flag (there is only one bin), the codeword of palette_run_prefix is determined by the process described in Table 3 below, and the first five bins are context encoded. The palette_run_suffix is bypass It is encoded as a bin. Finally, the component escape values corresponding to the escape samples for the entire CU are grouped and encoded in bypass mode. After signaling the index value, an additional syntax element copy_above_indices_for_final_run_flag is signaled. This syntax element eliminates the need to signal the execution value corresponding to the last execution in the block, along with the number of indices.
[0080] In the reference software of VVC (VTM), the dual tree is enabled for I slices, separating the coding unit splitting for the luminance component and the chrominance components. As a result, the palette is applied separately for the luminance (Y component) and the chrominance (Cb component and Cr component). When the dual tree is disabled, the palette is applied together for the Y component, Cb component, and Cr component.
[0081] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6]
[0082] [Table 2-1] [Table 2-2]
Table 2-3
Table 2-4
Table 2-5
[0083]
Table 3-1
Table 3-2
[0084] At the 15th JVET meeting, in order to simplify the buffer usage method and syntax in the palette mode of VTM6.0, line-based coefficient groups (CGs) have been proposed (document number JVET-O0120, accessible at http: / / phenix.int-evry.fr / jvet / ). A CU is divided into multiple line-based coefficient groups, each consisting of m samples, similar to the coefficient groups (CGs) used in the coding of transform coefficients. Here, the index execution, palette index value, and quantized color for the escape mode are sequentially coded / parsed for each CG. As a result, the pixels in the line-based CG can be reconstructed after parsing syntax elements such as the index execution, palette index value, and escape quantized color related to the CG. Thus, in VTM6.0, the buffer requirements in the palette mode where all syntax elements related to a CU need to be parsed (and stored) before reconstruction are greatly reduced. The buffer requirements in the palette mode where all syntax elements related to a CU need to be parsed (and stored) before reconstruction are greatly reduced.
[0085] In this application, each CU in the palette mode is divided into a plurality of segments of m samples (m = 8 in this test) as shown in FIG. 5D based on the horizontal scan mode.
[0086] In each segment, the encoding order for palette execution encoding is as follows. For each pixel, one context encoding bin run_copy_flag = 0 is signaled, indicating that the pixel is in the same mode as the previous pixel, i.e., both the previously scanned pixel and the current pixel are of execution type COPY_ABOVE or of execution type INDEX and have the same index value. Otherwise, run_copy_flag = 1 is signaled.
[0087] If the current pixel and the previous pixel are in different modes, one context encoding bin copy_above_palette indices_flag indicating the execution type of the pixel, i.e., INDEX or COPY_ABOVE, is signaled. In this case, since the INDEX mode is used by default, the decoder does not need to parse the execution type to determine whether the sample is in the first row (horizontal scan) or the first column (vertical scan). The decoder also does not need to parse the execution type to determine whether the previously parsed execution type is COPY_ABOVE.
[0088] After the palette execution encoding of the pixels in one segment, the index value (in INDEX mode) and the quantized escape color are encoded as bypass bins and grouped separately from the encoding / parsing of the context encoding bins to improve the internal processing capacity of each line-based CG. After the encoding is executed, since the index value is then encoded / parsed, the encoder does not need to signal the number of index values num_palette_indices_minus1 and the last execution type copy_above_indices_for_final_run_flag. The syntax of the CG palette mode is shown in Table 4.
[0089]
Table 4-1
Table 4-2
Table 4-3
Table 4-4
Table 4-5
[0090] FIG. 6 is a flowchart 600 showing an exemplary process by which a video decoder (e.g., video decoder 30) implements a technique for decoding video data according to some implementations of the present disclosure.
[0091] Regarding the palette mode of VVC, the palette mode may be applicable to CUs with 64×64 pixels or less. In some embodiments, a minimum palette mode block size is proposed to disable the palette mode for coding units smaller than the minimum palette mode block size in order to reduce complexity. For example, it has been proposed to disable the palette mode for all blocks smaller than a certain threshold, such as 16 samples. Since there are various color difference formats (e.g., 4:4:4, 4:2:2, 4:2:0) and various coding tree types (e.g., SINGLE_TREE, DUAL_TREE_LUMA, and DUAL_TREE_CHROMA), this threshold may vary. Note that "SINGLE_TREE" indicates that the luminance component and the color difference component of the image are similarly divided so that these two components share the same palette table and palette predictor under that palette mode. In contrast, "DUAL_TREE" indicates that the luminance component and the color difference component of the image are separately divided so that these two components obtain separate palette tables and palette predictors under that palette mode. For example, for the 4:2:0 format of YUV with the "DUAL_TREE" type where separate components are considered separately, the palette mode of the color difference component of CUs smaller than 16 samples may be disabled to reduce complexity. The following Table 5 shows an example of the proposed syntax.
[0092]
Table 5
[0093] In Table 5, pred_mode_plt_flag defines whether the palette mode is enabled (e.g., a value of 1) or disabled (e.g., a value of 0) for the coding unit. Parameters such as SubWidthC and SubHeightC are associated with the color difference format of the coding unit as follows.
Table 6
[0094] In black-and-white sampling, there is only one sample array, which is nominally regarded as the luminance array. In 4:2:0 sampling, each of the two chrominance arrays has half the height and half the width of the luminance array. In 4:2:2 sampling, each of the two chrominance arrays has the same height as the luminance array and half the width. In 4:4:4 sampling, each of the two chrominance arrays has the same height and the same width as the luminance array.
[0095] In another embodiment, in the case of a single tree, the palette mode is disabled for small-sized blocks depending on the luminance block size. In an example regarding the YUV 420 format, the palette mode for CUs smaller than 16 pixels is disabled in the case of a single tree depending on the luminance block size. In a specific example, since the activation of the palette is conditional on the size of the luminance samples and ignores the size of the chrominance samples, the palette mode can be enabled for an 8×4 CU containing 8×4 luminance samples and two 4×2 chrominance samples.
[0096] During decoding of the bitstream, video decoder 30 first receives (610) from the bitstream a plurality of syntax elements associated with the coding unit. The plurality of syntax elements indicate the size of the coding unit and the coding tree type. For example, the coding unit of the coding tree type can be one of SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA. Video decoder 30 then determines (620) the minimum palette mode block size of the coding unit according to the coding tree type of the coding unit. For example, as shown in Table 5 above, when the coding unit of the coding tree type is SINGLE_TREE or DUAL_TREE_LUMA, video decoder 30 sets the minimum palette mode block size to 16 samples. When the coding unit of the coding tree type is DUAL_TREE_CHROMA, video decoder 30 first determines the color difference format for the coding unit, and then sets the minimum palette mode block size according to the color difference format as shown in the above table. For example, when the color difference format is 4:4:4, the minimum palette mode block size is 16 samples; when the color difference format is 4:2:2, the minimum palette mode block size is 32 samples; and when the color difference format is 4:2:0, the minimum palette mode block size is 64 samples.
[0097] Video decoder 30 receives (640) a palette mode activation flag associated with an encoding unit from a bitstream in response to a determination that the size of the encoding unit is larger than the palette mode block size with the smallest size of the encoding units (630), and then decodes (650) the encoding unit from the bitstream according to the palette mode activation flag. In some embodiments, when the palette mode activation flag indicates that the palette mode is activated for the encoding unit, video decoder 30 generates (670) a palette table for the current unit from the bitstream, and then decodes (680) the encoding unit from the bitstream using the generated palette table as described above in connection with FIGS. 5A-5D.
[0098] FIG. 7 is a block diagram illustrating an exemplary context adaptive binary arithmetic coding (CABAC) engine according to some implementations of the present disclosure.
[0099] Context adaptive binary arithmetic coding (CABAC) is a form of entropy coding used in many video coding standards such as, for example, H.264 / MPEG-4 AVC, High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC). CABAC is based on arithmetic coding and conforms to the requirements of video coding standards through several technological innovations and changes. For example, CABAC encodes binary symbols, maintains low complexity, and enables probability modeling for the more frequently used bits of any symbol. The probability model is adaptively selected based on the local context, and since the coding modes are usually locally well correlated, it enables better probability modeling. Finally, CABAC uses range division without multiplication by using quantized probability ranges and probability states.
[0100] CABAC has multiple probability models for various contexts. CABAC is , initially, convert all non-binary symbols to binary. Then, the encoder selects the probability model to be used for each bin (or bit), and then optimizes the probability estimation using information from nearby elements. Finally, arithmetic coding is applied to compress the data.
[0101] Context modeling results in the estimation of the conditional probability of the encoded symbol. Utilizing an appropriate context model, redundancy between given symbols can be exploited by switching between different probability models according to the already encoded symbols in the neighborhood of the current symbol for encoding. The step of encoding the data symbol includes the following stages.
[0102] Binarization processing: CABAC uses binary arithmetic coding, which means that only binary decisions (1 or 0) are encoded. Non-binary symbols (such as transform coefficients or motion vectors) are "binarized", i.e., converted to binary codes, prior to arithmetic coding. This process is similar to the process of converting data symbols to variable-length codes, but the binary code is further encoded (by the arithmetic coder) before transmission. The stage is repeated for each bin (i.e., "bit") of the binarized symbol.
[0103] Selection of context model: A "context model" is a probability model for one or more bins of the binarized symbol. This model may be selected from the available models depending on the statistics of the most recently encoded data symbol. The context model stores the probability that each bin is "1" or "0".
[0104] Arithmetic coding: The arithmetic coder encodes each bin according to the selected probability model. Note that for each bin, there are exactly two sub-ranges (corresponding to "0" and "1").
[0105] Probability update: The selected context model is updated based on the actual encoded value (e.g., if the bin value is "1", the count of the frequency of "1" is incremented).
[0106] By decomposing the value of each non-binary syntax element into a series of bins, further processing of each bin value in CABAC can be selected to be in the normal mode or the bypass mode depending on the associated coding mode decision. The bins for which the bypass mode is selected are assumed to have a uniform distribution, and as a result, all normal binary arithmetic coding (and decoding) processes are simply bypassed. In the normal coding mode, each bin value is encoded by using a normal binary arithmetic coding engine, and the associated probability model is determined by a certain selection based on the type of the syntax element and the bin position of the binary representation of the syntax element, i.e., the bin index (binIdx), or is adaptively selected from two or more probability models depending on the associated side information (e.g., the spatial neighborhood, component, depth or size of the CU / PU / TU, or the position inside the TU). The selection of the probability model is referred to as context modeling. As an important design decision, the latter is generally applied only to the most frequently observed bins, and other bins that are not so frequently observed are generally processed using concatenation and typically using a zero-order probability model. In this way, CABAC enables selective adaptive probability modeling at the sub-symbol level, thus providing an efficient means for exploiting inter-symbol redundancy with a considerably reduced overall modeling cost or learning cost. It should be noted that for both certain selection and adaptive selection, in principle, a switch from one probability model to another can occur between any two consecutive normally encoded bins. Generally, in CABAC, the design of the context model finds an excellent compromise between the conflicting goals of preventing an overhead of unnecessary modeling cost and exploiting statistical dependencies to a considerable extent. This reflects the goal of finding.
[0107] The parameters of the probability model in CABAC are adaptive, which means that the adaptation of the model probability to the statistical variations of the bin sources is performed synchronously in a backward-adaptive manner for each bin in both the encoder and the decoder, and this process is called probability estimation. For this purpose, each probability model in CABAC can adopt one of 126 separate states, and the related model probability value p ranges from [0.01875; 0.98125]. The two parameters of each probability model are stored as 7-bit entries in the context memory, with 6 bits for each of the 63 probability states representing the model probability pLPS of the least probable symbol (LPS), and 1 bit for the value nMPS of the most probable symbol (MPS).
[0108] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable recording medium corresponding to a tangible medium such as a data recording medium, or a communication medium including any medium that facilitates transfer of a computer program from one location to another, for example, according to a communication protocol. Thus, a computer-readable medium generally may correspond to either (1) a non-transitory tangible computer-readable recording medium, or (2) a communication medium such as a signal or a carrier wave. The data recording medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.
[0109] The terminology used in the description of the embodiments herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The terms "comprising" and / or "comprises" when used herein specify the presence of the stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.
[0110] For the purpose of describing various elements, terms such as first, second, etc. may be used herein, but it should also be understood that these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, without departing from the scope of the embodiment, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes but not the same electrode.
[0111] The description of the present application is presented for purposes of illustration and description and is not intended to be exhaustive or limited to the invention in the disclosed form. Many modifications, variations, and alternative embodiments should be apparent to those skilled in the art who have the benefit of the teachings presented in the foregoing description and the related drawings. The embodiments are presented to best explain the principles of the invention and its practical application to enable others skilled in the art to understand the invention with respect to various embodiments, and also to include various modifications It is selected and described in order to make the best use of various implementation forms accompanied by states. Therefore, it should be understood that the scope of the claims is not limited to specific examples of the disclosed implementation forms and their modified forms, and other implementation forms are intended to be included within the scope of the appended claims.
Claims
1. A method for decrypting video data, comprising: deriving a plurality of variables related to an encoding unit, wherein the plurality of variables indicate the size of a luminance encoding block and an encoding tree type, the encoding unit is related to a predetermined splitting method, and the predetermined splitting method includes four-way splitting, three-way horizontal splitting, three-way vertical splitting, two-way horizontal splitting, or two-way vertical splitting; determining a threshold value of a palette mode block size of the encoding unit according to the encoding tree type; in response to a determination that at least the size of the luminance encoding block is larger than the threshold value of the palette mode block size, receiving a palette mode activation flag related to the encoding unit from a bitstream; decrypting the encoding unit from the bitstream according to the palette mode activation flag; and the step of determining a threshold value of a palette mode block size of the encoding unit according to the encoding tree type includes: when the encoding tree type is SINGLE_TREE or DUAL_TREE_LUMA, setting the threshold value of the palette mode block size to 16; when the encoding tree type is DUAL_TREE_CHROMA, setting the threshold value of the palette mode block size according to a chroma format.
2. The method according to claim 1, wherein the step of decrypting the encoding unit from the bitstream according to the palette mode activation flag includes: when the palette mode activation flag indicates that the palette mode is activated for the encoding unit, generating a palette table for the encoding unit from the bitstream; and decrypting the encoding unit from the bitstream using the generated palette table.
3. The method according to claim 2, further comprising determining whether an escape encoded sample is within the encoding unit from an escape flag.
4. The method according to claim 1, wherein when the chroma format is 4:4:4, the threshold value of the palette mode block size is 16.
5. The method according to claim 1, wherein when the color difference format is 4:2:2, the threshold value of the palette mode block size is 32.
6. The method according to claim 1, wherein when the color difference format is 4:2:0, the threshold value of the palette mode block size is 64.
7. An electronic device, one or more processing units, a memory connected to the one or more processing units, and a plurality of programs stored in the memory and comprising, when the plurality of programs are executed by the one or more processing units, causing the electronic device to perform the method according to any one of claims 1 to 6.
8. A non-transitory computer-readable recording medium storing a plurality of programs for execution by an electronic device having one or more processing units, wherein when the plurality of programs are executed by the one or more processing units, causing the electronic device to perform the method according to any one of claims 1 to 6.
9. A computer program including instructions for execution by a computer device having one or more processors, wherein when the instructions are executed by the one or more processors, causing the computer device to perform the method according to any one of claims 1 to 6.
10. A method of transmitting a bitstream, generating the bitstream by executing an encoding method, transmitting the bitstream to a decoder and comprising, wherein the encoding method comprises: deriving a plurality of variables related to an encoding unit, the plurality of variables indicating the size of a luminance encoding block and an encoding tree type, the encoding unit being related to a predetermined splitting method, the predetermined splitting method including 4-way splitting, horizontal 3-way splitting, vertical 3-way splitting, horizontal 2-way splitting or vertical 2-way splitting; determining a threshold value of the palette mode block size of the encoding unit according to the encoding tree type; in response to a determination that at least the size of the luminance encoding block is larger than the threshold value of the palette mode block size, determining a palette mode enabling flag related to the encoding unit, encoding the encoding unit according to the palette mode activation flag including the step of determining a threshold value of a palette mode block size of the encoding unit according to the encoding tree type when the encoding tree type is SINGLE_TREE or DUAL_TREE_LUMA setting the threshold value of the palette mode block size to 16 when the encoding tree type is DUAL_TREE_CHROMA a method including the step of setting the threshold value of the palette mode block size according to a color difference format A method according to claim 10, wherein when the color difference format is 4:4:4, the threshold value of the palette mode block size is 16 A method according to claim 10, wherein when the color difference format is 4:2:2, the threshold value of the palette mode block size is 32 A method according to claim 10, wherein when the color difference format is 4:2:0, the threshold value of the palette mode block size is 64
Citation Information
Patent Citations
Restriction on palette block size in video coding
US20160234494A1