Increase the decoding throughput of blocks decoded intra-frame

By applying image size limitations in video encoding and decoding, the problems of increased decoding complexity and reduced parallelism caused by too small block size in the prior art are solved, and higher parallel decoding capabilities and lower complexity are achieved.

CN113994674BActive Publication Date: 2025-06-24QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080044006.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-18
Filing Date
2020-06-19
Publication Date
2025-06-24
Estimated Expiration
2040-06-19

AI Technical Summary

Technical Problem

When processing video data blocks, it is difficult to effectively avoid splitting resulting in too small block size, resulting in increased decoding complexity and reduced parallelism.

Method used

By applying image size limits, ensure that the image size of the video data block is 8 and the maximum multiple of the minimum decoding unit size, preventing smaller chromatic blocks at the corners of the picture, thereby reducing block dependence and improving parallelism.

Benefits of technology

It effectively avoids splitting with too small block size, improves the parallel decoding capability of video data blocks, and has little impact on decoding accuracy and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113994674B_ABST
    Figure CN113994674B_ABST
Patent Text Reader

Abstract

A method for decoding video data, comprising: determining, by one or more processors implemented in a circuit, a picture size of a picture. The picture size applies a picture size limit to determine a width of the picture and a height of the picture as respective multiples of a maximum value between 8 and a minimum decoding unit size for the picture. The method further comprises: determining, by one or more processors, a partitioning for partitioning the picture into a plurality of blocks, and generating, by one or more processors, a predicted block for a block among the plurality of blocks. The method further comprises: decoding, by one or more processors, a residual block for the block, and combining, by one or more processors, the predicted block and the residual block to decode the block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 16 / 905,467, filed on Jun. 18, 2020, which claims the benefit of U.S. Provisional Application No. 62 / 864,855, filed on Jun. 21, 2019. The entire content of each application is incorporated herein by reference. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital live systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones (so-called “smart phones”), video teleconferencing devices, video streaming devices, etc. Digital video devices implement video decoding technologies, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions to these standards. Video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information by implementing such video decoding technologies.

[0004] Video decoding technologies include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. For block-based video decoding, a video slice (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture can be encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] In general, the present disclosure describes techniques for processing video data blocks (e.g., smaller intra-coded blocks). A video encoder may be configured to partition video data into multiple blocks. For example, instead of processing a large block having 64x64 samples (e.g., pixels), the video encoder may split the block into two or more smaller blocks, such as, for example, four 32x32 blocks, sixteen 16x16 blocks, or other block sizes. In some examples, the video encoder may be configured to split the block into relatively small sizes (e.g., 2x2 blocks, 2x4 blocks, 4x2 blocks, etc.). Similarly, a video decoder may be configured to determine a partitioning of the video data into multiple blocks.

[0006] According to example techniques of the present disclosure, a video decoder (e.g., a video encoder or a video decoder) may determine a picture size that applies a picture size limit to reduce or eliminate smaller chrominance blocks (e.g., 2x2 chrominance blocks, 2x4 chrominance blocks, 4x2 chrominance blocks) at the lower right corner of a picture. That is, during a partitioning of a picture of video data for splitting a larger video data block into smaller blocks, the picture size limit may prevent one or more splittings that would result in a relatively small block size at a corner of the picture (e.g., a video picture, a stripe of a video picture, or other video data). For example, a video encoder may limit the picture size of video data. In some examples, a video decoder may determine a picture size that applies the picture size limit. After partitioning the picture, the video decoder may determine a predicted block for a block of the picture. The predicted block may depend on neighboring blocks. For example, the video decoder may determine a predicted block for a block based on an upper neighboring block and a left neighboring block. By preventing splittings that result in a relatively small block size at a corner of the picture, the video decoder may determine predicted blocks for blocks of a picture of video data with fewer block dependencies, thereby potentially increasing the parallelism for coding (e.g., encoding or decoding) the blocks of the picture with little loss in prediction accuracy and / or complexity.

[0007] In one example, a method of decoding video data includes: determining, by one or more processors implemented in a circuit, a picture size of a picture, wherein the picture size applies a picture size limit to set each of a width and a height of the picture to a respective multiple of a maximum of 8 and a minimum coding unit size for the picture; determining, by the one or more processors, a partitioning of the picture into multiple blocks; generating, by the one or more processors, a predicted block for a block of the multiple blocks; decoding, by the one or more processors, a residual block for the block; and merging, by the one or more processors, the predicted block and the residual block to decode the block.

[0008] In another example, a method of encoding video data includes: setting, by one or more processors implemented in a circuit, a picture size of a picture, wherein setting the picture size includes applying a picture size limit to set each of a width of the picture and a height of the picture to a respective multiple of 8 and a maximum value of a minimum decoding unit size for the picture; dividing, by the one or more processors implemented in the circuit, the picture into a plurality of blocks; generating, by the one or more processors, a predicted block for a block among the plurality of blocks; generating, by the one or more processors, a residual block for the block based on a difference between the block and the predicted block; and encoding, by the one or more processors, the residual block.

[0009] In one example, a device for decoding video data includes one or more processors implemented in a circuit and configured to: determine a picture size of a picture, wherein the picture size applies a picture size limit to set each of a width of the picture and a height of the picture to a respective multiple of 8 and a maximum value of a minimum decoding unit size for the picture; determine a division for dividing the picture into a plurality of blocks; generate a predicted block for a block among the plurality of blocks; decode a residual block for the block; and combine the predicted block and the residual block to decode the block.

[0010] In another example, a device for encoding video data includes one or more processors implemented in a circuit and configured to: set a picture size of a picture, wherein, to set the picture size, the one or more processors are configured to apply a picture size limit to set each of a width of the picture and a height of the picture to a respective multiple of 8 and a maximum value of a minimum decoding unit size for the picture; divide the picture into a plurality of blocks; generate a predicted block for a block among the plurality of blocks; generate a residual block for the block based on a difference between the block and the predicted block; and encode the residual block.

[0011] In one example, a device for decoding video data includes: a unit for determining a picture size of a picture, wherein the picture size applies a picture size limit to set each of a width of the picture and a height of the picture to a respective multiple of 8 and a maximum value of a minimum decoding unit size for the picture; a unit for determining a division for dividing the picture into a plurality of blocks; a unit for generating a predicted block for a block among the plurality of blocks; a unit for decoding a residual block for the block; and a unit for combining the predicted block and the residual block to decode the block.

[0012] In another example, an apparatus for encoding video data includes: a unit for setting a picture size of a picture, wherein the unit for setting the picture size includes a unit for applying a picture size limit to set each of a width and a height of the picture to a respective multiple of a maximum value between 8 and a minimum decoding unit size for the picture; a unit for dividing the picture into a plurality of blocks; a unit for generating a prediction block for a block among the plurality of blocks; a unit for generating a residual block for the block based on a difference between the block and the prediction block; and a unit for encoding the residual block.

[0013] In one example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors to perform operations as follows: determining a picture size of a picture, wherein the picture size applies a picture size limit to set each of a width and a height of the picture to a respective multiple of a maximum value between 8 and a minimum decoding unit size for the picture; determining a partitioning for dividing the picture into a plurality of blocks; generating a prediction block for a block among the plurality of blocks; decoding a residual block for the block; and combining the prediction block and the residual block to decode the block.

[0014] In another example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors to perform operations as follows: setting a picture size of a picture, wherein, to set the picture size, the instructions cause the one or more processors to apply a picture size limit to set each of a width and a height of the picture to a respective multiple of a maximum value between 8 and a minimum decoding unit size for the picture; dividing the picture into a plurality of blocks; generating a prediction block for a block among the plurality of blocks; generating a residual block for the block based on a difference between the block and the prediction block; and encoding the residual block.

[0015] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may execute the techniques of the present disclosure.

[0017] Figure 2A and 2B is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and a corresponding coding tree unit (CTU).

[0018] Figures 3A - 3E is a conceptual diagram illustrating multiple example tree splitting patterns.

[0019] Figure 4 is a conceptual diagram showing an example reference sample array for intra prediction of chrominance components.

[0020] Figure 5 is a conceptual diagram showing examples of expanding a 2x2 block to a 4x4 block, expanding a 2x4 block to a 4x4 block, and expanding a 4x2 block to a 4x4 block.

[0021] Figure 6 is a block diagram showing an example video encoder that can perform the techniques of the present disclosure.

[0022] Figure 7 is a block diagram showing an example video decoder that can perform the techniques of the present disclosure.

[0023] Figure 8 is a flowchart showing an example method for encoding a current block.

[0024] Figure 9 is a flowchart showing an example method for decoding a current block of video data.

[0025] Figure 10 is a flowchart showing an example process of using picture size limitation according to the techniques of the present disclosure. Detailed Description

[0026] Generally, the present disclosure describes techniques for processing blocks of video data (e.g., intra-coded blocks). In an example of the present disclosure, a video encoder can be configured to partition video data into multiple blocks. For example, instead of processing a large block having 64x64 samples (e.g., pixels), the video encoder can split the block into two or more smaller blocks, such as, for example, four 32x32 blocks, sixteen 16x16 blocks, or other block sizes. In some examples, the video encoder can be configured to split the block into relatively small sizes (e.g., 2x2 blocks, 2x4 blocks, 4x2 blocks, etc.). For example, the video encoder can split a 16x8 block into two 8x8 blocks. Similarly, a video decoder can be configured to determine the partitioning of the video data into multiple blocks.

[0027] To reduce the decoding complexity and with little or no loss in decoding accuracy, a video decoder (e.g., a video encoder or a video decoder) may be configured to use a luminance component to represent the luminance of a video data block and a chrominance component to represent the color characteristics of the video data block. The chrominance component may include a blue minus luminance value ("Cb") and / or a red minus luminance value ("Cr"). For example, a video decoder (e.g., a video encoder or a video decoder) may be configured to represent an 8x8 block by an 8x8 luminance block of the luminance component (e.g., "Y"), a first 4x4 chrominance block of the chrominance component (e.g., "Cr"), and a second 4x4 chrominance block of the chrominance component (e.g., "Cb"). That is, the chrominance component of the video data block may be downsampled to have fewer samples than the luminance component of the video data block. In this way, downsampling the chrominance component can improve the decoding efficiency with little or no loss in decoding accuracy.

[0028] A video decoder (e.g., a video encoder or a video decoder) may be configured for a block to be intra-coded, where the prediction block depends on other blocks. For example, the video decoder may use an upper adjacent block and a left adjacent block to predict the current block to improve the decoding accuracy. Thus, the video decoder may not predict the current block in parallel with the prediction of the upper adjacent block and the left adjacent block. Instead, the video decoder may wait to predict the current block until the prediction of the upper adjacent block and the left adjacent block is completed. Block dependency may increase the decoding complexity, which increases as the block size decreases.

[0029] According to the techniques of the present disclosure, a video decoder (e.g., a video encoder or a video decoder) may apply a picture size limit to prevent splitting that results in a relatively small block size. As used herein, splitting may refer to dividing a block into smaller blocks. For example, the video decoder may be configured to apply a picture size limit to prevent splitting of a picture that would result in a smaller chrominance block at a corner of the picture (e.g., the lower right corner). Applying the picture limit may help improve the decoding parallelism of the decoded blocks while having no or little impact on the decoding accuracy and / or complexity.

[0030] After partitioning or splitting the video data, a video decoder (e.g., a video encoder or a video decoder) may generate prediction information for the blocks of a picture and determine a prediction block for the block based on the prediction information. Similarly, in the case of intra prediction, the prediction block may depend on adjacent blocks. For example, the video decoder may determine a prediction block for a block based on an upper adjacent block and a left adjacent block. By preventing splitting that results in a relatively small block size, the video decoder may determine the prediction information for the blocks of the picture of the video data with less block dependency, and thus potentially increase the number of blocks that can be decoded (e.g., encoded or decoded) in parallel with little or no loss in prediction accuracy and / or complexity.

[0031] Figure 1 is a block diagram showing an example video encoding and decoding system 100 that can implement the techniques of the present disclosure. The techniques of the present disclosure generally relate to decoding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0032] As Figure 1 shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 can include any of a variety of devices, including desktop computers, notebooks (i.e., laptops) computers, tablet computers, set-top boxes, telephone handheld devices such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and destination device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.

[0033] In Figure 1 the example, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to the present disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for extending chrominance split limits in a shared tree configuration, limiting picture size, and / or processing 2x2, 2x4, or 4x2 chrominance blocks at a corner of a picture. Thus, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, source device and destination device can include other components or arrangements. For example, source device 102 can receive video data from an external video source (such as an external camera). Similarly, destination device 116 can interface with an external display device instead of including an integrated display device.

[0034] As Figure 1The system 100 shown is merely an example. In general, any digital video encoding and / or decoding device can perform techniques for extending chroma split limits in a shared tree configuration, limiting picture size, and / or processing 2x2, 2x4, or 4x2 chroma blocks at a corner of a picture. The source device 102 and the destination device 116 are merely examples of such decoding devices, where the source device 102 generates encoded video data for transmission to the destination device 116. This disclosure refers to a "decoding" device as a device that performs data decoding (encoding and / or decoding). Thus, the video encoder 200 and the video decoder 300 represent examples of decoding devices, specifically, examples of a video encoder and a video decoder, respectively. In some examples, the devices 102 and 116 can operate in a substantially symmetric manner such that each of the devices 102 and 116 includes video encoding and decoding components. Accordingly, the system 100 can support one-way or two-way video transmission between the video devices 102 and 116, e.g., for video streaming, video playback, video broadcast, or video telephony.

[0035] In general, the video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequence of consecutive pictures (also referred to as "frames") of video data to the video encoder 200, where the video encoder 200 encodes the data for the pictures. The video source 104 of the source device 102 can include a video capture device, such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 104 can generate computer graphics-based data as the source video, or generate a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 can reorder pictures from the received order (sometimes referred to as "display order") into a decoding order for decoding. The video encoder 200 can generate a bitstream that includes the encoded video data. The source device 102 can then output the encoded video data via the output interface 108 onto a computer-readable medium 110 for reception and / or retrieval, e.g., by the input interface 122 of the destination device 116.

[0036] The memories 106 of the source device 102 and 120 of the destination device 116 represent general memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from the video source 104 and decoded raw video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable, e.g., by the video encoder 200 and the video decoder 300, respectively. Although in this example the memories 106 and 120 are shown separately from the video encoder 200 and the video decoder 300, the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store encoded video data, e.g., the encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, e.g., for storing raw video data, decoded video data, and / or encoded video data.

[0037] The computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to directly send the encoded video data to the destination device 116 in real time, e.g., via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, the output interface 108 may modulate the transmission signal including the encoded video data, and the input interface 122 may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other devices that may be used to facilitate communication from the source device 102 to the destination device 116.

[0038] In some examples, the computer-readable medium 110 may include a storage device 112. The source device 102 may output the encoded data from the output interface 108 to the storage device 112. Similarly, the destination device 116 may access the encoded data from the storage device 112 via the input interface 122. The storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.

[0039] In some examples, the computer-readable medium 110 can include a file server 114 or another intermediate storage device that can store the encoded video data generated by the source device 102. The source device 102 can output the encoded video data to the file server 114 or another intermediate storage device, which can store the encoded video generated by the source device 102. The destination device 116 can access the stored video data from the file server 114 via streaming or downloading. The file server 114 can be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 can access the encoded video data from the file server 114 via any standard data connection including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both, which is suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 can be configured to operate according to a streaming transport protocol, a download transport protocol, or a combination thereof.

[0040] The output interface 108 and the input interface 122 can represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 can be configured to transmit data, such as encoded video data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 can be configured to transmit data, such as encoded video data, according to other wireless standards such as the IEEE 802.11 specifications, the IEEE 802.15 specifications (e.g., ZigBee TM )、Bluetooth TM standards, etc. In some examples, the source device 102 and / or the destination device 116 can include corresponding System-on-Chip (SoC) devices. For example, the source device 102 can include an SoC device for performing the functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 can include an SoC device for performing the functions attributed to the video decoder 300 and / or the input interface 122.

[0041] The techniques of the present disclosure can be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (e.g., Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0042] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream can include signaling information defined by the video encoder 200 (which is also used by the video decoder 300), such as syntax elements having values for describing the characteristics and / or processing of video blocks or other decoded units (e.g., strips, pictures, groups of pictures, sequences, etc.). The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 can represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0043] Although not shown in Figure 1 In some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder and can include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream including audio and video in a common data stream. When applicable, the MUX-DEMUX unit can comply with the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).

[0044] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder circuits and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular telephones.

[0045] The video encoder 200 and the video decoder 300 can operate according to a video coding standard, such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or an extension thereof, such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder 200 and the video decoder 300 can operate according to other proprietary or industry standards, such as ITU-T H.266, also known as Versatile Video Coding (VVC). A draft of the VVC standard is described in Bross et al., “Versatile Video Coding (Draft 8)” (presented at the 17th meeting of the Joint Video Team (JVT) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Brussels, Belgium, January 7-17, 2020, JVET-Q2001-vA (hereinafter referred to as “VVC Draft 8”)). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0046] Generally, video encoder 200 and video decoder 300 may perform block - based decoding of pictures. The term "block" generally refers to a structure including data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two - dimensional matrix of samples of luminance data and / or chrominance data. Generally, video encoder 200 and video decoder 300 may decode video data represented in the YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 may decode luminance components and chrominance components, rather than decoding red, green, and blue (RGB) data of the samples for a picture, where the chrominance components may include a red chrominance component and a blue chrominance component. In some examples, video encoder 200 converts the received RGB - format data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to the RGB format. Alternatively, a pre - processing unit and a post - processing unit (not shown) may perform these conversions.

[0047] This disclosure generally relates to decoding (e.g., encoding and decoding) of pictures to include processes regarding encoding or decoding data of pictures. Similarly, this disclosure may relate to decoding of blocks of pictures to include processes regarding encoding or decoding data for blocks, e.g., predictive decoding and / or residual decoding. An encoded video bitstream generally includes a series of values for syntax elements used to represent decoding decisions (e.g., decoding modes) and partitioning of a picture into blocks. Thus, referring to decoding a picture or a block should generally be understood as decoding values of the syntax elements forming the picture or the block.

[0048] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (e.g., video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoder divides the CTU and CUs into four equal, non - overlapping squares, and each node of the quadtree has zero or four children. A node without children may be referred to as a "leaf node", and a CU of such a leaf node may include one or more PUs and / or one or more TUs. The video decoder may further divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the division of TUs. In HEVC, a PU represents inter - frame prediction data, and a TU represents residual data. An intra - predicted CU includes intra - prediction information, such as an intra - mode indication.

[0049] As another example, video encoder 200 and video decoder 300 may be configured to operate according to JEM or VVC. According to JEM or VVC, a video coder (such as video encoder 200) divides a picture into multiple coding tree units (CTUs). Video encoder 200 may divide a CTU according to a tree structure (such as a quadtree binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple partitioning types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level that is divided according to quadtree partitioning, and a second level that is divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0050] In the MTT partitioning structure, quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning may be used to partition a block. Ternary tree partitioning is a partitioning for splitting a block into three sub-blocks. In some examples, ternary tree partitioning splits a block into three sub-blocks, where the original block is not split through the center. The partitioning types in MTT (such as QT, BT, and TT) may be symmetric or asymmetric.

[0051] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance component and the chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for the respective chrominance components).

[0052] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures according to HEVC. For purposes of explanation, the description of the techniques of the present disclosure is presented with respect to QTBT partitioning. However, the techniques of the present disclosure may also be applied to video coders configured to use quadtree partitioning or other types of partitioning.

[0053] The present disclosure may interchangeably use "NxN" and "N×N" to denote the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions. For example, 16x16 samples or 16×16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non - negative integer value. The samples in a CU can be arranged in rows and columns. Additionally, a CU does not have to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include NxM samples, where M does not necessarily equal N.

[0054] Video encoder 200 encodes video data for a CU representing prediction information and / or residual information and other information. The prediction information indicates how to predict the CU to form a prediction block for the CU. The residual information generally represents the sample - by - sample difference between the samples of the CU before encoding and the prediction block.

[0055] To predict a CU, video encoder 200 can generally form a prediction block for the CU by inter - frame prediction or intra - frame prediction. Inter - frame prediction generally refers to predicting the CU based on the data of previously decoded pictures, while intra - frame prediction generally refers to predicting the CU based on the previously encoded data of the same picture. To perform inter - frame prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, for example, in terms of the difference between the CU and the reference block. Video encoder 200 can use the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to compute a difference metric to determine whether the reference block closely matches the current CU. In some examples, video encoder 200 can use uni - directional prediction or bi - directional prediction to predict the current CU.

[0056] Some examples of JEM and VVC also provide an affine motion compensation mode, which can be regarded as an inter - frame prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors representing non - translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0057] To perform intra prediction, the video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of JEM and VVC provide 67 intra prediction modes, including various directional modes, as well as planar mode and DC mode. Generally, the video encoder 200 selects an intra prediction mode that describes neighboring samples of the current block (e.g., the block of a CU), where samples of the current block are predicted from the neighboring samples. Assuming that the video encoder 200 decodes CTUs and CUs in raster scan order (from left to right, top to bottom), such samples can typically be above, above and to the left, or to the left of the current block in the same picture as the current block.

[0058] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for an inter prediction mode, the video encoder 200 may encode data representing which one of the various available inter prediction modes is used and the motion information for the corresponding mode. For example, for uni-directional inter prediction or bi-directional inter prediction, the video encoder 200 may encode the motion vector using advanced motion vector prediction (AMVP) or merge mode. The video encoder 200 may use a similar mode to encode the motion vector for an affine motion compensation mode.

[0059] After prediction such as intra prediction or inter prediction of a block, the video encoder 200 may compute residual data for the block. Residual data such as a residual block represents the sample-by-sample difference between the block and the prediction block for the block formed using the corresponding prediction mode. The video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 may apply a second transform after the first transform, such as a mode-dependent non-separable second-order transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0060] As described above, after any transform used to generate transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a right shift of the bits of the value to be quantized.

[0061] After quantization, the video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and thus lower frequency) coefficients at the front of the vector and lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, the video encoder 200 may utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then perform entropy coding on the quantized transform coefficients of the vector. In other examples, the video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 may perform entropy coding on the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). The video encoder 200 may also perform entropy coding on the values of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0062] To perform CABAC, the video encoder 200 may assign a context within a context model to the symbol to be sent. For example, the context may relate to whether adjacent values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0063] The video encoder 200 may also generate syntax data for the video decoder 300, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, in, for example, a picture header, a block header, a slice header, or other syntax data (such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS)). The video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.

[0064] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, e.g., syntax elements for describing the partitioning of a picture into blocks (e.g., CUs) and prediction information and / or residual information for the blocks. Eventually, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0065] Generally, the video decoder 300 performs a process reciprocal to the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values of the syntax elements for the bitstream in a manner substantially similar (although reciprocal) to the CABAC encoding process of the video encoder 200. The syntax elements can define: partitioning information for partitioning a picture into CTUs, and partitioning each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. The syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0066] The residual information can be represented by, e.g., quantized transform coefficients. The video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of a block to regenerate the residual block for the block. The video decoder 300 uses the signaling prediction mode (intra prediction or inter prediction) and associated prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Then, the video decoder 300 can (on a sample-by-sample basis) combine the prediction block and the residual block to regenerate the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the blocks.

[0067] To improve decoding accuracy, a video decoder (e.g., video encoder 200 or video decoder 300) may partition data blocks. For example, the video decoder may use quadtree splitting, binary tree splitting, or another splitting to partition the blocks. The video decoder (e.g., video encoder 200 or video decoder 300) may determine a single tree for video data (e.g., a video data stripe) based on the luminance component of the block. For example, a block may be represented by an 8x8 luminance block (e.g., "Y"), a first 4x4 chrominance block (e.g., "Cr"), and a second 4x4 chrominance block (e.g., "Cb"). In this example, the video decoder may generate a single tree to split the block, thereby splitting the 8x8 luminance block into two 4x4 luminance blocks. The video decoder may split the first 4x4 chrominance block (e.g., "Cr") into two 2x2 chrominance blocks and split the second 4x4 chrominance block (e.g., "Cb") into two 2x2 chrominance blocks according to the single tree. In this way, the video decoder may improve the accuracy of the predicted blocks generated for the block, which may improve the prediction accuracy of the video data.

[0068] However, when partitioning video data blocks (e.g., intra-coded blocks), the video decoder (e.g., video encoder 200 or video decoder 300) may split the block (e.g., the chrominance component of the block, referred to herein as the "chrominance block", and / or the luminance component of the block, referred to herein as the "luminance block") into smaller blocks (e.g., 2x2 blocks, 2x4 blocks, 4x2 blocks, etc.). Additionally, each of the smaller blocks may have a decoding dependency on adjacent blocks. For example, the video decoder may use samples from one or more adjacent blocks (e.g., the left adjacent block and / or the upper adjacent block) to determine the predicted block for each of the smaller blocks. Thus, the smaller blocks and the data dependencies may cause the video decoder to sequentially determine the predicted blocks for each of the smaller blocks, which may result in lower decoding parallelism.

[0069] According to example techniques of the present disclosure, a video decoder (e.g., video encoder 200) may be configured to set each of a width of a picture and a height of the picture to a respective multiple of the maximum of 8 and a minimum coding unit size for the picture. For example, video encoder 200 may be configured to calculate the width of the picture as X1*N, where X1 is a first multiple and N = max(8, minCuSize), and minCuSize is a minimum decoding unit value. The video encoder may be configured to calculate the height of the picture as X2*N, where X2 is a second multiple. That is, if 8 is greater than minCuSize, the picture width is limited to X1*8 and the picture height is limited to X2*8. However, if minCuSize is greater than 8, the picture width is limited to X1*minCuSize and the picture height is limited to X2*minCuSize. In this way, the video encoder can set the width and height of the picture to help ensure that the bottom-right block of the picture includes a width of at least 8 samples and a height of at least 8 samples (i.e., at least an 8x8 block), which can result in a chrominance block for the bottom-right block that includes a size of at least 4x4 when the video encoder applies a color format that down-samples the chrominance block (e.g., 4:2:2 or 4:2:0), and a size of at least 8x8 when the video encoder does not apply a color format that down-samples the chrominance block.

[0070] A video decoder (e.g., video decoder 300) may be configured to determine a picture size, where the picture size applies a picture size limit. The picture size limit may set each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and a minimum decoding unit size for the picture. For example, the video decoder may be configured to decode one or more partitioning syntax elements that indicate the picture size of the picture and a partitioning regarding partitioning the picture into multiple blocks. In this example, the picture size may help ensure that: the video decoder identifies a bottom-right block of the picture that includes a width of at least 64 luma samples (or 16 chroma samples), which may result in a chrominance block of the bottom-right block that includes a size of at least 16 samples. That is, the video decoder may not decode further split flags or other partitioning syntax elements for splitting the bottom-right block of the picture into less than 8x8. In this way, the video decoder can determine the picture size and determine the partitioning of the picture to help ensure that the bottom-right block of the picture includes a width of at least 8 samples and a height of at least 8 samples (i.e., at least an 8x8 block), which can result in a chrominance block for the bottom-right block that includes a size of at least 4x4 when the video encoder applies a color format that down-samples the chrominance block (e.g., 4∶2∶2 or 4∶2∶0), and a size of at least 8x8 when the video encoder does not apply a color format that down-samples the chrominance block.

[0071] After partitioning video data, a video decoder (e.g., video encoder 200 or video decoder 300) may generate prediction information for a block and determine a predicted block for the block based on the prediction information. Similarly, the predicted block may depend on neighboring blocks. For example, the video decoder may determine a predicted block for a current block based on an upper neighboring block and a left neighboring block. By preventing block splitting (e.g., for chrominance components and / or luminance components) that results in relatively small block sizes, the video decoder may determine prediction information for video data blocks with fewer block dependencies, thereby potentially increasing the number of blocks that can be decoded in parallel (e.g., encoded or decoded) with little loss in prediction accuracy and / or complexity.

[0072] The present disclosure may generally refer to "signaling" certain information, such as a syntax element. The term "signaling" generally may refer to the conveyance of a value for a syntax element and / or other data that is used to decode encoded video data. That is, video encoder 200 may signal a value for a syntax element in a bitstream. Generally, signaling refers to generating a value in the bitstream. As described above, source device 102 may transmit the bitstream to destination device 116 substantially in real time or not in real time, such as may occur when storing the syntax elements to storage device 112 for later retrieval by destination device 116.

[0073] Figure 2A and 2B is a conceptual diagram showing an example quadtree binary tree (QTBT) structure 130 and corresponding coding tree units (CTUs) 132. Solid lines represent quadtree splits and dashed lines represent binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type (i.e., horizontal or vertical) was used, where, in this example, 0 represents a horizontal split and 1 represents a vertical split. For a quadtree split, no indication of the split type is needed because a quadtree node splits a block horizontally and vertically into 4 equal-sized sub-blocks. Accordingly, video encoder 200 may encode and video decoder 300 may decode syntax elements (e.g., split information) for the region tree level (i.e., solid lines) in QTBT structure 130 and syntax elements (e.g., split information) for the prediction tree level (i.e., dashed lines) in QTBT structure 130. Video encoder 200 may encode and video decoder 300 may decode video data for CUs represented by terminal leaf nodes in QTBT structure 130, such as prediction data and transform data.

[0074] Generally, Figure 2BThe CTU 132 can be associated with parameters for defining the size of blocks corresponding to nodes in the QTBT structure 130 at the first and second levels. These parameters can include the CTU size (representing the size of the CTU 132 in the sample), the minimum quadtree size (MinQTSize, representing the minimum quadtree leaf node size allowed), the maximum binary tree size (MaxBTSize, representing the maximum binary tree root node size allowed), the maximum binary tree depth (MaxBTDepth, representing the maximum binary tree depth allowed), and the minimum binary tree size (MinBTSize, representing the minimum binary tree leaf node size allowed).

[0075] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each child node can be divided according to the quadtree partitioning. That is, the nodes at the first level are either leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 shows that such nodes include a parent node and child nodes with solid lines for branching. If the nodes at the first level are not larger than the allowed maximum binary tree root node size (MaxBTSize), these nodes can be further divided by the corresponding binary tree. The binary tree splitting of a node can be iteratively performed until the nodes generated by the splitting reach the allowed minimum binary tree leaf node size (MinBTSize) or the allowed maximum binary tree depth (MaxBTDepth). An example of the QTBT structure 130 shows that such nodes have dashed lines for branching. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra-picture prediction or inter-picture prediction) and transformation without any further partitioning. As described above, the CU can also be referred to as a "video block" or a "block".

[0076] In an example of the QTBT partitioning structure, the CTU size is set to 128x128 (luma samples and two corresponding 64x64 chroma samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a quadtree leaf node is 128x128, it will not be further split by the binary tree because its size exceeds MaxBTSize (64x64 in this example). Otherwise, the quadtree leaf node will be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. When the width of a binary tree node is equal to MinBTSize (4 in this example), it means that further horizontal splitting is not allowed. Similarly, the height of a binary tree node being equal to MinBTSize means that no further vertical splitting is allowed for that binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.

[0077] Figures 3A - 3E is a conceptual diagram showing examples of multiple tree splitting patterns. Figure 3A shows quadtree partitioning, Figure 3B shows vertical binary tree partitioning, Figure 3C shows horizontal binary tree partitioning, Figure 3D shows vertical ternary tree partitioning, Figure 3E shows horizontal ternary tree partitioning.

[0078] In VVC WD5, the CTU is split into CUs by using a quadtree structure represented as a decoding tree to adapt to various local features. A decision is made at the leaf CU level whether to use inter-picture (temporal) or intra-picture (spatial) prediction to decode the picture region. Depending on the PU split type, each leaf CU can be further split into one, two, or four PUs. Within a PU, the same prediction process is applied, and relevant information is sent to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU split type, the leaf CU can be partitioned into transform units (TUs) according to another quadtree structure (such as the decoding tree for CUs). A feature of the HEVC structure is that the HEVC structure has multiple partitioning concepts, including CUs, PUs, and TUs.

[0079] In VVC, a quadtree with a nested multi-type tree that uses binary and ternary tree splitting partitioning structures replaces the concept of multiple partitioning unit types. That is, VVC removes the separation of the CU, PU, and TU concepts (except when needed for a CU that is too large in size for the maximum transform length) and supports greater flexibility for CU partitioning shapes. In the coding tree structure, a CU can be square or rectangular. A coding tree unit (CTU) is first partitioned by a quadtree (also known as a four-tree) structure. Then, the quadtree leaf nodes are further partitioned by a multi-type tree structure.

[0080] Figure 3A is a conceptual diagram showing an example of quadtree partitioning, including a vertical binary tree split 140 ("SPLIT_BT_VER") and a horizontal binary tree split 141 ("SPLIT_BT_HOR"). Figure 3B is a conceptual diagram showing an example of a vertical binary tree partition including a vertical binary tree split 142. Figure 3C is a conceptual diagram showing an example of a horizontal binary tree partition including a horizontal binary tree split 143. Figure 3D is a conceptual diagram showing an example of a vertical ternary tree partition including vertical ternary tree splits 144, 145 ("SPLIT_TT_VER"). Figure 3E is a conceptual diagram showing an example of a horizontal ternary tree partition including horizontal ternary tree splits 146, 147 ("SPLIT_TT_HOR").

[0081] The multi-type tree leaf nodes are called coding units (CUs), and this partitioning is used for prediction and transform processing without further partitioning, unless the CU is too large for the maximum transform length. This means that: in some cases, the CU, PU, and TU have the same block size in a quadtree with a nested multi-type tree coding block structure. Exceptions may occur when the supported maximum transform length is less than the width or height of the color component of the CU.

[0082] A CTU can include a luma coding tree block (CTB) and two chroma coding tree blocks. At the CU level, a CU is associated with a luma coding block (CB) and two chroma coding blocks. Just as in VTM (the reference software for VVC), the luma tree and chroma tree are separated within a slice (referred to as a dual-tree structure), while they are shared across slices (referred to as a single-tree or shared-tree structure). The size of a CTU can be up to 128x128 (luma component), while the size of a coding unit can range from 4x4 to the size of the CTU. In this case, the size of a chroma block can be 2x2, 2x4, or 4x2 in the 4:2:0 color format.

[0083] Figure 4It is a conceptual diagram showing a reference sample array for intra prediction of chrominance components. A video decoder (e.g., video encoder 200 or video decoder 300) may use samples near the decoding block 640 for intra prediction of the block. Generally, the video decoder uses the reconstructed reference sample lines closest to the left and upper boundaries of the decoding block 150 as reference samples for intra prediction. For example, the video decoder may use the reconstructed samples of the upper row 152 and / or the left row 154. However, VVC also enables other samples near the decoding block 150 to be used as reference samples (e.g., top-left, bottom-left, top-right). For example, the video decoder may use the reconstructed samples of the top-left pixel 158, the left-bottom row 160, and / or the right-upper row 156.

[0084] In VVC, a video decoder (e.g., video encoder 200 or video decoder 300) may use only the reference lines with MRLIdx equal to 0, 1, and 3 for the luminance component. For the chrominance component, the video decoder may use only the reference line with MRLIdx equal to 0, as Figure 4 shown. The video decoder may decode the index into a reference line that is used to decode the block (indicating the values 0, 1, and 2 of the lines with MRLIdx 0, 1, and 3 respectively) using truncated unary codewords. For reference lines with MRLIdx > 0, the video decoder may not use the planar mode and the DC mode. In some examples, the video decoder may add only the available samples near the decoding block to the reference array for intra prediction.

[0085] To improve the processing throughput of intra decoding, several methods have been proposed. In "CE3-related: Shared reference samples for multiple chroma intra CBs" (JVET-M0169) by Z.-Y. Lin, T.-D. Chuang, C.-Y. Chen, Y.-W. Huang, S.-M. Lei and "Non-CE3: Intra chroma partitioning and prediction restriction" (JVET-M0065) by T. Zhou, T. Ikai, smaller block sizes, e.g., 2x2, 2x4, and 4x2, are disabled in the dual tree. For the single tree, sharing reference samples for smaller blocks is proposed (JVET-M0169).

[0086] In some hardware video encoders and video decoders, when a picture has a large number of small blocks, the processing throughput is reduced. This reduction in processing throughput may be caused by the use of small intra blocks because inter blocks of small size can be processed in parallel while intra blocks have data dependencies between adjacent blocks (e.g., the generation of predictors for intra blocks requires reconstructed samples from the upper and left boundaries of adjacent blocks) and must be processed sequentially.

[0087] In HEVC, the worst - case processing throughput occurs when processing 4x4 chroma intra blocks. In VVC, the size of the smallest chroma intra block is 2x2, and due to the adoption of new tools, the reconstruction process of chroma intra blocks may become complex.

[0088] In U.S. patent application Ser. No. 16 / 813,508, “RECONSTRUCTION OF BLOCKS OF VIDEO DATA USING BLOCK SIZE RESTRICTION”, filed on Mar. 9, 2020, U.S. Provisional Patent Application Ser. No. 62 / 817,457, “ENABLING PARALLEL RECONSTRUCTION OF INTRA - CODED BLOCKS”, filed on Mar. 12, 2019, and U.S. Provisional Patent Application Ser. No. 62 / 824,688, “ENABLING PARALLEL RECONSTRUCTION OF INTRA - CODED BLOCKS”, filed on Mar. 27, 2019, several techniques for increasing the worst - case throughput have been proposed, each of which is incorporated by reference. In these patent applications, there are generally three main methods, including: removing intra prediction dependencies, intra prediction mode restrictions, and restrictions on chroma splitting that results in smaller blocks. In particular, for chroma splitting restrictions in a shared - tree configuration, the chroma block can be not split while splitting the corresponding luma block region.

[0089] In VVC, the picture width and height of luma are restricted to be multiples of the minimum coding unit size, and the minimum luma coding unit size can be 4.

[0090] pic_width_in_luma_samples specifies the width of each decoded picture in luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0091] pic_height_in_luma_samples specifies the height of each decoded picture in luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0092] In this case, 4x4 or 4x8 or 8x4 luma regions may appear at the corners of the picture. In other words, in this case, there may be 2x2 or 2x4 or 4x2 chroma blocks at the lower right corner of the picture.

[0093] The present disclosure presents techniques for extending chroma split limits in a shared tree configuration to enable parallel processing of smaller intra-coded blocks and thereby increase processing throughput. In addition, several methods for processing 2x2 and 2x4 chroma blocks at the corners of a picture are presented.

[0094] The present disclosure extends the chroma split limits in a shared tree configuration presented in U.S. Provisional Patent Application No. 62 / 817,457, "ENABLING PARALLEL RECONSTRUCTION OF INTRA-CODED BLOCKS", filed on March 12, 2019, and U.S. Provisional Patent Application No. 62 / 824,688, "ENABLING PARALLEL RECONSTRUCTION OF INTRA-CODED BLOCKS", filed on March 27, 2019.

[0095] For example, a video decoder (e.g., video encoder 200 or video decoder 300) may limit the modes for non-split chroma blocks. That is, in some examples, the video decoder may be configured to apply a shared tree (also referred to herein as a "single tree"). In this example, a first luma block (e.g., an 8x8 luma block) may be split into second luma blocks (e.g., four 4x4 luma blocks) according to the shared tree. However, due to chroma split limits (e.g., non-split chroma blocks), the corresponding chroma block (e.g., a 4x4 chroma block) may not be further split.

[0096] When all blocks in a corresponding luminance region are intra-coded, the video coder may force the chrominance blocks to be intra-coded. For example, the video coder (e.g., video encoder 200 or video decoder 300) may determine that the chrominance blocks are not split due to chrominance splitting restrictions (e.g., 4x4 chrominance blocks) and that all corresponding luminance blocks (e.g., four 4x4 luminance blocks) are intra-coded. In response to determining that the chrominance blocks are not split due to chrominance splitting restrictions and that all corresponding luminance blocks are intra-coded, the video coder may force the chrominance blocks (e.g., 4x4 chrominance blocks) to be intra-coded. In this way, the video coder may reduce the complexity of the video coder.

[0097] In some examples, when a corresponding luminance region contains both inter and intra (including intra and IBC modes), the video coder (e.g., video encoder 200 or video decoder 300) may be configured to encode chrominance using a default mode. For example, the video coder (e.g., video encoder 200 or video decoder 300) may determine that the chrominance blocks are not split due to chrominance splitting restrictions (e.g., 4x4 chrominance blocks) and that all corresponding luminance blocks (e.g., four 4x4 luminance blocks) are intra-coded. In response to determining that the chrominance blocks are not split due to chrominance splitting restrictions and that all corresponding luminance blocks are intra-coded, the video coder may force the chrominance blocks (e.g., 4x4 chrominance blocks) to be decoded in the default mode. For example, the default mode may be intra-coding. In some examples, the default mode may be inter-coding. In this example, the motion vector of the chrominance block may be the motion vector between the luminance blocks, or the average motion vector between the luminance blocks. When using the default mode, the video encoder (e.g., video encoder 200) may not signal an indication of the default mode, and the video decoder (e.g., video decoder 300) may use the configuration data to infer that the decoding mode is the default mode.

[0098] The video coder (e.g., video encoder 200 or video decoder 300) may limit the picture size to avoid 2x2 and 2x4 / 4x2 chrominance blocks at the corners of the picture. That is, the video coder may be configured to determine that the chrominance blocks at the bottom right corner of the picture include a size of at least 4x4. In some examples, the video coder may be configured to determine that the luminance blocks at the bottom right corner of the picture include a size of at least 64 pixels.

[0099] For example, the block width and height of a picture can be multiples of the maximum of 8 and the minimum CU size. For example, the video encoder 200 can be configured to calculate the width of a picture as X1*N, where X1 is the first multiple and N = max(8, minCuSize), and minCuSize is the minimum decoding unit value. The video encoder can be configured to calculate the height of the picture as X2*N, where X2 is the second multiple. That is, if 8 is greater than minCuSize, the picture width is limited to X1*8, and the picture height is limited to X2*8. However, if minCuSize is greater than 8, the picture width is limited to X1*minCuSize, and the picture height is limited to X2*minCuSize.

[0100] That is, the video encoder (e.g., video encoder 200) can be configured to apply a picture size limit to set each of the width and height of the picture to a respective multiple of the maximum of 8 and the minimum decoding unit size (e.g., MinCbSizeY) for the picture. The minimum decoding size for the picture can include the minimum width of the decoding unit for the picture or the minimum height of the decoding unit for the picture. The video decoder (e.g., video decoder 300) can be configured to determine the picture size of the picture, where the picture size applies the picture size limit to set each of the width and height of the picture to a respective multiple of the maximum of 8 and the minimum decoding unit size for the picture. Although the previous examples relate to pictures, a picture can include one or more stripes, each stripe including one or more blocks. That is, the blocks of the picture can be included in the stripes of the picture.

[0101] The corresponding text regarding the picture width and height in VVC WD can be modified to:

[0102] pic_width_in_luma_samples specifies the width of each decoded picture in luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of the maximum of 8 and MinCbSizeY.

[0103] pic_height_in_luma_samples specifies the height of each decoded picture in luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of the maximum of 8 and MinCbSizeY.

[0104] Where MinCbSizeY specifies the minimum width and / or the minimum height of the decoding unit for the picture.

[0105] That is, a video encoder (e.g., video encoder 200) may be configured to set the width of a picture to include a first number of luma samples (e.g., pic_width_in_luma_samples), where the first number is a first multiple of the maximum of 8 and the minimum coding unit size for the picture. In some examples, the video encoder may be configured to set the height of the picture (e.g., pic_height_in_luma_samples) to include a second number of luma samples, where the second number is a second multiple of the maximum of 8 and the minimum coding unit size for the picture. It should be understood that the values of pic_width_in_luma_samples and pic_height_in_luma_samples may be the same or different. The video encoder may signal a syntax element that indicates the values for pic_width_in_luma_samples and / or pic_height_in_luma_samples. The techniques described herein may reduce the number of possible values (e.g., prevent values less than 8), thereby potentially reducing the size of the bitstream with little to no loss in prediction accuracy and / or complexity.

[0106] A video decoder (e.g., video decoder 300) may be configured to determine the picture size. For example, the picture size may apply a picture size limit to set the width of the picture to include a first number of luma samples, where the first number is a first multiple of the maximum of 8 and the minimum coding unit size for the picture. For example, the video decoder may decode the value of a syntax element that indicates pic_width_in_luma_samples to determine the width of the picture. In some examples, the picture size may apply a picture size limit to set the height to include a second number of luma samples, where the second number is a second multiple of the maximum of 8 and the minimum coding unit size for the picture. For example, the video decoder may decode the value of a syntax element that indicates pic_height_in_luma_samples to determine the height of the picture.

[0107] After partitioning video data, a video decoder (e.g., video encoder 200 or video decoder 300) may generate prediction information for a block and determine a predicted block for the block based on the prediction information. Similarly, the predicted block may depend on neighboring blocks. For example, the video decoder may determine a predicted block for a current block based on an upper neighboring block and a left neighboring block. By preventing block splitting (e.g., for chrominance components and / or luminance components) that results in relatively small block sizes, the video decoder may determine prediction information for video data blocks with fewer block dependencies, thereby potentially increasing the number of blocks that can be decoded (e.g., encoded or decoded) in parallel with little loss in prediction accuracy and / or complexity.

[0108] The video decoder (e.g., video encoder 200 or video decoder 300) may not decode (e.g., encode or decode) 2x2 and 2x4 / 4x2 blocks at the diagonal. To reconstruct these blocks, the video decoder may apply a filling method. In some examples, the video decoder may use the reconstructed pixels of the upper left pixel of the current block to reconstruct the block. In some examples of filling, the video decoder may reconstruct the block by, for example, repeating an adjacent reconstructed left column. In some examples of filling, the video decoder may reconstruct the block by, for example, repeating an adjacent reconstructed upper row.

[0109] The video decoder (e.g., video encoder 200 or video decoder 300) may expand 2x2 and 2x4 / 4x2 blocks to a size of 4x4 samples by filling 2x2 and 2x4 / 4x2 blocks located at the lower right corner of the picture. The filling region may contain zeros or another appropriate constant value (e.g., half of the maximum sample value, e.g., 512 for 10-bit samples), or repeat or reproduce block samples, etc. For example, the video decoder may fill the 2x2 and 2x4 / 4x2 blocks with a filling region that may contain zeros or another appropriate constant value (e.g., half of the maximum sample value, e.g., 512 for 10-bit samples), or repeat or reproduce block samples, etc.

[0110] The video decoder (e.g., video encoder 200 or video decoder 300) may decode (e.g., encode or decode) the resulting 4x4 blocks like other 4x4 blocks. The video decoder may be configured to crop the 4x4 block to its original 2x2 or 2x4 / 4x2 size at the picture corner after reconstructing the 4x4 block.

[0111] A video decoder (e.g., video encoder 200 or video decoder 300) may expand the block sizes of 2x2 and 2x4 / 4x2 blocks at the corners to 4x4. In this case, the video decoder may set the residuals of the expanded region to be equal to a default value (e.g., 0). Compared to encoding a 4x4 block, the transform and quantization of the expanded 4x4 block may remain unchanged. During the reconstruction process regarding the block sizes of 2x2 and 2x4 / 4x2 blocks at the corners, the video decoder may be configured to not use the prediction values at the expanded region, but use the unexpanded prediction values.

[0112] Figure 5 is a conceptual diagram showing examples of expanding a 2x2 block 170 to a 4x4 block, expanding a 2x4 block 172 to a 4x4 block, and expanding a 4x2 block 174 to a 4x4 block. Figure 5 The blocks with gray shapes may represent actual data, while Figure 5 the white areas may represent the expanded regions.

[0113] Figure 6 is a block diagram showing an example video encoder 200 that may perform the techniques of the present disclosure. Figure 6 is provided for explanatory purposes and should not be considered a limitation on the techniques widely illustrated and described in the present disclosure. For explanatory purposes, the present disclosure describes the video encoder 200 in the context of video coding standards (e.g., the HEVC video coding standard and the H.266 video coding standard being developed). However, the techniques of the present disclosure are not limited to these video coding standards and generally apply to video encoding and decoding.

[0114] In Figure 6 the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filtering unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or processing circuits. Additionally, the video encoder 200 may include additional or alternative processors or processing circuits for performing these and other functions.

[0115] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104( Figure 1 ). The DPB 218 may act as a reference picture memory that stores reference video data for use by the video encoder 200 to predict subsequent video data. The video data memory 230 and the DPB 218 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or by separate memory devices. In various examples, as shown, the video data memory 230 may be on-chip with other components of the video encoder 200 or off-chip relative to those components.

[0116] In the present disclosure, a reference to the video data memory 230 should not be construed as being limited to a memory internal to the video encoder 200, unless specifically described as such, or a memory external to the video encoder 200, unless specifically described as such. Instead, a reference to the video data memory 230 should be understood as a reference to a memory that stores video data received by the video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage of outputs from the various units of the video encoder 200.

[0117] Figure 6 The various units are shown to assist in understanding the operations performed by the video encoder 200. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides a specific function and is preset on the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by instructions of the software or firmware. Fixed-function circuitry may execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuitry is typically invariant. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0118] Video encoder 200 may include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operations of video encoder 200 are performed using software executed by programmable circuitry, memory 106( Figure 1 ) may store the object code of the software that video encoder 200 receives and executes, or another memory within video encoder 200 (not shown) may store such instructions.

[0119] Video data memory 230 is configured to store received video data. Video encoder 200 may obtain pictures of the video data from video data memory 230 and provide the video data to residual generation unit 204 and mode selection unit 202. The video data in video data memory 230 may be the original video data to be encoded.

[0120] Mode selection unit 202 includes motion estimation unit 222, motion compensation unit 224, and intra prediction unit 226. Mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As an example, mode selection unit 202 may include a palette unit, a block copy unit (which may be part of motion estimation unit 222 and / or motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0121] Mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include dividing a CTU into CUs, the prediction mode for a CU, the transform type for the residual data of a CU, the quantization parameter for the residual data of a CU, etc. Mode selection unit 202 may ultimately select the combination of encoding parameters having a better rate-distortion value compared to other tested combinations.

[0122] Video encoder 200 may divide a picture obtained from video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a strip. Mode selection unit 202 may divide the CTUs of the picture according to a tree structure (such as the QTBT structure or quadtree structure of HEVC as described above). As described above, video encoder 200 may form one or more CUs by dividing CTUs according to a tree structure. Such a CU may generally also be referred to as a "video block" or "block". In some examples, mode selection unit 202 may be configured to determine a plurality of sub-blocks of a non-split chrominance block of video data based on a chrominance split limit and process the plurality of sub-blocks to generate prediction information for the non-split chrominance block. In some examples, mode selection unit 202 may be configured to limit the picture size to avoid 2x2 and 2x4 / 4x2 chrominance blocks at a corner of the picture. In some examples, mode selection unit 202 may be configured to prohibit processing of 2x2 and 2x4 / 4x2 blocks at a corner of the picture.

[0123] Generally, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or an overlapping portion of a PU and a TU in HEVC). For inter prediction of the current block, motion estimation unit 222 may perform a motion search to identify one or more reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218) that closely match. Specifically, motion estimation unit 222 may calculate, for example, a value representing the similarity between a potential reference block and the current block according to the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may generally perform these calculations using the per-sample differences between the current block and the reference block under consideration. Motion estimation unit 222 may identify the reference block having the lowest value obtained from these calculations, and this lowest value indicates the reference block that most closely matches the current block.

[0124] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in the current picture. Then, the motion estimation unit 222 may provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. Then, the motion compensation unit 224 may use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to obtain the data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate the values for the prediction block according to one or more interpolation filters. In addition, for bi-directional inter prediction, the motion compensation unit 224 may obtain the data for two reference blocks identified by the respective motion vectors and combine the obtained data, for example, by per-sample averaging or weighted averaging.

[0125] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block according to the samples adjacent to the current block. For example, for the directional mode, the intra prediction unit 226 may typically mathematically combine the values of the adjacent samples and fill the current block with these calculated values in the defined direction to produce a prediction block. As another example, for the DC mode, the intra prediction unit 226 may calculate the average value of the adjacent samples of the current block and generate a prediction block that includes this obtained average value for each sample of the prediction block.

[0126] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the unencoded original version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the per-sample difference between the current block and the prediction block. The obtained per-sample difference defines the residual block for the current block. In some examples, the residual generation unit 204 may also determine the differences between the sampled values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, one or more subtractor circuits that perform binary subtraction may be used to form the residual generation unit 204.

[0127] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of a CU may refer to the size of the luminance coding block of the CU, and the size of a PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a specific CU is 2Nx2N, the video encoder 200 may support a PU size of 2Nx2N or NxN for intra prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 may also support an asymmetric division for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0128] In an example where the mode selection unit does not further divide a CU into PUs, each CU may be associated with a luminance coding block and a corresponding chrominance coding block. As described above, the size of a CU may refer to the size of the luminance coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0129] The mode selection unit 202 may limit the picture size. For example, the mode selection unit 202 may be configured to set each of the width and the height of the picture to a respective multiple of the maximum of 8 and the size of the smallest coding unit for the picture. In this way, the mode selection unit 202 may take into account downsampling the chrominance component of the block, which may reduce the coding complexity with little or no loss of coding accuracy. For example, the mode selection unit 202 may limit the picture size to prevent splitting the chrominance component of a block partition into relatively small chrominance blocks (e.g., 2x2 chrominance blocks, 2x4 chrominance blocks, or 4x2 chrominance blocks). Limiting the picture size may help reduce the coding complexity resulting from block dependencies while having no or little impact on coding accuracy.

[0130] For other video coding techniques (such as, as a few examples, intra-block copy mode coding, affine mode coding, and linear model (LM) mode coding), the mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the coding technique. In some examples, such as palette mode coding, the mode selection unit 202 may not generate a prediction block but instead generates a syntax element indicating a way to reconstruct the block based on a selected palette. In this mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for coding.

[0131] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the prediction block and the current block.

[0132] The transform processing unit 206 applies one or more transforms to the residual block to generate a transform coefficient block (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, such as a primary transform and a secondary transform, such as a rotation transform. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0133] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause loss of information, and thus, the quantized transform coefficients may have lower precision compared to the original transform coefficients generated by the transform processing unit 206.

[0134] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block according to the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although with a certain degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to produce the reconstructed block.

[0135] The filtering unit 216 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce blocky artifacts along the edges of the CU. In some examples, the operation of the filtering unit 216 may be skipped.

[0136] Video encoder 200 stores the reconstructed blocks in DPB 218. For example, in an example where the operation of filter unit 216 is not required, reconstruction unit 214 may store the reconstructed blocks into DPB 218. In an example where the operation of filter unit 216 is required, filter unit 216 may store the filtered reconstructed blocks into DPB 218. Motion estimation unit 222 and motion compensation unit 224 may obtain reference pictures formed by the reconstructed (and potentially filtered) blocks from DPB 218 to perform inter prediction on the blocks of the subsequently encoded pictures. Additionally, intra prediction unit 226 may use the reconstructed blocks in DPB 218 of the current picture to perform intra prediction on other blocks in the current picture.

[0137] Generally, entropy coding unit 220 may perform entropy coding on the syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 may perform entropy coding on the quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 may perform entropy coding on the prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from mode selection unit 202. Entropy coding unit 220 may perform one or more entropy coding operations on the syntax elements as another example of video data to generate the entropy-coded data. For example, entropy coding unit 220 may perform context adaptive variable length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 may operate in a bypass mode where the syntax elements are not entropy coded.

[0138] Video encoder 200 may output a bitstream that includes the entropy-coded syntax elements required to reconstruct the blocks of a slice or picture. Specifically, entropy coding unit 220 may output the bitstream.

[0139] The operations described above are described with respect to blocks. Such description should be understood as operations for luma decoding blocks and / or chroma decoding blocks. As described above, in some examples, the luma decoding blocks and chroma decoding blocks are the luma and chroma components of a CU. In some examples, the luma decoding blocks and chroma decoding blocks are the luma and chroma components of a PU.

[0140] In some examples, for chrominance decoding blocks, operations performed with respect to luminance decoding blocks need not be repeated. As an example, for identifying motion vectors (MVs) and reference pictures for chrominance blocks, operations for identifying MVs and reference pictures for luminance decoding blocks need not be repeated. Instead, MVs for luminance decoding blocks can be scaled to determine MVs for chrominance blocks, and the reference pictures can be the same. As another example, for luminance decoding blocks and chrominance decoding blocks, the intra prediction process can be the same.

[0141] Video encoder 200 represents an example of a device configured to encode video data, the device including a memory configured to store video data, and one or more processing units implemented in circuitry and configured to: determine a plurality of sub-blocks of a non-split chrominance block of the video data based on a chrominance splitting limit, and encode the plurality of sub-blocks to generate prediction information for the non-split chrominance block. In some examples, video encoder 200 can be configured to limit the picture size to avoid 2x2 and 2x4 / 4x2 chrominance blocks at a corner of the picture. In some examples, video encoder 200 and video decoder 300 can be configured to prohibit processing of 2x2 and 2x4 / 4x2 blocks at a corner of the picture.

[0142] Video encoder 200 represents an example of a device configured to encode video data, the device including a memory (e.g., video data memory 230) configured to store video data and one or more processors implemented in circuitry. Mode selection unit 202 can be configured to limit the picture size of the video data. To limit the picture size of the video data, mode selection unit 202 can be configured to set each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the size of the smallest decoding unit for the picture. Mode selection unit 202 can be configured to determine a predicted block for a block of the picture. Residual generation unit 204 can be configured to generate a residual block for the block based on the difference between the block and the predicted block. Entropy encoding unit 220 can encode the residual block.

[0143] Figure 7 is a block diagram illustrating an example video decoder 300 that can perform the techniques of the present disclosure. Figure 7 is provided for explanatory purposes and does not limit the techniques generally illustrated and described in the present disclosure. For explanatory purposes, the present disclosure describes video decoder 300 in accordance with techniques of JEM, VVC, and HEVC. However, the techniques of the present disclosure can be performed by video decoding devices configured to conform to other video coding standards.

[0144] In Figure 7In the example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or in processing circuitry. Additionally, video decoder 300 may include additional or alternative processors or processing circuitry for performing these and other functions.

[0145] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units for performing prediction according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, a block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 may include more, fewer, or different functional components. In some examples, the prediction processing unit 304 may be configured to determine multiple sub-blocks of a non-split chrominance block of video data based on chrominance split limits and process the multiple sub-blocks to generate prediction information for the non-split chrominance block. In some examples, the prediction processing unit 304 may be configured to limit the picture size to avoid 2x2 and 2x4 / 4x2 chrominance blocks at a corner of the picture. In some examples, the prediction processing unit 304 may be configured to prohibit processing of 2x2 and 2x4 / 4x2 blocks at a corner of the picture.

[0146] The CPB memory 320 may store video data to be decoded by components of video decoder 300, such as an encoded video bitstream. The video data stored in the CPB memory 320 may be, for example, from a computer-readable medium 110 ( Figure 1)Obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from an encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of decoded pictures, such as temporary data representing the outputs of the respective units from the video decoder 300. The DPB 314 generally stores decoded pictures, and the video decoder 300 may output and / or use the decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed of any of a variety of memory devices (such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices). The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip or off-chip relative to the other components of the video decoder 300.

[0147] Additionally or alternatively, in some examples, the video decoder 300 may obtain decoded video data from the memory 120( Figure 1 ). That is, the memory 120 may store data together with the CPB memory 320 as described above. Similarly, when some or all of the functions of the video decoder 300 are implemented in software to be executed by the processing circuitry of the video decoder 300, the memory 120 may store instructions to be executed by the video decoder 300.

[0148] Figure 7 The various units shown are presented to assist in understanding the operations performed by the video decoder 300. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 6 , fixed-function circuitry refers to circuitry that provides a specific function and is pre-set on the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by instructions of the software or firmware. Fixed-function circuitry may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuitry are generally invariant. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0149] The video decoder 300 may include an ALU, an EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuitry, on-chip or off-chip memory may store the instructions (e.g., object code) of the software that the video decoder 300 receives and executes.

[0150] The entropy decoding unit 302 may receive the encoded video data from the CPB and perform entropy decoding on the video data to regenerate syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate the decoded video data based on the syntax elements extracted from the bitstream.

[0151] Generally, the video decoder 300 reconstructs pictures block by block. The video decoder 300 may perform the reconstruction operation on each block individually (where the block being currently reconstructed, i.e., the decoded block, may be referred to as the "current block").

[0152] The entropy decoding unit 302 may perform entropy decoding on the syntax elements, where the syntax elements define the quantized transform coefficients of the quantized transform coefficient block and the transform information (such as quantization parameter (QP) and / or transform mode indication). The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by the inverse quantization unit 306. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse-quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0153] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.

[0154] The prediction processing unit 304 may determine a picture size limit. For example, the picture size limit may include setting each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the minimum decoding unit size for the picture. In this way, the prediction processing unit 304 may take into account downsampling of the chrominance components of the blocks, which may reduce the decoding complexity with little or no loss of decoding accuracy. For example, the prediction processing unit 304 may apply a picture size limit that prevents splitting the chrominance components of a block partitioned into relatively small chrominance blocks (e.g., 2x2 chrominance blocks, 2x4 chrominance blocks, or 4x2 chrominance blocks). Applying the picture size limit may help reduce the decoding complexity resulting from block dependencies while having no or little impact on the decoding accuracy.

[0155] In addition, the prediction processing unit 304 generates a prediction block based on the prediction information syntax element entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-predicted, the motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate the reference picture in the DPB 314 from which to obtain the reference block and the motion vector identifying the position of the reference block in the reference picture relative to the current block in the current picture. The motion compensation unit 316 may generally perform the inter-frame prediction process in a manner substantially similar to the manner described with respect to the motion compensation unit 224 ( Figure 6 ).

[0156] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Similarly, the intra-prediction unit 318 may perform the intra-prediction process in a manner substantially similar to the manner described with respect to the intra-prediction unit 226 ( Figure 6 ). The intra-prediction unit 318 may obtain data of adjacent samples of the current block from the DPB 314.

[0157] The reconstruction unit 310 may use the prediction block and the residual block to reconstruct the current block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0158] The filtering unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operation of the filtering unit 312 is not necessarily performed in all examples.

[0159] Video decoder 300 may store the reconstructed blocks in DPB 314. For example, in an example where the operations of filter unit 312 are not performed, reconstruction unit 310 may store the reconstructed blocks into DPB 314. In an example where the operations of filter unit 312 are performed, filter unit 312 may store the filtered reconstructed blocks into DPB 314. As described above, DPB 314 may provide reference information to prediction processing unit 304, such as samples of the current picture for intra prediction and samples of the previously decoded pictures for subsequent motion compensation. In addition, video decoder 300 may output the decoded pictures from DPB 314 for subsequent presentation on a display device such as Figure 1 display device 118.

[0160] In this way, video decoder 300 represents an example of an apparatus that includes a memory (e.g., video data memory 230) configured to store video data and one or more processing units implemented in circuitry and configured to determine a picture size limit. The picture size limit may include setting each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the minimum coding unit size for the picture. Mode selection unit 202 may be configured to determine a prediction block for a block of a picture. Residual generation unit 204 may be configured to generate a residual block for the block based on the difference between the block and the prediction block. Entropy encoding unit 220 may encode the residual block. Prediction processing unit 304 may be configured to determine a prediction block for a block of a picture. Entropy decoding unit 302 may be configured to decode the residual block for the block. Reconstruction unit 310 may be configured to merge the prediction block and the residual block to decode the block.

[0161] Figure 8 is a flowchart illustrating an example method for encoding a current block. The current block may include a current CU. Although described with respect to video encoder 200 ( Figure 1 and 6 ), other devices may be configured to perform methods similar to Figure 8 the method.

[0162] Mode selection unit 202 may set the picture size (348) of the picture. For example, mode selection unit 202 may apply the picture size limit to set each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the minimum coding unit size for the picture. By preventing splits that result in relatively small block sizes, mode selection unit 202 may determine prediction blocks for blocks of a picture (e.g., a video picture, a slice, or another portion) of video data with fewer block dependencies, thereby potentially increasing the number of blocks that can be decoded (e.g., encoded or decoded) in parallel with little loss in prediction accuracy and / or complexity.

[0163] The mode selection unit 202 may predict the current block (350). The mode selection unit 202 may limit the partitioning of the picture. For example, the mode selection unit 202 may set each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the minimum decoding unit size for the picture, which may help prevent splitting of the block that would result in a smaller chrominance block at a corner of the picture.

[0164] For example, the mode selection unit 202 may form a prediction block for the current block. Then, the residual generation unit 204 may calculate a residual block (352) for the current block. To calculate the residual block, the residual generation unit 204 may calculate the difference between the original unencoded block and the prediction block for the current block. The transform processing unit 206 and the quantization unit 208 may then transform and quantize the transform coefficients (354) of the residual block. The entropy encoding unit 220 may scan the quantized transform coefficients (356) of the residual block. During or after the scanning, the entropy encoding unit 220 may perform entropy encoding on the transform coefficients (358). For example, the entropy encoding unit 220 may use CAVLC or CABAC to encode the transform coefficients. The entropy encoding unit 220 may output the entropy-encoded data (360) of the block.

[0165] Figure 9 is a flowchart showing an example method for decoding a current block of video data. The current block may include a current CU. Although described with respect to the video decoder 300 ( Figure 1 and 7 ), other devices may be configured to perform a method similar to Figure 9 's method.

[0166] The prediction processing unit 304 may determine the picture size (368) of the picture. For example, the prediction processing unit 304 may determine the picture size that applies a picture size limit to set each of the width of the picture and the height of the picture to a respective multiple of the maximum of 8 and the minimum decoding unit size for the picture. Configuring the prediction processing unit 304 to determine the picture size for applying the picture size limit may help prevent splitting that results in a relatively small block size. Preventing splitting that results in a relatively small block size may help determine the prediction block for a block (e.g., a video picture, a stripe, or another portion) of the picture of the video data with fewer block dependencies, and thus potentially increase the number of blocks that can be decoded (e.g., encoded or decoded) in parallel with little loss in prediction accuracy and / or complexity.

[0167] The entropy decoding unit 302 may receive the entropy-coded data for the current block, such as the entropy-coded prediction information and the entropy-coded data of the coefficients for the residual block corresponding to the current block (370). The entropy decoding unit 302 may perform entropy decoding on the entropy-coded data to determine the prediction information for the current block and regenerate the coefficients of the residual block (372). The prediction processing unit 304 may predict the current block (374), for example, using an intra or inter prediction mode indicated by the prediction information for the current block, to calculate a prediction block for the current block. The prediction processing unit 304 may use a picture size constraint to determine the partitioning of the video data. The picture size constraint may include setting each of the width of the picture and the height of the picture to a multiple of the maximum of 8 and the minimum coding unit size of the picture. For example, the prediction processing unit 304 may apply the picture size constraint to determine a partition for preventing splitting of a block, where the splitting would result in smaller block chroma blocks at the corners of the picture.

[0168] Then, the entropy decoding unit 302 may inverse scan the regenerated coefficients (376) to create a block of quantized transform coefficients. The inverse quantization unit 306 and the inverse transform processing unit 308 may perform inverse quantization and inverse transform on the transform coefficients to generate a residual block (378). The reconstruction unit 310 may decode the current block (380) by combining the prediction block and the residual block.

[0169] Figure 10 is a flowchart showing an example process of using a picture size constraint according to the techniques of the present disclosure. The current block may include the current CU. Although described with respect to the video decoder 300 ( Figure 1 and Figure 7 ), other devices may be configured to perform methods similar to the method of Figure 9 .

[0170] A video coder (e.g., the video encoder 200 or the video decoder 300, or more specifically, e.g., the mode selection unit 202 of the video encoder 200 or the prediction processing unit 304 of the video decoder 300) may determine a minimum CU size (390). For example, the video coder may determine MinCbSizeY. The video coder may determine a multiplication value as max(8, minCuSize) (391). For example, if 8 is greater than minCuSize, the video coder may determine the multiplication value as 8, and if minCuSize is greater than 8, the video coder may determine the multiplication value as minCuSize.

[0171] A video decoder may use a multiplication value to determine the picture size (392). For example, the video decoder may determine the width for the picture based on an integer (e.g., X1) and a multiplication value (e.g., max(8, minCuSize)), and determine the height for the picture based on an integer (e.g., X2) and a multiplication value (e.g., max(8, minCuSize)).

[0172] The video decoder may use the picture size to identify the bottom-right block of the picture (393). For example, the video decoder may determine a partitioning (e.g., partitioning or decoding one or more syntax values indicating the partitioning) that splits the picture into an integer number of blocks based on the picture size, including the bottom-right block of the picture.

[0173] The video decoder may determine the size of the bottom-right block based on the picture size (394). For example, the video decoder may determine a partitioning (e.g., partitioning or decoding one or more syntax values indicating the partitioning) that sets the bottom-right block to have a width of at least 8 samples and a height of at least 8 samples based on the picture size.

[0174] The video decoder may decode (e.g., encode or decode) the bottom-right block (395). For example, the video decoder may generate a prediction block for the bottom-right block. In an example where the video decoder is a video encoder, the video encoder may generate a residual block for the bottom-right block. In this example, the video encoder may generate a residual block for the bottom-right block based on the difference between the bottom-right block and the predicted bottom-right block, and encode the residual block. In an example where the video decoder is a video decoder, the video decoder may generate a prediction block for the bottom-right block. In this example, the video decoder may decode the residual block for the bottom-right block and merge the prediction block and the residual block to decode the bottom-right block.

[0175] In Figure 10 the example, a video decoder (e.g., video encoder 200 or video decoder 300) may perform encoding and / or decoding of the bottom-right corner of a picture, which may include at least 64 luma samples. In some examples, a video decoder (e.g., video encoder 200 or video decoder 300) may use a local dual-tree scheme to encode and / or decode the bottom-right block, where the luma blocks for the bottom-right corner of the picture may be split into luma sub-blocks (e.g., an 8x8 block may be split into two 4x8 blocks). The video decoder may encode and / or decode all luma sub-blocks using the same inter prediction mode. In some examples, the video decoder may use an intra prediction mode, intra block copy (IBC), and / or palette mode to encode and / or decode the luma sub-blocks.

[0176] In response to determining that all luma sub-blocks are encoded or decoded using an inter-frame mode, a video coder (e.g., video encoder 200 or video decoder 300) may split chroma blocks at the lower right corner of a picture into chroma sub-blocks when splitting a luma block, and inherit motion information from the luma block to encode and / or decode the chroma sub-blocks. However, in response to determining that the luma sub-blocks are encoded or decoded using an intra-frame mode, intra-block copy (IBC), and / or palette mode, the video coder may not split the chroma blocks, and the video coder may use intra prediction to encode and / or decode the chroma blocks.

[0177] The following provides a non-limiting illustrative list of examples of the techniques of the present disclosure.

[0178] Example 1: A method of processing video data, the method comprising: determining, by a video coder, based on a chroma splitting limit, a plurality of sub-blocks of a non-split chroma block of an inter-frame decoded strip of video data; and processing, by the video coder, the plurality of sub-blocks to generate prediction information for the non-split chroma block.

[0179] Example 2: The method according to Example 1, comprising: forcing intra-frame decoding for each of the plurality of sub-blocks when all blocks in a corresponding luma region are intra-frame decoded.

[0180] Example 3: The method according to any one of Examples 1-2, comprising: encoding each sub-block using a default mode when a corresponding luma region contains both inter-frame decoded luma blocks and intra-frame decoded luma blocks.

[0181] Example 4: The method according to Example 3, wherein the default mode includes an intra-frame mode.

[0182] Example 5: The method according to Example 3, wherein the default mode includes an inter-frame mode.

[0183] Example 6: The method according to Example 5, further comprising: determining a motion vector for one or more of the plurality of sub-blocks according to a motion vector of one of the inter-frame decoded luma blocks in the corresponding luma region.

[0184] Example 7: The method according to Example 5, further comprising: determining a motion vector for one or more of the plurality of sub-blocks according to an average of motion vectors of each of the inter-frame decoded luma blocks in the corresponding luma region.

[0185] Example 8. A method for processing video data, the method comprising: restricting the picture size to avoid 2x2 and 2x4 / 4x2 chrominance blocks at the corners of the picture; determining, by a video decoder, a plurality of blocks of the picture of the video data; and processing, by the video decoder, the plurality of blocks.

[0186] Example 9. The method according to Example 8, wherein restricting the picture size includes setting the width of the picture and the height of the picture to multiples of the maximum of 8 and the minimum CU size.

[0187] Example 10. A method for processing video data, the method comprising: determining, by a video decoder, a plurality of blocks of a picture of the video data; and processing, by the video decoder, the plurality of blocks, wherein processing the plurality of blocks includes prohibiting processing of 2x2 and 2x4 / 4x2 blocks at the corners of the picture.

[0188] Example 11. The method for processing video data according to Example 10, wherein prohibiting processing of 2x2 and 2x4 / 4x2 blocks at the corners of the picture includes: reconstructing the 2x2 and 2x4 / 4x2 blocks at the corners of the picture.

[0189] Example 12. The method for processing video data according to Example 10, wherein prohibiting processing of 2x2 and 2x4 / 4x2 blocks at the corners of the picture includes: expanding the 2x2 and 2x4 / 4x2 blocks to a size of 4x4 samples by filling the 2x2 and 2x4 / 4x2 blocks located at the lower right corner of the picture.

[0190] Example 13. The method for processing video data according to Example 10, wherein prohibiting processing of 2x2 and 2x4 / 4x2 blocks at the corners of the picture includes: expanding the block size of the 2x2 and 2x4 / 4x2 blocks at the corners to 4x4, and setting the residual of the expanded region to a default value.

[0191] Example 14. A device for decoding video data, the device comprising one or more units for performing the method according to any one of Examples 1-13.

[0192] Example 15. The device according to Example 14, wherein the one or more units include one or more processors implemented in a circuit.

[0193] Example 16. The device according to any one of Examples 14 and 15, further comprising a memory for storing video data.

[0194] Example 17. The device according to any one of Examples 14-16, further comprising: a display configured to display the decoded video data.

[0195] Example 18. The apparatus according to any one of Examples 14-17, wherein the apparatus comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0196] Example 19. The apparatus according to any one of Examples 14-18, wherein the apparatus comprises a video decoder.

[0197] Example 20. The apparatus according to any one of Examples 14-19, wherein the apparatus comprises a video encoder.

[0198] Example 21. A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the method according to any one of Examples 1-13.

[0199] It should be appreciated that, according to the examples, the specific acts or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted together (e.g., not all described acts or events are necessary for the practice of the technique). Additionally, in a particular example, the acts or events may be performed concurrently, e.g., via multi-threading, interrupt processing, or multiple processors, rather than sequentially.

[0200] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, the computer-readable medium generally corresponds to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0201] By way of example and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transitory tangible storage media. As used herein, disk and optical disks include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0202] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms "processor" and "processing circuitry" can refer to any one of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, these techniques can be fully implemented in one or more circuits or logic elements.

[0203] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a group of ICs (e.g., a chipset). In the present disclosure, various components, modules, or units are described to emphasize the functional aspects of the devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0204] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for decoding video data, the method comprising: determining, by one or more processors implemented in a circuit, a picture size of a picture according to a picture size limit, wherein the picture size limit includes: determining a maximum value between 8 and a minimum coding unit size for the picture, and determining each of a width and a height of the picture as a respective multiple of the determined maximum value; determining, by the one or more processors, a partitioning for partitioning the picture into a plurality of blocks; generating, by the one or more processors, a predicted block for a block among the plurality of blocks; decoding, by the one or more processors, a residual block for the block; and combining, by the one or more processors, the predicted block and the residual block to decode the block.

2. The method according to claim 1, wherein The minimum coding unit size includes: a minimum width of a coding unit for the picture or a minimum height of the coding unit for the picture.

3. The method according to claim 1, wherein Determining the partitioning includes: determining that a chrominance block at a lower right corner of the picture has a size of at least 4x4.

4. The method according to claim 1, wherein Determining the partitioning includes: determining that a luminance block at a lower right corner of the picture has a size of at least 8x8.

5. The method according to claim 1, Among them, wherein the picture size applies the picture size limit to set the width of the picture to include a first number of luminance samples, the first number being a first multiple of a maximum value between 8 and the minimum coding unit size for the picture; and wherein the picture size applies the picture size limit to set the height of the picture to include a second number of luminance samples, the second number being a second multiple of a maximum value between 8 and the minimum coding unit size for the picture.

6. The method according to claim 1, wherein, The block is included in a strip of the picture.

7. A method for encoding video data, the method comprising: setting, by one or more processors implemented in a circuit, a picture size of a picture, wherein setting the picture size includes: determining a maximum value between 8 and a minimum coding unit size for the picture, and applying a picture size limit to set each of a width and a height of the picture as a respective multiple of the determined maximum value; partitioning, by the one or more processors, the picture into a plurality of blocks; generating, by the one or more processors, a predicted block for a block among the plurality of blocks; generating, by the one or more processors, a residual block for the block based on a difference between the block and the predicted block; and encoding, by the one or more processors, the residual block.

8. The method according to claim 7, wherein The minimum coding unit size includes: a minimum width of a coding unit for the picture or a minimum height of the coding unit for the picture.

9. The method according to claim 7, wherein The partitioning includes: determining that a chrominance block at a lower right corner of the picture has a size of at least 4x4.

10. The method according to claim 7, wherein, The partitioning includes: determining that a luminance block at a lower right corner of the picture has a size of at least 8x8.

11. The method according to claim 7, wherein, Setting the picture size includes: Set the width of the picture to include a first number of luminance samples, the first number being a first multiple of the maximum of 8 and the minimum coding unit size for the picture; and Set the height of the picture to include a second number of luminance samples, the second number being a second multiple of the maximum of 8 and the minimum coding unit size for the picture.

12. The method according to claim 7, wherein The block is included in a strip of the picture.

13. A device for decoding video data, the device comprising one or more processors implemented in circuitry and configured to perform the following operations: Determine the picture size of a picture according to picture size limitations to determine the maximum of 8 and the minimum coding unit size for the picture, and determine the width of the picture and the height of the picture each as a respective multiple of the determined maximum; Determine a partitioning for dividing the picture into a plurality of blocks; Generate a predicted block for a block among the plurality of blocks; Decode a residual block for the block; And Merge the predicted block and the residual block to decode the block.

14. The device according to claim 13, wherein, The minimum coding unit size includes: the minimum width of a coding unit for the picture or the minimum height of the coding unit for the picture.

15. The device according to claim 13, wherein, To determine the partitioning, the one or more processors are configured to: determine that a chrominance block at the lower right corner of the picture includes a size of at least 4x4.

16. The apparatus according to claim 13, wherein, To determine the partitioning, the one or more processors are configured to: determine that a luminance block at the lower right corner of the picture includes a size of at least 8x8.

17. The device according to claim 13, Among them, The picture size applies the picture size limitations to set the width of the picture to include a first number of luminance samples, the first number being a first multiple of the maximum of 8 and the minimum coding unit size for the picture; And wherein the picture size applies the picture size limitations to set the height of the picture to include a second number of luminance samples, the second number being a second multiple of the maximum of 8 and the minimum coding unit size for the picture.

18. The apparatus according to claim 13, wherein, The block is included in a strip of the picture.

19. The apparatus according to claim 13, further comprising: A display configured to display the picture.

20. The apparatus according to claim 13, wherein The device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

21. A device for encoding video data, the device comprising one or more processors implemented in circuitry and configured to perform the following operations: Set the image size of the picture, where To set the picture size, the one or more processors are configured to: determine the maximum of 8 and the minimum coding unit size for the picture, and apply picture size limitations to set the width of the picture and the height of the picture each as a respective multiple of the determined maximum; Divide the picture into a plurality of blocks; Generate a predicted block for a block among the plurality of blocks; Generate a residual block for the block based on the difference between the block and the predicted block; and Encode the residual block.

22. The device according to claim 21, wherein, The minimum decoding unit size includes: the minimum width of the decoding unit for the picture or the minimum height of the decoding unit for the picture.

23. The device according to claim 21, wherein, For partitioning, the one or more processors are configured to: determine that the chrominance block at the lower right corner of the picture has a size of at least 4x4.

24. The apparatus according to claim 21, wherein, For partitioning, the one or more processors are configured to: determine that the luma block at the lower right corner of the picture has a size of at least 8x8.

25. The device according to claim 21, wherein, For setting the picture size, the one or more processors are configured to: Set the width of the picture to include a first number of luma samples, where the first number is a first multiple of the maximum value of 8 and the minimum decoding unit size for the picture; and Set the height of the picture to include a second number of luma samples, where the second number is a second multiple of the maximum value of 8 and the minimum decoding unit size for the picture.

26. The device according to claim 21, wherein, The block is included in a strip of the picture.

27. An apparatus for decoding video data, the apparatus comprising: A unit for determining the picture size of a picture, where the picture size applies a picture size limit, where the picture size limit includes: determining the maximum value of 8 and the minimum decoding unit size for the picture, and determining the width of the picture and the height of the picture as respective multiples of the determined maximum value; A unit for determining a partitioning for partitioning the picture into a plurality of blocks; A unit for generating a predicted block for a block among the plurality of blocks; A unit for decoding a residual block for the block; and A unit for combining the predicted block and the residual block to decode the block.

28. An apparatus for encoding video data, the apparatus comprising: A unit for setting the picture size of a picture, where the unit for setting the picture size includes: a unit for determining the maximum value of 8 and the minimum decoding unit size for the picture, and a unit for applying a picture size limit to set the width of the picture and the height of the picture as respective multiples of the determined maximum value; A unit for partitioning the picture into a plurality of blocks; A unit for generating a predicted block for a block among the plurality of blocks; A unit for generating a residual block for the block based on the difference between the block and the predicted block; and A unit for encoding the residual block.

29. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the following operations: Determine the image size of the picture, where, The picture size applies a picture size limit to determine the maximum value of 8 and the minimum decoding unit size for the picture, and to determine the width of the picture and the height of the picture as respective multiples of the determined maximum value; Determine a partitioning for partitioning the picture into a plurality of blocks; Generate a predicted block for a block among the plurality of blocks; Decode a residual block for the block; And Combine the predicted block and the residual block to decode the block.

30. A non-transitory computer-readable storage medium having instructions stored thereon, which when executed cause one or more processors to perform the following operations: Set the image size of the picture, where, To set the picture size, the instructions cause the one or more processors to: determine the maximum of 8 and the minimum decoding unit size for the picture, and apply a picture size limit to set the width and the height of the picture to respective multiples of the determined maximum; Divide the picture into a plurality of blocks; Generate prediction blocks for the blocks among the plurality of blocks; Generate residual blocks for the blocks based on the difference between the blocks and the prediction blocks; and Encode the residual blocks.

Citation Information

Patent Citations

  • Reconstruction of blocks of video data using block size restriction

    US20200296367A1