Using a non-rectangular prediction mode to reduce the storage of motion fields for video data prediction

By adopting non-rectangular segmentation mode in video encoding and decoding and disabling mixing operations, the problem of insufficient data storage and bit rate optimization in the prior art is solved, and more efficient video data storage and transmission is achieved.

CN113892264BActive Publication Date: 2025-06-10QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080039988.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-04
Filing Date
2020-06-05
Publication Date
2025-06-10
Estimated Expiration
2040-06-05

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies use non-rectangular segmentation mode, data storage and bit rate optimization are insufficient, resulting in inefficient video data storage and transmission efficiency.

Method used

By segmenting and predicting video data blocks in a non-rectangular segmentation mode (such as triangle segmentation mode) during video encoding and decoding, and disabling mixing operations in storage and transmission, only the set of motion information related to the segmentation below is stored and used.

Benefits of technology

The data storage amount during and after the video encoding and decoding process is reduced, the bit rate of the codeced video data is reduced, and the storage and transmission efficiency of the video data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113892264B_ABST
    Figure CN113892264B_ABST
Patent Text Reader

Abstract

An exemplary video codec device is configured to: encode and decode a first set of motion information of a current block of video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning pattern, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding and decoding the first set of motion information, encode and decode a second set of motion information of the current block that refers to the reference picture list and is associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, store the second set of motion information of the current block; and use the stored second set of motion information to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of U.S. Application No. 16 / 893,052, filed on Jun. 4, 2020, U.S. Provisional Application No. 62 / 857,584, filed on Jun. 5, 2019, and U.S. Provisional Application No. 62 / 861,811, filed on Jun. 14, 2019, the entire contents of each of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to video coding and decoding, including video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called “smart phones”, video teleconferencing devices, video streaming devices, etc. Digital video devices implement video coding and decoding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC) and extensions of such standards. By implementing such video coding and decoding techniques, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004] Video coding and decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in a video sequence. For block-based video coding and decoding, a video strip (e.g., a video picture or a portion of a video picture) can be divided into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) strip of a picture are encoded using spatial prediction with respect to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) strip of a picture can use spatial prediction with respect to reference samples in adjacent blocks in the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] Generally, this disclosure describes techniques for video coding and decoding using inter-frame prediction. More specifically, during inter-frame prediction, non-rectangular partitioning modes (e.g., triangular partitioning mode (TPM) or other geometric partitioning modes) can be used to partition and predict blocks of video data. This disclosure describes techniques that can be used to simplify the data storage for non-rectangular partitioning modes (e.g., TPM). Thus, the techniques of this disclosure can reduce the amount of data stored during and after the video coding and decoding process, thereby improving the storage of video data and the bit rate of the decoded video data. In addition, the techniques of this disclosure can be used to perform non-rectangular partitioning on blocks that form part of a unidirectional inter-frame prediction strip (i.e., a P-strip). That is, in addition to the other techniques described in this disclosure, a video codec can disable a hybrid operation during non-rectangular partitioning (e.g., TPM) of blocks of a P-strip.

[0006] In one example, a method for coding and decoding video data includes: coding a first set of motion information of a current block of video data, the current block being partitioned into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after coding the first set of motion information, coding a second set of motion information of the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, storing the second set of motion information of the current block; and using the stored second set of motion information to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.

[0007] In another example, a device for coding and decoding video data includes: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: code a first set of motion information of a current block of video data, the current block being partitioned into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after coding the first set of motion information, code a second set of motion information of the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, storing the second set of motion information of the current block in the memory; and using the stored second set of motion information to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.

[0008] In another example, instructions are stored on a computer-readable storage medium that, when executed, cause a processor to encode and decode a first set of motion information for a current block of video data, where the current block is divided into a first partition and a second partition according to a non-rectangular partitioning pattern, and the first set of motion information references a picture list and is associated with the first partition; after encoding and decoding the first set of motion information, encode and decode a second set of motion information for the current block, where the second set of motion information references the picture list and is associated with the second partition; in response to both the first set of motion information and the second set of motion information referencing the picture list, store the second set of motion information for the current block; and use the stored second set of motion information to predict subsequent motion information for a subsequent block of the video data adjacent to the current block.

[0009] In another example, a device for encoding and decoding video data includes: means for encoding and decoding a first set of motion information for a current block of the video data, where the current block is divided into a first partition and a second partition according to a non-rectangular partitioning pattern, and the first set of motion information references a picture list and is associated with the first partition; means for encoding and decoding a second set of motion information for the current block after encoding and decoding the first set of motion information, where the second set of motion information references the picture list and is associated with the second partition; means for storing the second set of motion information for the current block in response to both the first set of motion information and the second set of motion information referencing the picture list; and means for using the stored second set of motion information to predict a subsequent set of motion information for a subsequent block of the video data adjacent to the current block.

[0010] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may perform the techniques of the present disclosure.

[0012] Figure 2A and 2B is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and corresponding coding tree unit (CTU).

[0013] Figure 3 is a block diagram illustrating an example of a block divided using a triangle partitioning pattern (TPM).

[0014] Figure 4 is a block diagram illustrating spatial and temporal neighboring blocks used to construct a motion prediction candidate list.

[0015] Figure 5 is a block diagram illustrating example weights that can be used in the blending process of the TPM.

[0016] Figure 6 is a block diagram illustrating example weights that can be used in the blending process for a TPM prediction block having a stride width equal to two samples.

[0017] Figure 7 is a block diagram illustrating an example video encoder that can perform the techniques of the present disclosure.

[0018] Figure 8 is a block diagram illustrating an example video decoder that can perform the techniques of the present disclosure.

[0019] Figure 9 is a flowchart illustrating an example method for encoding a current block according to the techniques of the present disclosure.

[0020] Figure 10 is a flowchart illustrating an example method for decoding a current block according to the techniques of the present disclosure.

[0021] Figure 11 is a flowchart illustrating an example method for decoding video data according to the techniques of the present disclosure. DETAILED DESCRIPTION

[0022] Video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including their scalable video coding (SVC) and multi-view video coding (MVC) extensions. The latest joint draft of MVC was described in the ITU-T Recommendation H.264, "Advanced video coding for generic audiovisual services" in March 2010. In addition, there is a newly developed video coding standard by the Joint Collaborative Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG), namely High Efficiency Video Coding (HEVC). The latest draft of HEVC can be obtained from phenix.int-evry.fr / jct / doc_end_user / documents / 12_Geneva / wg11 / JCTVC-L1003-v34.zip.

[0023] In addition, the following describes the video coding and decoding standards to be developed: "Versatile Video Coding (Draft 5)" by Bross et al., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 14th meeting: Geneva, Switzerland, March 19 - 27, 2019, JVET-N1001-v3 (hereinafter referred to as "VVC Draft 5"). The following describes VVC and Test Model 4 (VTM 4): "Algorithm Description of Versatile Video Coding and Test Model 5 (VTM 5)" by Chen et al., document JVET-N1002, May 21, 2019.

[0024] Video coding and decoding devices (e.g., video encoders and video decoders) implement compression techniques to, for example, perform spatial and temporal prediction to reduce or remove redundancies inherent in the input video signal. To reduce temporal redundancy (i.e., the similarity between video signals in adjacent frames), motion estimation is performed to track the motion of video objects. Motion estimation can be performed on blocks of variable size. The object displacement as a result of motion estimation is typically referred to as a motion vector. The motion vector can have half-pixel, quarter-pixel, 1 / 16-pixel accuracy (or any finer accuracy). This allows the video codec to track the motion field with higher accuracy than integer pixel positions and thus obtain better predicted blocks. When using motion vectors with fractional pixel values, interpolation operations are performed.

[0025] After motion estimation, the video encoder can use a certain rate-distortion model to determine the best-performing motion vectors. Then, the video encoder can use the best motion vectors to form predicted video blocks through motion compensation. The video encoder can form a residual video block by subtracting the predicted video block from the original video block. Then, the video encoder can apply a transform to the residual block. Then, the video encoder can quantize the resulting transform coefficients and perform entropy coding on the quantized transform coefficients (and other video data, such as motion information defining the motion vectors) to further reduce the bit rate.

[0026] Figure 1 is a block diagram illustrating an example video coding and decoding system 100 that can execute the techniques of the present disclosure. The techniques of the present disclosure are generally directed to coding and / or decoding video data. Generally, video data includes any data for processing video. Thus, video data can include raw unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).

[0027] As Figure 1As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a target device 116. Specifically, source device 102 provides the video data to target device 116 via a computer-readable medium 110. Source device 102 and target device 116 can include any of a variety of devices, including desktop computers, notebooks (i.e., laptop computers), tablets, set-top boxes, handheld phones (e.g., smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and target device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.

[0028] In Figure 1 the example, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Target device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to the present disclosure, video encoder 200 of source device 102 and video decoder 300 of target device 116 can be configured to apply techniques for reducing storage of a predicted motion field using a triangle partitioning mode. Thus, source device 102 represents an example of a video encoding device, while target device 116 represents an example of a video decoding device. In other examples, the source device and the target device can include other components or arrangements. For example, source device 102 can receive video data from an external video source such as an external camera. Similarly, target device 116 can interface with an external display device rather than include an integrated display device.

[0029] As Figure 1 shown, system 100 is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for reducing storage of a predicted motion field using a triangle partitioning mode. Source device 102 and target device 116 are merely examples of such encoding and decoding devices, where source device 102 generates encoded and decoded video data for transmission to target device 116. The present disclosure refers to an “encoding and decoding” device as a device that performs encoding and decoding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of encoding and decoding devices, specifically, a video codec and a video decoder, respectively. In some examples, devices 102, 116 can operate in a substantially symmetric manner such that each of devices 102, 116 includes video encoding and decoding components. Thus, system 100 can support one-way or two-way video transmission between video devices 102, 116, e.g., for video streaming, video playback, video broadcasting, or video telephony.

[0030] Typically, video source 104 represents a video data source (i.e., raw, unencoded video data) and provides a series of consecutive pictures (also referred to as "frames") of the video data to video encoder 200, which encodes the data of the pictures. The video source 104 of source device 102 may include a video capture device, such as a camera, a video archive containing previously captured raw video, and / or a video feed interface that receives video from a video content provider. As a further alternative, the video source 104 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may rearrange the pictures from the received order (sometimes referred to as the "display order") into a codec order for encoding and decoding. The video encoder 200 may generate a bitstream containing the encoded video data. The source device 102 may then output the encoded video data via output interface 108 to a computer-readable medium 110 for reception and / or retrieval by an input interface 122 of, for example, a destination device 116.

[0031] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memories. In some examples, the memories 106, 120 may store raw video data (e.g., raw video from video source 104) and raw decoded video data from video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300. Although shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store encoded video data, such as data output from video encoder 200 and input to video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, e.g., to store raw, decoded, and / or encoded video data.

[0032] Computer-readable medium 110 can represent any type of medium or device that can transfer encoded video data from source device 102 to target device 116. In one example, computer-readable medium 110 represents a communication medium such that source device 102 can send encoded video data to target device 116 in real time, for example, via a radio-frequency network or a computer-based network. According to a communication standard (such as a wireless communication protocol), output interface 108 can modulate a transmission signal that contains the encoded video data, and input interface 122 can demodulate the received transmission signal. The communication medium can include any wireless or wired communication medium, such as the radio-frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network (such as a local area network, a wide area network, or a global network (such as the Internet)). The communication medium can include routers, switches, base stations, or any other device that helps facilitate communication from source device 102 to target device 116.

[0033] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, target device 116 can access the encoded data from storage device 112 via input interface 122. Storage device 112 can include any one of a variety of distributed or locally accessible data storage media, such as a hard disk, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0034] In some examples, source device 102 can output encoded video data to file server 114 or another intermediate storage device that can store the encoded video generated by source device 102. Target device 116 can access the stored video data from file server 114 via streaming or downloading. File server 114 can be any type of server device that can store encoded video data and send the encoded video data to target device 116. File server 114 can represent a web server (such as for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. Target device 116 can access the encoded video data from file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (such as a Wi-Fi connection), a wired connection (such as a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on file server 114. File server 114 and input interface 122 can be configured to operate according to a streaming protocol, a download transfer protocol, or a combination thereof.

[0035] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (e.g., encoded video data) according to cellular communication standards such as, for example, 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, or similar standards. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data, such as encoded video data, according to other wireless standards, such as IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee TM ), Bluetooth TM standards, or similar standards. In some examples, source device 102 and / or destination device 116 may include respective system-on-a-chip (SoC) devices. For example, source device 102 may include an SoC device that performs the functionality belonging to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device that performs the functionality belonging to video decoder 300 and / or input interface 122.

[0036] The techniques of the present disclosure may be applied to video encoding and decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (e.g., Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0037] Input interface 122 of destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by video encoder 200, such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coding / decoding units (e.g., strips, pictures, groups of pictures, sequences, etc.), and this signaling information is also used by video decoder 300. Display device 118 displays decoded pictures of the decoded video data to a user. Display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0038] Although not shown in Figure 1shown, but in some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder, and can include appropriate MUX-DEMUX units or other hardware and / or software to handle a multiplexed stream that contains both audio and video in a common data stream. If applicable, the MUX-DEMUX unit can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0039] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, both of which can be integrated as part of a combined encoder / decoder (CODEC) in their respective devices. Devices that include the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices such as cellular phones.

[0040] The video encoder 200 and the video decoder 300 can operate according to a video coding standard (e.g., ITU-T H.265, also known as High Efficiency Video Coding (HEVC)) or an extension thereof (e.g., multi-view and / or scalable video coding extensions). Alternatively, the video encoder 200 and the video decoder 300 can operate according to other proprietary or industry standards such as the Joint Exploration Test Model (JEM) or ITU-T H.266, also known as Versatile Video Coding (VVC). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0041] Generally, video encoder 200 and video decoder 300 can perform block-based encoding and decoding of pictures. The term "block" generally refers to a structure containing data to be processed (e.g., encoded, decoded, or in other forms used in the encoding process and / or decoding process). For example, a block can contain a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 can encode and decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can encode and decode luminance and chrominance components, rather than encoding and decoding red, green, and blue (RGB) data of the samples of a picture, where the chrominance components can include two types, i.e., red chrominance component and blue chrominance component. In some examples, video encoder 200 converts the received RGB-formatted data into YUV representation before encoding, and video decoder 300 converts the YUV representation into RGB format. Alternatively, preprocessing and postprocessing units (not shown) can perform these conversions.

[0042] This disclosure generally can relate to encoding and decoding (e.g., encoding and decoding) of pictures, including the process of encoding or decoding picture data. Similarly, this disclosure can relate to encoding and decoding of blocks of pictures, including the process of encoding or decoding data of the blocks, such as prediction and / or residual encoding and decoding. An encoded video bitstream generally contains a series of values of syntax elements for representing encoding and decoding decisions (e.g., encoding and decoding modes) and the segmentation of pictures into blocks. Therefore, a reference to encoding and decoding a picture or a block generally should be understood as encoding and decoding the values of the syntax elements used to form the picture or the block.

[0043] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video codec (e.g., video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video codec divides the CTU and CUs into four equal and non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes can be referred to as a "leaf node", and the CU of this leaf node can contain one or more PUs and / or one or more TUs. The video codec can also divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the segmentation of TUs. In HEVC, a PU represents inter-prediction data, while a TU represents residual data. A CU for intra-prediction contains intra-prediction information, such as an intra-mode indication.

[0044] As another example, video encoder 200 and video decoder 300 may be configured to operate according to JEM or VVC. According to JEM or VVC, a video codec (e.g., video encoder 200) divides a picture into multiple coding tree units (CTUs). Video encoder 200 may divide a CTU according to a tree structure, such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple partitioning types, such as the distinction between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning, and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0045] In the MTT partitioning structure, a block may be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning. Ternary tree partitioning is a partitioning that divides a block into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0046] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance component and the chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for the respective chrominance components).

[0047] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, or MTT partitioning according to HEVC or other partitioning structures. For purposes of explanation, the description of the techniques of the present disclosure is presented with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video codecs configured to use quadtree partitioning or other types of partitioning.

[0048] The present disclosure may interchangeably use "N×N" and "N by N" to denote the sample dimension of a block (e.g., a CU or other video block) in terms of vertical and horizontal dimensions, such as 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU can be arranged in rows and columns. Additionally, a CU does not have to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include N×M samples, where M does not necessarily equal N.

[0049] Video encoder 200 processes the video data of a CU representing prediction and / or residual information and other information. The prediction information indicates how to predict the CU to form a prediction block of the CU. The residual information generally represents the sample-by-sample difference between the CU samples before encoding and the prediction block.

[0050] To predict a CU, video encoder 200 can generally form a prediction block of the CU through inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU from the data of previously coded and decoded pictures, while intra-frame prediction generally refers to predicting the CU from the previously coded and decoded data of the same picture. To perform inter-frame prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, e.g., in terms of the difference between the CU and the reference block. Video encoder 200 can use the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to compute a difference metric to determine whether the reference block closely matches the current CU. In some examples, video encoder 200 can use uni-directional prediction or bi-directional prediction to predict the current CU.

[0051] Some examples of JEM and VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors representing non-translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0052] In an inter prediction mode, video encoder 200 and video decoder 300 may further encode and decode motion vectors used to predict blocks of video data. For example, in the merge mode, video encoder 200 and video decoder 300 may identify motion vectors of neighboring blocks of a current block for predicting the current block. As another example, in advanced motion vector prediction (AMVP), video encoder 200 and video decoder 300 may identify motion vector prediction values of neighboring blocks and then encode and decode data representing a motion vector difference relative to a motion vector predictor and other motion information (e.g., reference picture lists and reference picture indices).

[0053] Generally, the merge mode and the AMVP mode are used to predict motion information of rectangular blocks of video data. Generally, the motion information of such rectangular blocks is the same for the entire block. However, in some cases, a non-rectangular partitioning mode (e.g., triangular partitioning mode (TPM)) may be used to partition a block. In such a non-rectangular (also referred to as geometric) partitioning mode, a block may be partitioned into two different partitions, and each partition may have its own motion information. The AMVP and merge modes used to encode and decode motion information generally assume that neighboring blocks will have only one uniform motion information. However, in a non-rectangular partitioning mode, neighboring blocks may have two different sets of motion information (one for each partition). To use the motion information of such blocks as a reference for encoding and decoding the motion information of subsequent blocks, video encoder 200 and video decoder 300 determine which set of motion information to store for subsequent reference.

[0054] According to the techniques of the present disclosure, video encoder 200 and video decoder 300 may partition a block into a first partition and a second partition using a non-rectangular partitioning mode such as TPM. Generally, the first partition may be regarded as the partition having a larger number of samples along the top edge of the current block, and the second partition may be regarded as the partition having a larger number of samples along the bottom edge of the current block. Thus, the first partition may be considered to be higher than the second partition in the current block. Video encoder 200 and video decoder 300 may encode and decode the motion information of the first partition (i.e., the upper partition) before encoding and decoding the motion information of the second partition. Similarly, a bitstream containing video data may contain the encoded motion information for the first partition before containing the encoded motion information for the second partition.

[0055] According to the technology of the present disclosure, when the block is partitioned into the first partition and the second partition as described above, and when the motion information of the first partition and the motion information of the second partition refer to the same reference picture list, the video encoder 200 and the video decoder 300 may store only the motion information of the second partition for use as a reference when encoding and decoding the motion information of the subsequent block into the current block. For example, the video encoder 200 may store the motion information of the second partition in the memory 106, and the video decoder 300 may store the motion information of the second partition in the memory 120. Alternatively, the video encoder 200 and the video decoder 300 may store the motion information of the second partition in the memory of the video encoder 200 and the video decoder 300 themselves ( Figure 1 (not shown).

[0056] Therefore, when encoding and decoding a subsequent block, the video encoder 200 and the video decoder 300 may construct a motion vector predictor candidate list for the subsequent block, which may include motion information of the second partition (i.e., the lower partition) of the previous block (referred to as the current block above) instead of the motion information of the first partition. In addition, the video encoder 200 and the video decoder 300 may use the stored motion information for the second partition of the previous block to encode and decode the motion information of the subsequent block. The subsequent block may be a single block partition, or may be partitioned using a non-rectangular partition pattern such as TPM.

[0057] By storing motion information for the second partition, rather than the motion information for the first partition, the video encoder 200 and the video decoder 300 can continue to otherwise perform a conventional merge mode or an AMVP mode when encoding and decoding subsequent motion information. Thus, the techniques for constructing and using motion vector prediction candidate lists can remain the same, thereby preventing rework of other structures of the video encoder 200 and the video decoder 300. Furthermore, there is no need to allocate additional memory space to account for blocks encoded and decoded using a non-rectangular partitioning mode (e.g., TPM) to store two sets of motion information for the partitioning of the block. Thus, these techniques can allow for efficient implementation of non-rectangular partitioning modes.

[0058] To perform intra prediction, the video encoder 200 can select an intra prediction mode to generate a prediction block. Some examples of JEM and VVC provide sixty-seven intra prediction modes, including modes for various directions, as well as a planar mode and a DC mode. Typically, the video encoder 200 selects an intra prediction mode that describes neighboring samples of a current block (e.g., a block of a CU) from which to predict samples of the current block. Assuming that the video encoder 200 encodes and decodes CTUs and CUs in a raster scan order (from left to right, from top to bottom), such samples can typically be above the current block in the same picture as the current block, above to the left of the current block, or to the left of the current block.

[0059] Video encoder 200 encodes data representing the prediction mode of the current block. For example, for an inter prediction mode, video encoder 200 may encode data representing which one of the various available inter prediction modes is used, and encode motion information for the corresponding mode. For uni - directional or bi - directional inter prediction, for example, video encoder 200 may use advanced motion vector prediction (AMVP) or merge mode to encode the motion vectors. Video encoder 200 may use a similar mode to encode the motion vectors for the affine motion compensation mode.

[0060] After prediction, such as intra prediction or inter prediction of a block, video encoder 200 may calculate residual data for the block. The residual data (e.g., residual block) represents the sample - by - sample difference between the block and the predicted block for that block, which is formed by using the corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), integer transform, wavelet transform, or conceptually similar transform to the residual video data. Additionally, video encoder 200 may apply a second - order transform after the first transform, such as a mode - dependent non - separable second - order transform (MDNSST), signal - dependent transform, Karhunen - Loeve transform (KLT), etc. Video encoder 200 produces transform coefficients after applying one or more transforms.

[0061] As described above, after any transform that produces transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to possibly reduce the amount of data used to represent the coefficients, thereby providing further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the coefficients. For example, video encoder 200 may round an n - bit value down to an m - bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a right - shift of the bits of the value to be quantized.

[0062] After quantization, video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix containing the quantized transform coefficients. The scan may be designed to place higher energy (and thus lower frequency) coefficients at the front of the vector and lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, video encoder 200 may utilize a predetermined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy encode the values of syntax elements that describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.

[0063] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be transmitted. The context may relate to, for example, whether the neighboring values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0064] Video encoder 200 may also generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for video decoder 300 in, for example, a picture header, a block header, a slice header, or other syntax data (such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS)). Video decoder 300 may similarly decode this syntax data to determine how to decode the corresponding video data.

[0065] In this way, video encoder 200 may generate a bitstream that contains encoded video data, such as syntax elements that describe the segmentation of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, video decoder 300 may receive the bitstream and decode the encoded video data.

[0066] Generally, video decoder 300 performs a process opposite to that performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 may use CABAC, in a manner substantially similar but opposite to the CABAC encoding process of video encoder 200, to decode the values of the syntax elements of the bitstream. The syntax elements may define the segmentation information of a picture into CTUs, and the segmentation of each CTU according to a corresponding segmentation structure (e.g., QTBT structure) to define the CUs of the CTU. The syntax elements may also define the prediction and residual information for blocks (e.g., CUs) of the video data.

[0067] The residual information can be represented by, for example, quantized transform coefficients. The video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of a block to regenerate a residual block for the block. The video decoder 300 uses a signaling notified prediction mode (intra prediction or inter prediction) and relevant prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Then the video decoder 300 can (on a sample-by-sample basis) combine the prediction block and the residual block to regenerate the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the blocks.

[0068] The present disclosure can generally relate to “signaling” specific information, such as syntax elements. The term “signaling” can generally refer to the conveyance of values of syntax elements and / or other data used to decode encoded video data. That is, the video encoder 200 can signal the values for the syntax elements in the bitstream. Generally, signaling refers to generating values in the bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 substantially in real time (or non-real time, e.g., may occur when storing the syntax elements to the storage device 112 for later retrieval by the destination device 116).

[0069] Figure 2A and 2B is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure 130 and corresponding coding tree unit (CTU) 132. Solid lines represent quadtree partitioning, and dashed lines indicate binary tree partitioning. In each partitioning node (i.e., non-leaf node) of the binary tree, a flag is signaled to indicate which partitioning type (i.e., horizontal or vertical) is used, where in this example, 0 indicates horizontal partitioning and 1 indicates vertical partitioning. For quadtree partitioning, since a quadtree node partitions a block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the partitioning type. Accordingly, the video encoder 200 can encode syntax elements (e.g., partitioning information) for the region tree level (i.e., solid lines) of the QTBT structure 130 and syntax elements (e.g., partitioning information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130, and the video decoder 300 can decode these syntax elements. The video encoder 200 can encode video data (e.g., prediction and transform data) for a CU represented by a terminal leaf node of the QTBT structure 130, and the video decoder 300 can decode this video data.

[0070] Generally, Figure 2BThe CTU 132 can be associated with parameters that define the size of blocks corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters can include the CTU size (representing the size of the CTU 132 in samples), the minimum quadtree size (MinQTSize, representing the minimum allowable size of a quadtree leaf node), the maximum binary tree size (MaxBTSize, representing the maximum allowable size of a binary tree root node), the maximum binary tree depth (MaxBTDepth, representing the maximum allowable depth of a binary tree), and the minimum binary tree size (MinBTSize, representing the minimum allowable size of a binary tree leaf node).

[0071] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each of these child nodes can be divided according to quadtree partitioning. That is, the nodes at the first level are either leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 represents such nodes as including a parent node and child nodes with solid lines for the branches. If the nodes at the first level are not larger than the maximum allowable binary tree root node size (MaxBTSize), they can be further divided by the corresponding binary tree. The binary tree partitioning of a node can be iterated until the resulting nodes reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents such nodes as having dashed lines for the branches. The binary tree leaf nodes are referred to as coding units (CUs), which are used for prediction (e.g., intra-picture prediction or inter-picture prediction) and transformation without any further partitioning. As discussed above, the CU can also be referred to as a "video block" or a "block".

[0072] In an example of the QTBT partitioning structure, the CTU size is set to 128×128 (luma samples and two corresponding 64×64 chroma samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If a leaf quadtree node is 128×128, since its size exceeds MaxBTSize (64×64 in this example), this leaf quadtree node will not be further partitioned by the binary tree. Otherwise, the leaf quadtree node will be further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree, and the binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further partitioning is not permitted. When the width of a binary tree node equals MinBTSize (4 in this example), it means that further horizontal partitioning is not permitted. Similarly, a binary tree node with a height equal to MinBTSize means that further vertical partitioning is not permitted for this binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.

[0073] Figure 3 is a block diagram illustrating an example of a block partitioned using a triangle partitioning pattern (TPM). In Figure 3 example, block 134 is partitioned into partitions 136A and 136B along a diagonal from left to right, and block 138 is partitioned into partitions 140A and 140B along a diagonal from right to left. In these examples, partition 136A can be considered above partition 136B, and partition 140A can be considered above partition 140B. That is, partition 136A has more samples along the upper edge of block 134 than partition 136B, and partition 140A has more samples along the upper edge of block 138 than partition 140B. Thus, for block 134, video encoder 200 and video decoder 300 can encode and decode the motion information of partition 136A before the motion information of partition 136B, and for block 138, video encoder 200 and video decoder 300 can encode and decode the motion information of partition 140A before the motion information of partition 140B.

[0074] Similarly, according to the technology of the present disclosure, if the motion information of both partition 136A and partition 136B refers to the same reference picture list, the video encoder 200 and the video decoder 300 may store the motion information of partition 136B as the reference motion information of block 134 when encoding and decoding the motion information of the subsequent blocks of block 134. Similarly, according to the technology of the present disclosure, if the motion information of both partition 140A and partition 140B refers to the same reference picture list, the video encoder 200 and the video decoder 300 may store the motion information of partition 140B as the reference motion information of block 138 when encoding and decoding the motion information of the subsequent blocks of block 138.

[0075] As introduced in JVET-N1002, the triangular partitioning mode can be applied to CUs encoded or decoded in skip or merge mode, but not to the merge mode with motion vector difference (MMVD) or combined inter and intra prediction (CIIP) mode. For CUs that meet those conditions, the video encoder 200 can encode and the video decoder 300 can decode a flag indicating whether the triangular partitioning mode is applied.

[0076] When using the triangular partitioning mode (TPM), the video codec can evenly divide the CU into two triangular partitions by dividing it along the left-to-right diagonal of block 134 (which can be referred to as "diagonal partitioning") or along the right-to-left diagonal of block 138 (which can be referred to as "anti-diagonal partitioning"). The video encoder 200 and the video decoder 300 can use their own motion information to inter-predict each triangular partition (e.g., partitions 136A, 136B, 140A, 140B). According to JVET-N1002, only unidirectional prediction is allowed for each partition. That is, in JVET-N1002, each partition has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, like conventional bidirectional prediction, each CU only requires two motion compensation predictions.

[0077] The video encoder 200 and the video decoder 300 can derive the unidirectional prediction motion of each partition from the unidirectional prediction candidate list constructed by using the unidirectional prediction candidate list construction process described below.

[0078] The CU level flag may indicate whether the current CU is encoded / decoded using the triangular partitioning mode. If the triangular partitioning mode is used, then the video encoder 200 and the video decoder 300 may further encode / decode a flag indicating the direction (diagonal or anti-diagonal) of the triangular partitioning and two merge indices (one for each partition). After predicting each of the triangular partitions, the video encoder 200 and the video decoder 300 may adjust the sample values along the diagonal or anti-diagonal edge using a hybrid process with adaptive weights. The resulting hybrid prediction samples represent the prediction signal for the entire CU (i.e., the prediction block), and the video encoder 200 and the video decoder 300 may apply the transform and quantization processes to the entire CU as in other prediction modes. Finally, according to JVET-N1002, the motion field of the CU predicted using the triangular partitioning mode is stored in 4×4 units, as in the technique of blending along the triangular partitioning edge discussed below.

[0079] TPM is an example of a non-rectangular partitioning mode or a geometric partitioning mode. Other non-rectangular / geometric partitioning modes include, for example, partitioning the block using a polyline, which may be predefined or defined in the encoded / decoded data of the block. As another example of a geometric partitioning mode, a single line may be used to partition the block, where the line does not have to touch both of the opposite corner points of the block, e.g., as described in U.S. Patent No. 9,020,030, published on April 28, 2015. The line may be defined using slope and intercept values, or using two points representing the start and end points of the line relative to the block.

[0080] Figure 4 is a block diagram illustrating the spatial and temporal neighboring blocks for constructing the motion prediction candidate list. Figure 4 Shows the current block 150, which includes five spatial neighboring blocks (NBs) 152A - 152E and two temporal neighboring blocks (TNBs) 154A, 154B.

[0081] The uni-directional prediction candidate list includes five uni-directional prediction motion vector candidates. The video encoder 200 and the video decoder 300 may select from among five spatial neighboring blocks ( Figure 4 labeled as NB 152A - 152E in Figure 4Seven adjacent blocks (labeled TNB 154A, 154B in the figure) are used to derive a unidirectional prediction candidate list. The video encoder 200 and the video decoder 300 can collect the motion vectors of the seven adjacent blocks and place the motion vectors in the unidirectional prediction candidate list in the following order: First, the motion vectors of the unidirectional prediction adjacent blocks; then, for the bi-directional prediction adjacent blocks, the list zero (L0) motion vectors (i.e., the L0 motion vector part of the bi-directional prediction MV), the list one (L1) motion vectors (i.e., the L1 motion vector part of the bi-directional prediction MV), and the average motion vector of the L0 and L1 motion vectors of the bi-directional prediction MV. If the number of candidates is less than 5, zero motion vectors can be added to the end of the list.

[0082] The video encoder 200 and the video decoder 300 can further infer the motion of the TPM from the merge candidate list (e.g., the unidirectional prediction candidate list discussed above). Table 1 below shows example motion information for the TPM as described below:

[0083] Table 1

[0084] merge index L0 MV L1 MV 0 X 1 X 2 X 3 X 4 X

[0085] Given a merge candidate index, the video encoder 200 and the video decoder 300 can derive the unidirectional prediction motion vector for triangular partitioning from the merge candidate list. For a candidate in the merge list, its LX MV (where X is equal to the parity of the merge candidate index value) can be used as the unidirectional prediction motion vector for the triangular partitioning mode. These motion vectors are marked with "x" in Table 1. In the absence of the corresponding LX motion vector, the L(1-X) motion vector of the same candidate in the extended merge prediction candidate list can be used as the unidirectional prediction motion vector for the triangular partitioning mode. For example, assume the merge list consists of 5 bi-directional prediction motion sets. The TPM candidate list can consist of the L0 / L1 / L0 / L1 / L0 MVs of the 0th / 1st / 2nd / 3rd / 4th merge candidate from the first to the last. Then, the TPM mode requires signals for two different merge indices, one for triangular partitioning, to indicate the use of candidates in the TPM candidate list.

[0086] Figure 5 is a block diagram illustrating example weights that can be used in the blending process of the TPM. After predicting each triangular partition using its own motion, the video encoder 200 and the video decoder 300 can apply blending to the two prediction signals to derive samples around the diagonal or anti-diagonal edges. Figure 5 Shows example weights used in the blending process.

[0087] In JVET-N1002, motion vectors of CUs encoded and decoded in a triangular partitioning mode are stored in 4×4 units. Depending on the position of each 4×4 unit, a unidirectional prediction or a bidirectional prediction motion vector is stored. Let Mv1 and Mv2 be the unidirectional prediction motion vectors of partition 1 and partition 2, respectively. (Note that partition 1 and 2 are triangular blocks located at the upper right corner and the lower left corner, respectively, when the CU is partitioned from the upper left to the lower right (i.e., 45° partitioning, i.e., diagonal partitioning), and become triangular blocks located at the upper left corner and the lower right corner, respectively, when the CU is partitioned from the upper right to the lower left (i.e., 135° partitioning, i.e., anti-diagonal partitioning)). If the 4×4 unit is located in the Figure 5 non-weighted region shown in the example of Figure 5 , then Mv1 or Mv2 is stored for this 4×4 unit. Otherwise, if the 4×4 unit is located in this weighted region, then a bidirectional prediction motion vector is stored. The bidirectional prediction motion vector can be derived from Mv1 and Mv2 according to the following procedure of JVET-N1002:

[0088] 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional prediction motion vector.

[0089] 2) Otherwise, if Mv1 and Mv2 are from the same list, and without loss of generality, assume they are both from L0. In this case,

[0090] a. If the reference picture of Mv2 (or Mv1) appears in L1, then Mv2 (or Mv1) is converted to an L1 motion vector using the reference picture in L1. Then the two motion vectors are combined to form a bidirectional prediction motion vector;

[0091] b. Otherwise, instead of the bidirectional prediction motion, only the unidirectional prediction motion Mv1 is stored.

[0092] This disclosure recognizes that according to the procedure described in JVET-N1002, the current triangular motion compensation (MC) (i.e., triangular partitioning mode) does not use motion information from the motion buffer, which is contrary to all other prediction modes and is not conducive to hardware implementation. This disclosure describes techniques that can modify the TPM procedure described in JVET-N1002 such that triangular MC can be performed by using only the motion information in the motion buffer.

[0093] In addition, the present disclosure recognizes that the current motion storage design for the triangular MC of JVET-N1002 results in corner cases (in 4×N / N×4 blocks) without motion compensation for split 0 and split 1. That is, a corner case can be defined as a block having a size that meets a threshold (e.g., 4×N or N×4, where N is a positive integer value). The present disclosure further recognizes that the generation of bidirectional prediction motion vectors to be stored in the hybrid region of triangular partitioning to be described in JVET-N1002 is complex. This is especially true for step 2 discussed above, where two motion vectors come from the same reference picture list and one of them is mapped to the other list. The present disclosure describes techniques for simplifying this process. For example, in a corner case, when the motion information of both the lower split and the upper split of the current block refers to the same reference picture list, the video codec may store the set of motion information of the lower split of the current block.

[0094] The techniques of the present disclosure can be used to remove bidirectional prediction motion vectors for storage. That is, the video encoder 200 and the video decoder 300 can be configured to store only the unidirectional prediction motion vectors of the triangular partitioning mode, rather than storing the bidirectional prediction motion vectors. The video encoder 200 and the video decoder 300 can be configured individually or in any combination according to any one of the various examples discussed below:

[0095] · In some examples, the storage of bidirectional prediction motion vectors is completely removed to store only the motion vector Mv1, as in P1.

[0096] · In some examples, the storage of bidirectional prediction motion vectors is completely removed to store only the motion vector Mv2, as in P2.

[0097] · In some examples, the storage of bidirectional prediction motion vectors is completely removed to store Mv1 or Mv2, depending on the position within the block (e.g., blocks in the upper half store Mv1, and blocks in the lower half store Mv2).

[0098] · In some examples, the storage of bidirectional prediction motion vectors is completely removed to store Mv1 or Mv2, depending on the triangular partitioning direction (e.g., store MV1 for 45° partitioning and store MV2 for 135° partitioning).

[0099] · In some examples, the bidirectional prediction motion vectors are removed according to the block size (e.g., only for the corner cases of 4×N and N×4 blocks).

[0100] · In some examples, the bidirectional prediction motion vectors are removed according to the block size and the position within the block (e.g., only for the corner cases of 4×N and N×4, and for the first and last PUs).

[0101] Video encoder 200 and video decoder 300 may also be configured with a modified algorithm for generating bidirectional prediction motion vectors, which may simplify the process of generating bidirectional prediction motion vectors. Video encoder 200 and video decoder 300 may be configured individually or in any combination according to any one of the various examples discussed below:

[0102] · In some examples, when both Mv1 and Mv2 are from the same list, only Mv1 is stored.

[0103] · In some examples, when both Mv1 and Mv2 are from the same list, only Mv2 is stored.

[0104] · In some examples, when both Mv1 and Mv2 are from the same list, Mv1 or Mv2 is stored depending on the position within the block (e.g., Mv1 is stored for blocks in the upper half and Mv2 is stored for blocks in the lower half).

[0105] · In some examples, when both Mv1 and Mv2 are from the same list, Mv1 or Mv2 is stored depending on the triangle partitioning direction (e.g., Mv1 is stored for a 45° (diagonal) partition and Mv2 is stored for a 135° (anti - diagonal) partition).

[0106] In some examples, when the bidirectional prediction merge candidate in the merge list has a non - 0.5BCW weight value, video encoder 200 and video decoder 300 do not use the motion information corresponding to the reference picture list coupled with the lower weight value as a valid TPM candidate. Specifically, when the following conditions are met, the motion information of the reference picture list Lx (where x is 0 or 1) corresponding to the bidirectional prediction merge candidate may be included in the TPM candidate list:

[0107] · If the bidirectional prediction merge candidate has a 0.5BCW weight value, x is determined by a parity process (i.e., the same as the above - mentioned TPM candidate list construction process).

[0108] · Otherwise, if the bidirectional prediction merge candidate has a non - 0.5BCW weight value, x is determined by the list assigned the larger BCW weight value.

[0109] Video encoder 200 and video decoder 300 may encode and decode sequence, picture, slice group, stripe, and / or CTU - level flags in the bitstream, which indicate the use of the above - mentioned techniques for selecting valid TPM candidates.

[0110] In addition, the present disclosure recognizes that triangle partitioning patterns are not allowed in P - stripes (i.e., unidirectional inter - prediction stripes). There are two reasons. First, the hybrid processes discussed above (e.g., regarding Figure 5)Two motion vectors are used. However, in a P slice, only one motion vector is available because for a P slice, bidirectional prediction is not enabled and not allowed. Second, the motion field storage can store bidirectional motion vectors (e.g., motion vectors from list 0 and list 1), which is not allowed in a P slice.

[0111] Video encoder 200 and video decoder 300 can be configured to enable a non-rectangular partitioning mode, such as a triangular partitioning mode, for blocks (e.g., coding units (CUs)) of a P slice. Specifically, using the motion vector storage technique of the simplified triangular partitioning mode described above (also described in U.S. Provisional Patent Application No. 62 / 857,584 filed on June 5, 2019) to store only unidirectional motion vectors, and by disabling the hybrid operation of the bidirectional motion vector triangular partitioning mode, the triangular partitioning mode can be performed for blocks of a P slice. That is, video encoder 200 and video decoder 300 can be configured to use the triangular partitioning mode to predict blocks of a P slice by performing the technique for simplified motion vector storage discussed above and by disabling the hybrid operation.

[0112] Figure 6 is a block diagram illustrating example weights that can be used in a hybrid process for a TPM prediction block having a stride width equal to two samples. According to the techniques of the present disclosure, video encoder 200 and video decoder 300 can disable the hybrid operation applied to 4×4 units along the boundary between two triangular blocks in a CU. When it is indicated that the hybrid operation is disabled, video encoder 200 and video decoder 300 can apply any one or all of the following techniques. Regardless of which technique is applied, the weighted value assigned to each corresponding sample on P2 can be set to be equal to 1 minus the weighted value assigned to the corresponding sample on P1.

[0113] · In some examples, when hybridizing samples from partitioned P1 and P2, video encoder 200 and video decoder 300 can reset the weighted value assigned to each corresponding sample on P1 to be equal to:

[0114] о (#a) 8 / 8 if they are greater than 4 / 8,

[0115] о (#b) 4 / 8 if they are equal to 4 / 8, or

[0116] о (#c) 0 / 8 if they are less than 4 / 8.

[0117] · In some examples, the configuration of #b (i.e., 4 / 8) is replaced with 8 / 8.

[0118] · In some examples, the configuration of #b (i.e., 4 / 8) is replaced with 0 / 8.

[0119] · In some examples, the configuration of #b (i.e., 4 / 8) is replaced by 0 / 8 or 8 / 8, depending on the partitioning direction of the triangle (e.g., 8 / 8 if partitioned at 45°, and 0 / 8 if partitioned at 135°).

[0120] · In some examples, when the width of the stride with a weight of 4 / 8 as in #b is greater than 1 sample (e.g., if the aspect ratio is N or 1 / N, the stride width equals N samples, where N = 2, 4, 8, …), if the samples on half of the stride are spatially closer to the corner point of P1, a weighted value equal to 8 / 8 is assigned to them, and a weighted value of 0 / 8 is assigned to the samples on the other half of the stride. For example, Figure 6 illustrates that if these samples are located closer to the corner point of P1, the weighted value assigned to each corresponding sample on P1 is reset to equal 8 / 8, while the rest are assigned 0 / 8.

[0121] Video encoder 200 and video decoder 300 may encode sequence, picture, slice group, stripe, and / or CTU-level flags in the bitstream, which indicate whether to enable or disable these simplified hybrid techniques.

[0122] In some examples, video encoder 200 and video decoder 300 may disable motion vectors pointing to fractional pixel (pel) positions. Video encoder 200 and video decoder 300 may encode and decode sequence, slice group, stripe, and / or CTU-level flags in the bitstream, which indicate using the above techniques to enable or disable fractional pixel motion vectors. When this new flag is enabled, all motion vectors should have integer precision, and thus fractional interpolation of sharp prediction signals can be avoided. How this new flag can change the operation of video encoder 200 and video decoder 300 with various inter prediction modes is elaborated in detail below:

[0123] · Conventional inter mode: The CABAC engine skips the encoding / parsing of bits representing fractional precision MVDs from the AMVR (Adaptive Motion Vector Resolution) syntax. Thus, AMVR only supports non-fractional pixel precision.

[0124] · Regular affine mode: The CABAC engine skips the encoding / parsing of bits representing fractional precision MVDs from the AMVR syntax. Thus, AMVR only supports non-fractional pixel precision. Additionally, in some examples, before using the derived affine motion for motion compensation, it can be clipped or rounded (by a predefined offset value) to non-fractional precision.

[0125] · Regular merge mode: Before using its candidate motion for motion compensation, its candidate motion can be clipped or rounded (by a predefined offset value) to non-fractional precision.

[0126] · TPM merge mode: Before using the reference merge candidates to construct the TPM candidate list, the reference merge candidates can be truncated or rounded (by a predefined offset value) to non-fractional precision. Additionally, in some examples, it can be inferred that the flag indicating the use of a simplified mixing method is enabled.

[0127] · MMVD mode: Before using the reference merge candidates for the basis vectors used to form the MMVD mode, they can be cropped or rounded (by a predetermined offset value) to non-fractional precision. Additionally, in some examples, the fractional offset values in the MMVD distance table can be disabled.

[0128] · CIIP mode: Before using the reference merge candidates for motion compensation, they can be truncated or rounded (by a predetermined offset value) to non-fractional precision. In some examples, when a new flag is enabled, the CIIP mode can be completely disabled in the bitstream.

[0129] Additionally, in some examples, when the flag indicates disabling fractional precision motion vectors, the adaptive loop filter and the deblocking filter can be disabled.

[0130] Figure 7 is a block diagram of an example video encoder 200 that can perform the techniques of the present disclosure. Provided Figure 7 is for explanatory purposes and should not be considered a limitation on the techniques widely exemplified and described in the present disclosure. For explanatory purposes, the present disclosure describes the video encoder 200 in the context of video coding standards (such as the HEVC video coding standard and the H.266 video coding standard under development). However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.

[0131] In Figure 7 the example of, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 can be implemented in one or more processors or processing circuits. Additionally, the video encoder 200 can include additional or alternative processors or processing circuits to perform these and other functions.

[0132] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104( Figure 1 ). The DPB 218 may be used as a reference picture memory that stores reference video data for use by the video encoder 200 in predicting subsequent video data. The video data memory 230 and the DPB 218 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or separate memory devices. In various examples, as illustrated, the video data memory 230 may be on-chip with other components of the video encoder 200 or off-chip relative to those components.

[0133] In this disclosure, a reference to the video data memory 230 should not be construed as limited to a memory internal to the video encoder 200 unless specifically described as such, or limited to a memory external to the video encoder 200 unless specifically described as such. Instead, a reference to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage of outputs from various units of the video encoder 200.

[0134] is illustrated Figure 7 Various units to assist in understanding the operations performed by the video encoder 200. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides a specific function and is preset in the operations it can perform. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in the operations it can perform. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuitry may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by fixed-function circuitry are generally immutable. In some examples, one or more units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.

[0135] Video encoder 200 may include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where software executed by programmable circuitry is used to perform the operations of video encoder 200, memory 106( Figure 1 ) may store the object code of the software that video encoder 200 receives and executes, or another memory (not shown) within video encoder 200 may store such instructions.

[0136] Video data memory 230 is configured to store the received video data. Video encoder 200 may retrieve pictures of video data from video data memory 230 and provide the video data to residual generation unit 204 and mode selection unit 202. The video data in video data memory 230 may be the original video data to be encoded.

[0137] Mode selection unit 202 includes motion estimation unit 222, motion compensation unit 224, and intra prediction unit 226. Mode selection unit 202 may include additional functional units to perform video prediction according to other prediction modes. As an example, mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of motion estimation unit 222 and / or motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0138] Mode selection unit 202 generally coordinates multiple encoding processes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include the CTU to CU partitioning, the prediction mode for the CU, the transform type for the CU residual data, the quantization parameter for the CU residual data, etc. Mode selection unit 202 may ultimately select a combination of encoding parameters having a better rate-distortion value than other tested combinations.

[0139] Video encoder 200 may divide the pictures retrieved from video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. Mode selection unit 202 may divide the CTUs of the picture according to a tree structure (such as the QTBT structure or quadtree structure of HEVC described above). As described above, video encoder 200 may form one or more CUs by dividing CTUs according to a tree structure. This CU may also generally be referred to as a "video block" or "block".

[0140] Typically, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., an overlapping portion of a current CU or a PU and a TU in HEVC). To perform inter prediction for the current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 may calculate values representing how similar a potential reference block is to the current block, for example, based on the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 generally may perform these calculations using the sample-by-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block having the lowest value obtained from these calculations, which indicates the reference block that most closely matches the current block.

[0141] The motion estimation unit 222 may form one or more motion vectors (MVs), which define the position of the reference block in the reference picture relative to the current block in the current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve the data of the reference block. As another example, if the motion vectors have fractional sample accuracy, the motion compensation unit 224 may interpolate the values for the prediction block according to one or more interpolation filters. Additionally, for bi-directional inter prediction, the motion compensation unit 224 may retrieve the data for two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.

[0142] The motion compensation unit 224 may use a geometric or non-rectangular partitioning mode (e.g., a triangle partitioning mode (TPM)) to generate a prediction block, either alone or in any combination, according to any one of the various techniques of the present disclosure discussed above. For example, the motion compensation unit 224 may use a non-rectangular partitioning mode to divide the current block into two partitions, each of which has its own respective set of motion information (received from the motion estimation unit 222). The motion compensation unit 224 may use the first set of motion information to form a first prediction block for the first partition and use the second set of motion information to form a second prediction block for the second partition. In some examples, the motion compensation unit 224 may perform a blending process to blend the sample values along the boundary between the first partition and the second partition, for example, as described above with respect to Figure 5and 6 As discussed above. Alternatively, if the current block forms part of a uni - directional inter - prediction strip (i.e., a P - strip), then the motion compensation unit 224 may disable the hybrid operation discussed above when generating the prediction block for the current block.

[0143] As another example, for intra - prediction or intra - prediction coding / decoding, the intra - prediction unit 226 may generate a prediction block based on samples adjacent to the current block. For example, for the directional mode, the intra - prediction unit 226 may typically mathematically combine the values of adjacent samples and fill in the calculated values in the defined direction across the current block to produce the prediction block. As another example, for the DC mode, the intra - prediction unit 226 may calculate the average of the adjacent samples of the current block and generate a prediction block to contain the average value obtained for each sample of the prediction block.

[0144] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original, un - decoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample - by - sample difference between the current block and the prediction block. The resulting sample - by - sample difference defines the residual block of the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block to generate the residual block using residual differential pulse coding modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0145] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of a CU may refer to the size of the luminance coding / decoding block of the CU, while the size of a PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra - prediction, and 2N×2N, 2N×N, N×2N, N×N or similar symmetric PU sizes for inter - prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning of PU sizes of 2N×nU, 2N×nD, nL×2N and nR×2N for inter - prediction.

[0146] In an example where the mode selection unit does not further divide a CU into PUs, each CU may be associated with a luminance coding / decoding block and a corresponding chrominance coding / decoding block. Similarly, the size of a CU may refer to the size of the luminance coding / decoding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N or N×2N.

[0147] For other video coding and decoding techniques, such as intra block copy mode coding and decoding, affine mode coding and decoding, and linear model (LM) mode coding and decoding, as a few examples, the mode selection unit 202 generates a prediction block for the current block being encoded via corresponding units associated with the coding and decoding techniques. In some examples, such as palette mode coding and decoding, the mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating the manner in which the block is to be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.

[0148] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the prediction block and the current block.

[0149] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, such as a primary transform and a secondary transform, such as a rotation transform. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0150] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0151] The inverse quantization unit 210 and the inverse transform processing unit 212 can respectively apply inverse quantization and inverse transform to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block (although there may be a certain degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 can add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0152] The filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, the filter unit 216 can perform a deblocking operation to reduce block effect artifacts along the CU edge. In some examples, the operation of the filter unit 216 can be skipped.

[0153] The video encoder 200 stores the reconstructed block in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed block in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed block in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference picture from the DPB 218, which is formed by the reconstructed (and possibly filtered) block, to perform inter prediction on the blocks of the subsequent encoded pictures. In addition, the intra prediction unit 226 can use the reconstructed blocks in the DPB 218 of the current picture to perform intra prediction on other blocks in the current picture.

[0154] Generally, the entropy coding unit 220 can perform entropy coding on the syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 can perform entropy coding on the quantized transform coefficient block from the quantization unit 208. As another example, the entropy coding unit 220 can perform entropy coding on the prediction syntax elements from the mode selection unit 202 (e.g., motion information for inter prediction or intra mode information for intra prediction). The entropy coding unit 220 can perform one or more entropy coding operations on the syntax elements as another example of video data to generate the entropy-coded data. For example, the entropy coding unit 220 can perform context-adaptive variable length coding and decoding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding and decoding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval segmentation entropy (PIPE) coding and decoding operations, exponential Golomb coding operations, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 can operate in a bypass mode, in which the syntax elements are not entropy coded.

[0155] The entropy coding unit 220 may perform entropy coding on the motion information of the current block. For example, if a non-rectangular partitioning mode (e.g., TPM) is used to predict the current block, the entropy coding unit 220 may perform entropy coding on the motion information of each resulting partition. Generally, the entropy coding unit 220 may perform entropy coding on the motion information of the upper partition before performing entropy coding on the motion information of the lower partition. The upper partition may be considered as the partition having more samples along the upper edge of the current block, and the lower partition may be considered as the partition having more samples along the lower edge of the block.

[0156] When entropy coding the motion information, the motion information generally may indicate a motion vector (e.g., X component and Y component), a reference picture list including the reference pictures referred to by the motion vector, and a reference picture index identifying the reference pictures in the reference picture list. In the merge mode, the entropy coding unit 220 may perform entropy coding on the merge index identifying the motion vector predictor candidates (i.e., merge candidates) in the motion vector predictor candidate list, and may use each of the motion vector, the reference picture list, and the reference picture index from the motion vector predictor candidates. In the AMVP mode, the entropy coding unit 220 may perform entropy coding on the motion vector predictor candidate, the motion vector difference data (e.g., X offset and Y offset) relative to the motion vector predictor candidate, the reference picture list identifier, and the reference picture index.

[0157] To perform entropy coding on the motion information in this way, the entropy coding unit 220 may retrieve the previously decoded motion information candidates stored in, for example, the DPB 218. According to the techniques of the present disclosure, when the current block is partitioned using a non-rectangular partitioning mode and the motion information of the two partitions refers to a common reference picture list, the video encoder 200 may store the motion information of the second partition in the DPB 218 for the current block (without storing the motion information of the first partition). Thus, when encoding or decoding a subsequent block of the current block, the entropy coding unit 220 may select the motion information of the second partition as a motion vector predictor candidate to predict the motion information of the subsequent block.

[0158] The video encoder 200 may output a bitstream that includes the entropy-coded syntax elements required to reconstruct the blocks of the strip or picture. Specifically, the entropy coding unit 220 may output the bitstream.

[0159] The above operations are described with respect to blocks. This description should be understood as operations for the luminance coding / decoding blocks and / or the chrominance coding / decoding blocks. As described above, in some examples, the luminance coding / decoding blocks and the chrominance coding / decoding blocks are the luminance component and the chrominance component of the CU. In some examples, the luminance coding / decoding blocks and the chrominance coding / decoding blocks are the luminance component and the chrominance component of the PU.

[0160] In some examples, for chrominance coding / decoding blocks, operations performed for luma coding / decoding blocks need not be repeated. As an example, operations of identifying motion vectors (MVs) and reference pictures for luma coding / decoding blocks need not be repeated for identifying MVs and reference pictures for chrominance blocks. Instead, the MVs for luma coding / decoding blocks can be scaled to determine the MVs for chrominance blocks, and the reference pictures can be the same. As another example, for luma coding / decoding blocks and chrominance coding / decoding blocks, the intra prediction process can be the same.

[0161] In this way, video encoder 200 represents an example of a device configured to encode and decode video data, including a memory configured to store the video data; and one or more processors implemented in circuitry and configured to: encode a first set of motion information for a current block of the video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding the first set of motion information, encode a second set of motion information for the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, store the second set of motion information for the current block in the memory; and use the stored second set of motion information to predict subsequent motion information for a subsequent block of the video data adjacent to the current block.

[0162] Figure 8 is a block diagram illustrating an example video decoder 300 that can execute the techniques of the present disclosure. Provided Figure 8 is for explanatory purposes and does not limit the techniques of the broad examples and descriptions in the present disclosure. For explanatory purposes, the present disclosure describes video decoder 300 according to techniques of JEM, VCC, and HEVC. However, the techniques of the present disclosure can be performed by video coding / decoding devices configured for other video coding / decoding standards.

[0163] In Figure 8 an example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or processing circuitry. Additionally, video decoder 300 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0164] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units to perform prediction according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0165] The CPB memory 320 may store video data to be decoded by components of the video decoder 300, such as an encoded video bitstream. For example, the video data stored in the CPB memory 320 may be obtained from a computer-readable medium 110 ( Figure 1 ). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the decoded pictures, such as temporary data representing the outputs of various units of the video decoder 300. The DPB 314 generally stores decoded pictures, and the video decoder 300 may output the decoded pictures and / or use them as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed of any of various memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or different memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300 or off-chip relative to those components.

[0166] Additionally or alternatively, in some examples, the video decoder 300 may retrieve the encoded and decoded video data from a memory 120 ( Figure 1 ). That is, the memory 120 may store data as discussed above with respect to the CPB memory 320. Similarly, when some or all of the functions of the video decoder 300 are implemented in software to be executed by the processing circuitry of the video decoder 300, the memory 120 may store the instructions to be executed by the video decoder 300.

[0167] is illustrated Figure 8 the various units shown to assist in understanding the operations performed by the video decoder 300. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 7, A fixed - function circuit refers to a circuit that provides a specific function and is preset in terms of the operations it can perform. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality among the operations it can execute. For example, a programmable circuit can execute software or firmware to cause the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed - function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed - function circuit are generally immutable. In some examples, one or more units can be different circuit blocks (fixed - function or programmable), and in some examples, one or more units can be integrated circuits.

[0168] Video decoder 300 may include an ALU, an EFU, digital circuits, analog circuits, and / or a programmable core formed by a programmable circuit. In an example where the operation of video decoder 300 is performed by software executed on the programmable circuit, on - chip or off - chip memory may store the instructions (e.g., object code) of the software that video decoder 300 receives and executes.

[0169] The entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to regenerate syntax elements. The prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0170] Generally, video decoder 300 reconstructs pictures on a block - by - block basis. Video decoder 300 can perform the reconstruction operation for each block individually (where the block currently being reconstructed, i.e., the decoded block, can be referred to as the "current block").

[0171] The entropy decoding unit 302 can perform entropy decoding on the syntax elements that define the quantized transform coefficients of the quantized transform coefficient block and the transform information (e.g., quantization parameter (QP) and / or (multiple) transform mode indicators). The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the quantization level, and similarly, determine the inverse quantization level to be applied by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bit - by - bit left - shift operation to inverse - quantize the quantized transform coefficients. The inverse quantization unit 306 can thereby form a transform coefficient block containing transform coefficients.

[0172] After the inverse quantization unit 306 forms a transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the coefficient block.

[0173] The entropy decoding unit 302 may entropy decode the motion information of the current block. For example, if the current block is predicted using a non-rectangular partitioning mode (e.g., TPM), then the entropy decoding unit 302 may entropy decode the motion information of each resulting partition. Generally, the entropy decoding unit 302 may entropy decode the motion information of the upper partition before entropy decoding the motion information of the lower partition. That is, the bitstream may include the motion information of the upper partition before including the motion information of the lower partition. The upper partition may be considered a partition having more samples along the upper edge of the current block, while the lower partition may be considered a partition having more samples along the lower edge of the block.

[0174] When entropy decoding motion information, the motion information may generally indicate a motion vector (e.g., an X component and a Y component), a reference picture list including the reference picture to which the motion vector refers, and a reference picture index identifying the reference picture in the reference picture list. In the merge mode, the entropy decoding unit 302 may entropy decode a merge index identifying a motion vector predictor candidate (i.e., a merge candidate) in the motion vector predictor candidate list and use each of the motion vector, the reference picture list, and the reference picture index from the motion vector predictor candidate. In the AMVP mode, the entropy decoding unit 302 may entropy decode a motion vector predictor candidate, motion vector difference data (e.g., an X offset and a Y offset) relative to the motion vector predictor candidate, a reference picture list identifier, and a reference picture index.

[0175] To entropy decode motion information in this way, the entropy decoding unit 302 may retrieve previously decoded motion information candidates stored in, for example, the DPB 314. According to the techniques of the present disclosure, when the current block is partitioned using a non-rectangular partitioning mode and the motion information of the two partitions refers to a common reference picture list, the video decoder 300 may store the motion information of the second partition in the DPB 314 for the current block (without storing the motion information of the first partition). Thus, when encoding or decoding a subsequent block of the current block, the entropy decoding unit 302 may select the motion information of the second partition as a motion vector predictor candidate to predict the motion information of the subsequent block.

[0176] In addition, the prediction processing unit 304 generates a prediction block based on a prediction information syntax element entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter prediction, the motion compensation unit 316 can generate a prediction block. In this case, the prediction information syntax element can indicate a reference picture in the DPB 314 from which a reference block is retrieved, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture.

[0177] The motion compensation unit 316 can generally perform an inter prediction process in a manner generally similar to that described with respect to the motion compensation unit 224 ( Figure 7 ). The motion compensation unit 316 can generate a prediction block using a non-rectangular (e.g., geometric) partitioning mode (e.g., triangle partitioning mode (TPM)) alone or in any combination according to any one of the various techniques of the present disclosure as discussed above. In some examples, the motion compensation unit 316 can perform a blending operation to smooth the sample values along the partitioning boundary of two partitions, e.g., as explained above with respect to Figure 5 and 6 . Alternatively, if the current block forms part of a unidirectional inter prediction strip (i.e., a P-strip), then the motion compensation unit 316 can disable the blending operation discussed above when generating a prediction block for the current block.

[0178] As another example, if the prediction information syntax element indicates that the current block is intra prediction, the intra prediction unit 318 can generate a prediction block according to the intra prediction mode indicated by the prediction information syntax element. Similarly, the intra prediction unit 318 can generally perform an intra prediction process in a manner substantially similar to that described with respect to the intra prediction unit 226 ( Figure 7 ). The intra prediction unit 318 can retrieve data of adjacent samples of the current block from the DPB 314.

[0179] The reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0180] The filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 can perform a deblocking operation to reduce block effect artifacts along the edges of the reconstructed block. The operation of the filter unit 312 is not necessarily performed in all examples.

[0181] Video decoder 300 may store the reconstructed blocks in DPB 314. As described above, DPB 314 may provide reference information to prediction processing unit 304, such as samples of the current picture for intra prediction and samples of previously decoded pictures for subsequent motion compensation. Additionally, video decoder 300 may output decoded pictures from DPB 314 for subsequent presentation on a display device, such as Figure 1 display device 118.

[0182] In this way, video decoder 300 represents an example of a device for encoding and decoding video data, which includes a memory configured to store video data; and one or more processors, implemented in circuitry and configured to: encode and decode a first set of motion information for a current block of the video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding and decoding the first set of motion information, encode and decode a second set of motion information for the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, store the second set of motion information for the current block in the memory; and use the stored second set of motion information to predict subsequent motion information for a subsequent block of the video data adjacent to the current block.

[0183] Figure 9 is a flowchart illustrating an exemplary method for encoding a current block according to the techniques of the present disclosure. The current block may include a current CU. Although described with respect to video encoder 200 ( Figure 1 and 7 ), it should be understood that other devices may be configured to perform methods similar to the Figure 9 method.

[0184] In this example, video encoder 200 initially predicts a current block (350). For example, video encoder 200 may form a prediction block for the current block. Video encoder 200 may use a non-rectangular partitioning mode (e.g., triangular partitioning mode (TPM)) to form the prediction block according to any of the various techniques of the present disclosure. Additionally, if the current block forms part of a uni-directional inter prediction strip (i.e., a P-strip), then video encoder 200 may further disable the hybrid operation as discussed above when generating the prediction block for the current block. Video encoder 200 may then compute a residual block (352) for the current block. To compute the residual block, video encoder 200 may compute the difference between the original unencoded block and the prediction block for the current block. Then, video encoder 200 may transform and quantize the coefficients of the residual block (354). Next, video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, video encoder 200 may entropy encode the coefficients (358). For example, video encoder 200 may use CAVLC or CABAC to encode the coefficients. Video encoder 200 may then output the entropy encoded data of the block (360).

[0185] In this way, Figure 9 The method of represents an example of a method for encoding and decoding video data, which includes: encoding a first set of motion information for a current block of video data, the current block being partitioned into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding the first set of motion information, encoding a second set of motion information for the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, storing the second set of motion information for the current block; and using the stored second set of motion information to predict subsequent motion information for a subsequent block of the video data adjacent to the current block.

[0186] Figure 10 is a flowchart illustrating an example method for decoding a current block according to the techniques of the present disclosure. The current block may include a current CU. Although described with respect to video decoder 300( Figure 1 and 8 ), it should be understood that other devices may be configured to perform methods similar to Figure 10 .

[0187] Video decoder 300 may receive entropy-coded data of a current block, such as entropy-coded prediction information and entropy-coded data (370) of coefficients of a residual block corresponding to the current block. Video decoder 300 may entropy-decode the entropy-coded data to determine prediction information for the current block and to regenerate coefficients of the residual block (372). Video decoder 300 may predict the current block (374) using, for example, an intra or inter prediction mode indicated by the prediction information of the current block to calculate a predicted block of the current block. Video decoder 300 may form the predicted block using, for example, a non-rectangular partitioning mode (e.g., triangular partitioning mode (TPM)) according to any of the various techniques of the present disclosure. Further, if the current block forms part of a unidirectional inter prediction strip (i.e., a P-strip), then video decoder 300 may further disable the hybrid operation as discussed above when generating the predicted block of the current block. Video decoder 300 may then inverse-scan the regenerated coefficients (376) to create a block of quantized transform coefficients. Video decoder 300 may then inverse-quantize and inverse-transform the coefficients to produce a residual block (378). Video decoder 300 may finally decode the current block by combining the predicted block and the residual block (380).

[0188] In this way, Figure 10 The method of represents an example of a method of encoding and decoding video data, which includes: encoding a first set of motion information of a current block of video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding the first set of motion information, encoding a second set of motion information of the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, storing the second set of motion information of the current block; and using the stored second set of motion information to predict subsequent motion information of a subsequent block adjacent to the current block of the video data.

[0189] Figure 11 is a flowchart illustrating an example method of decoding video data according to the techniques of the present disclosure. For purposes of example, the method of is explained with respect to the video decoder 300 of Figure 1 and 8 The video encoder 200 may perform a generally similar method for encoding video data, as explained below. Figure 11

[0190] ​Initially, video decoder 300 uses a non-rectangular partitioning mode (400) (e.g., TPM) to partition a current block. Video decoder 300 may also decode a first motion vector (MV) of a first partition of the current block (402). Specifically, a video bitstream containing the current block may include a first motion information set and a second motion information set, where the first motion information set exists in the bitstream before the second motion information set. Thus, video decoder 300 may parse the bitstream and obtain the first motion information set before the second motion information set. Since the first motion information set is encountered before the second motion information set, video decoder 300 may determine that the first motion information set corresponds to an upper partition of the current block (i.e., a partition having more samples along the upper boundary of the current block). The upper partition is also referred to hereinafter as the "first partition".

[0191] To decode the first motion information set, video decoder 300 may construct a motion vector predictor candidate list according to, for example, the merge mode or the AMVP mode. The first motion information set may include a candidate index identifying a candidate in the motion vector predictor candidate list. If the first motion information set is encoded in the merge mode, video decoder 300 may determine the motion vector, the reference picture list, and the reference picture index of the first motion information set from the motion vector predictor candidate identified by the candidate index. If the first motion information set is encoded in the AMVP mode, video decoder 300 may decode the motion vector difference, the reference picture list identifier, and the reference picture index from the data of the bitstream, and reconstruct the motion vector by adding the motion vector difference to the components of the motion vector of the motion vector predictor candidate identified by the candidate index. Video decoder 300 may then use the first motion information set containing the first motion vector to predict the first partition of the current block (404).

[0192] Video decoder 300 may also decode a second motion vector of the current block corresponding to a second partition (406). Video decoder 300 may decode the second motion vector by decoding the second motion information set as discussed above with respect to the first motion information set. Video decoder 300 may then use the second motion vector to predict the second partition of the current block (408), where the second partition corresponds to a lower partition (i.e., a partition having more samples along the lower boundary of the current block).

[0193] In this example, video decoder 300 determines that a first motion vector and a second motion vector reference a common reference picture list (410). That is, in this example, video decoder 300 determines that the reference picture list identifier of the first motion information set and the reference picture list identifier of the second motion information set correspond to the same reference picture list. The reference picture list can be, for example, RefPicList0 or RefPicList1. In response to determining that two partitions of the current block (predicted using a non-rectangular partitioning mode) are predicted using the same reference picture list (i.e., the first motion information set and the second motion information set each contain the same reference picture list identifier), video decoder 300 stores the second MV of the current block (412), for example, in DPB 314. Video decoder 300 may store the second MV for all sub-blocks of the current block and not store the first MV of the current block to be used as a subsequent motion vector prediction sub-candidate. Although not described in Figure 11 , video decoder 300 may continue to use the first predicted partition and the second predicted partition to decode the current block, for example, as discussed above with respect to Figure 10 .

[0194] Thus, when decoding a subsequent block (e.g., a spatially adjacent block), video decoder 300 may add the stored second MV as a candidate motion vector predictor (MVP) candidate for the subsequent block (414). For example, the “current block” mentioned above may correspond to Figure 4 ’s NB 152A, and the “subsequent block” may correspond to Figure 4 ’s current block 150, in which case NB 152A is adjacent to current block 150. Thus, video decoder 300 may use the motion information of NB 152A as a motion vector prediction sub-candidate to be added to the motion vector prediction candidate list of current block 150.

[0195] Video decoder 300 may then use the second motion vector to decode the motion information of the subsequent block (416). For example, again with respect to Figure 4, the video decoder 300 may construct a list of motion vector predictor candidates including the motion information of NB 152 and TNB 154A or TNB 154B. The video decoder 300 may further decode a motion vector predictor candidate index to identify one of NB152 or TNB 154A or TNB 154B as a motion vector predictor candidate for the current block 150. Again, assuming that the "current block" mentioned above corresponds to NB 152A, the "subsequent block" mentioned above corresponds to the current block 150, and the candidate index identifies NB 152A, the video decoder 300 may use the motion information (including the motion vector) of NB 152A to decode the motion information of the current block 150. For example, in the merge mode, the video decoder 300 may use the motion vector of NB 152A (i.e., the stored second motion vector of the "current block" mentioned above) as the motion vector of the current block 150, and in the AMVP mode, the video decoder 300 may use the motion vector of NB 152A as a motion vector predictor and decode a motion vector difference representing an offset of a component to be applied to the motion vector predictor to form the motion vector of the current block 150.

[0196] In this way, Figure 11 The method of represents an example of a method for encoding and decoding video data, which includes: encoding and decoding a first set of motion information of a current block of video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning mode, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding and decoding the first set of motion information, encoding and decoding a second set of motion information of the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, storing the second set of motion information of the current block; and using the stored second set of motion information to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.

[0197] As described above, the video encoder 200 may be configured to perform a method substantially similar to Figure 11 the method. In a similar method, the video encoder 200 may encode a first motion vector and a second motion vector after using the motion vectors to predict the first partition and the second partition, and the video encoder 200 may use the second motion vector to encode the motion information of the subsequent block, for example, in the merge mode or the AMVP mode.

[0198] It should be recognized that, according to the examples, certain actions or events of any of the techniques described herein may be performed in a different sequence, may be added together, combined, or omitted (e.g., not all described actions or events are necessary for the practice of the techniques). Additionally, in some examples, the actions or events may be performed simultaneously, such as by multithreaded processing, interrupt processing, or multiple processors, rather than sequentially.

[0199] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this way, the computer-readable medium generally corresponds to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0200] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave from a website, server, or other remote source, then the medium includes coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but rather are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks generally reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0201] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms "processor" and "processing circuitry" may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Moreover, the techniques may be implemented entirely in one or more circuits or logic elements.

[0202] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a group of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but are not necessarily implemented by distinct hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperating hardware units including one or more processors as described above, in conjunction with appropriate software and / or firmware.

[0203] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for encoding and decoding video data, the method comprises: encoding and decoding a first motion information set of a current block of video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning mode, the first motion information set referring to a reference picture list and being associated with the first partition; after encoding and decoding the first motion information set, encoding and decoding a second motion information set of the current block, the second motion information set referring to the reference picture list and being associated with the second partition; in response to both the first motion information set and the second motion information set referring to the reference picture list, storing the second motion information set of the current block without storing the first motion information set of the current block; and using the stored second motion information set to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.

2. The method according to claim 1, wherein storing the second motion information set includes storing the second motion information set for all sub-blocks of the current block.

3. The method according to claim 1, further comprising determining that the size of the current block meets a threshold, wherein storing the second motion information set includes storing the second motion information set in response to determining that the size of the current block meets the threshold.

4. The method according to claim 3, wherein the threshold includes 4×N or N×4, and N is a positive integer value.

5. The method according to claim 1, wherein predicting subsequent motion information of a subsequent block includes: forming a motion prediction candidate list for the subsequent block, including adding the second motion information set to the motion prediction candidate list; selecting the second motion information set from the motion prediction candidate list; and using the second motion information set to predict the subsequent motion information of the subsequent block.

6. The method according to claim 1, wherein the current block forms part of a unidirectional inter prediction strip P-strip, and the method further comprises disabling a hybrid operation for the current block.

7. The method according to claim 1, wherein the first partition includes a first triangular partition, the second partition includes a second triangular partition, and the non-rectangular partitioning mode includes a triangular partitioning mode TPM.

8. The method according to claim 1, further comprising generating a prediction block of the current block using the first motion information set and the second motion information set without performing a hybrid operation between the first partition and the second partition.

9. The method according to claim 1, wherein generating a prediction block includes: resetting a hybrid weight greater than 4 / 8 to be equal to 8×8; resetting a hybrid weight less than 4 / 8 to be equal to 0×8; and using the reset hybrid weights to combine the first partition and the second partition.

10. The method according to claim 9, further comprising resetting a hybrid weight equal to 4 / 8 to be equal to 8 / 8.

11. The method according to claim 9, further comprising resetting a hybrid weight equal to 4 / 8 to be equal to 0 / 8.

12. The method according to claim 9, further comprising: Resetting a hybrid weight equal to 4 / 8 to be equal to 0 / 8 or 8 / 8 according to the partitioning direction of the current block.

13. The method according to claim 8, wherein, generating a prediction block includes generating the prediction block using conventional inter prediction, and the method further includes skipping context adaptive binary arithmetic coding (CABAC) of bits representing fractional precision motion vector difference (MVD) values for an adaptive motion vector resolution (AMVR) syntax element.

14. The method according to claim 8, wherein, generating a prediction block includes generating the prediction block using an affine prediction mode, and the method further includes skipping context adaptive binary arithmetic coding (CABAC) of bits representing fractional precision motion vector difference (MVD) values for an adaptive motion vector resolution (AMVR) syntax element.

15. The method according to claim 1, further comprising: Generating a prediction block for the current block using the first motion information set and the second motion information set; Performing inverse quantization and inverse transformation on the quantized transform block to generate a residual block for the current block; and Combining the samples of the residual block with the samples of the prediction block to decode the current block.

16. The method according to claim 1, further comprising: Generating a prediction block for the current block using the first motion information set and the second motion information set; Subtracting the samples of the prediction block from the samples of the current block to generate a residual block for the current block; and Performing transformation and quantization on the residual block to encode the current block.

17. An apparatus for encoding and decoding video data, the apparatus comprising: A memory configured to store video data; and One or more processors implemented in circuitry and configured to: Encode and decode a first motion information set of a current block of the video data, the current block being partitioned into a first partition and a second partition according to a non-rectangular partitioning pattern, the first motion information set referring to a reference picture list and being associated with the first partition; After encoding and decoding the first motion information set, encode and decode a second motion information set of the current block, the second motion information set referring to the reference picture list and being associated with the second partition; In response to both the first motion information set and the second motion information set referring to the reference picture list, storing the second motion information set of the current block in the memory without storing the first motion information set of the current block; and Using the stored second motion information set to predict subsequent motion information of a subsequent block of the video data adjacent to the current block.

18. The apparatus according to claim 17, wherein, The one or more processors are configured to store the second motion information set for all sub-blocks of the current block.

19. The apparatus according to claim 17, wherein, The one or more processors are configured to determine that the size of the current block meets a threshold, and store the second motion information set in response to determining that the size of the current block meets the threshold.

20. The apparatus according to claim 17, wherein, to predict the subsequent motion information of the subsequent block, the one or more processors are configured to: form a motion prediction candidate list for the subsequent block and add the second set of motion information to the motion prediction candidate list; select the second set of motion information from the motion prediction candidate list; and use the second set of motion information to predict the subsequent motion information of the subsequent block.

21. The apparatus according to claim 17, wherein, the one or more processors are further configured to generate a predicted block of the current block using the first set of motion information and the second set of motion information.

22. The apparatus according to claim 17, wherein, the one or more processors are further configured to generate a predicted block of the current block using the first set of motion information and the second set of motion information without performing a mixing operation between the first partition and the second partition.

23. The apparatus according to claim 17, wherein, the one or more processors are further configured to: generate a predicted block of the current block using the first set of motion information and the second set of motion information; perform inverse quantization and inverse transformation on the quantized transform block to generate a residual block of the current block; and combine the samples of the residual block with the samples of the predicted block to decode the current block.

24. The apparatus according to claim 17, wherein, the one or more processors are further configured to: generate a predicted block of the current block using the first set of motion information and the second set of motion information; subtract the samples of the predicted block from the samples of the current block to generate a residual block of the current block; and perform transformation and quantization on the residual block to encode the current block.

25. The apparatus according to claim 17, further comprising a display configured to display decoded video data.

26. The apparatus according to claim 17, wherein, the apparatus includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

27. The apparatus according to claim 17, wherein, the apparatus includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.

28. A computer-readable storage medium having instructions stored thereon that, when executed, cause a processor to: encode and decode a first set of motion information of a current block of video data, the current block being partitioned into a first partition and a second partition according to a non-rectangular partitioning pattern, the first set of motion information referring to a reference picture list and being associated with the first partition; after encoding and decoding the first set of motion information, encode and decode a second set of motion information of the current block, the second set of motion information referring to the reference picture list and being associated with the second partition; in response to both the first set of motion information and the second set of motion information referring to the reference picture list, store the second set of motion information of the current block without storing the first set of motion information of the current block; and Use the stored second motion information set to predict the subsequent motion information of subsequent blocks of the video data adjacent to the current block.

29. The computer-readable storage medium according to claim 28, wherein, the instructions for causing the processor to store the second motion information set include instructions for causing the processor to store the second motion information set for all sub-blocks of the current block.

30. The computer-readable storage medium according to claim 28, further comprising instructions for causing the processor to determine that the size of the current block meets a threshold, wherein, the instructions for causing the processor to store the second motion information set include instructions for causing the processor to store the second motion information set in response to determining that the size of the current block meets the threshold.

31. The computer-readable storage medium according to claim 28, wherein, the instructions for causing the processor to predict the subsequent motion information of the subsequent block include instructions for causing the processor to perform the following operations: Form a motion prediction candidate list for the subsequent block and add the second motion information set to the motion prediction candidate list; Select the second motion information set from the motion prediction candidate list; and Use the second motion information set to predict the subsequent motion information of the subsequent block.

32. The computer-readable storage medium according to claim 28, further comprising instructions for causing the processor to generate a predicted block of the current block using the first motion information set and the second motion information set.

33. The computer-readable storage medium according to claim 28, further comprising instructions for causing the processor to generate a predicted block of the current block using the first motion information set and the second motion information set without performing a mixing operation between the first partition and the second partition.

34. The computer-readable storage medium according to claim 28, further comprising instructions for causing the processor to perform the following steps: Generate a predicted block of the current block using the first motion information set and the second motion information set; Inverse quantize and inverse transform the quantized transform block to generate a residual block of the current block; and Combine the samples of the residual block with the samples of the predicted block to decode the current block.

35. The computer-readable storage medium according to claim 28, further comprising instructions for causing the processor to perform the following steps: Generate a predicted block of the current block using the first motion information set and the second motion information set; Subtract the samples of the predicted block from the samples of the current block to generate a residual block of the current block; and Transform and quantize the residual block to encode the current block.

36. An apparatus for encoding and decoding video data, the apparatus comprises: Components for encoding and decoding a first motion information set of a current block of video data, the current block being divided into a first partition and a second partition according to a non-rectangular partitioning pattern, the first motion information set referring to a reference picture list and being associated with the first partition; A component for encoding and decoding a second motion information set of the current block after encoding and decoding the first motion information set, the second motion information set referring to the reference picture list and being associated with the second segmentation; A component for storing the second motion information set of the current block without storing the first motion information set of the current block in response to both the first motion information set and the second motion information set referring to the reference picture list; And A component for predicting subsequent motion information of a subsequent block of the video data adjacent to the current block using the stored second motion information set.

37. The apparatus according to claim 36, Wherein, The component for predicting subsequent motion information of a subsequent block includes: A component for forming a motion prediction candidate list for the subsequent block, including adding the second motion information set to the motion prediction candidate list; A component for selecting the second motion information set from the motion prediction candidate list; and A component for predicting the subsequent motion information of the subsequent block using the second motion information set.

38. The apparatus according to claim 36, further Including: A component for generating a predicted block of the current block using the first motion information set and the second motion information set; A component for inverse quantizing and inverse transforming the quantized transform block to generate a residual block of the current block; And A component for combining the samples of the residual block with the samples of the predicted block to decode the current block.

39. The apparatus according to claim 36, further Including: A component for generating a predicted block of the current block using the first motion information set and the second motion information set; A component for subtracting the samples of the predicted block from the samples of the current block to generate a residual block of the current block; And A component for transforming and quantizing the residual block to encode the current block.

40. A computer program product, including a computer-readable medium having instructions stored thereon, Wherein, The instructions can be executed by one or more processors of an apparatus for encoding and decoding video data, so that the processors execute the method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Smoothing overlapped regions resulting from geometric motion partitioning

    US9020030B2