Geometric segmentation mode with coordinated motion field storage and motion compensation

By storing bidirectional predicted motion vectors for video blocks in geometric segmentation mode, the problem of high inter prediction complexity is solved, and the video encoding and decoding efficiency and robustness are improved.

CN114303371BActive Publication Date: 2025-05-16QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080061011.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-26
Filing Date
2020-08-27
Publication Date
2025-05-16
Estimated Expiration
2040-08-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are of high complexity in inter-frame prediction, especially in geometric segmentation (GEO) mode, which makes it difficult to effectively store and utilize motion vectors, affecting the encoding and decoding efficiency.

Method used

Subblock subsets are determined by dividing the current block into subblocks in geometric segmentation mode and storing a corresponding bidirectional predicted motion vector for each subblock, wherein each subblock includes at least one sample point corresponding to the predicted sample point in the final prediction block.

Benefits of technology

The operation efficiency of the video codec is improved, the video codec gain is provided, and the robustness of the de-blocking filtering and motion vector candidate list is improved by optimizing the storage and use of motion vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114303371B_ABST
    Figure CN114303371B_ABST
Patent Text Reader

Abstract

A technique for processing video data is described. The technique includes determining a first partition and a second partition of a current block encoded and decoded in a geometric partitioning mode, determining a first prediction block and a second prediction block based on a first motion vector and a second motion vector, blending the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and the second prediction block to generate a final prediction block, dividing the current block into a plurality of subblocks, determining a subset of subblocks, wherein each subblock includes at least one sample corresponding to a prediction sample in a final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block, and storing a corresponding bidirectional prediction motion vector for each subblock in the determined subset of subblocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Patent Application No. 17 / 003,733, filed on August 26, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 894,575, filed on August 30, 2019, the entire contents of each application being incorporated herein by reference. Technical Field

[0002] The present disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be integrated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), portable or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio telephones, so-called "smart phones", video teleconferencing devices, video streaming devices, etc. Digital video devices implement video codec technologies, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Codec (AVC), ITU-T H.265 / High Efficiency Video Codec (HEVC), and extensions of such standards. By implementing such video codec technologies, video devices can more efficiently send, receive, encode, decode and / or store digital video information.

[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as codec tree units (CTUs), codec units (CUs), and / or codec nodes. For video blocks in an intra-frame codec (I) slice of a picture, spatial prediction relative to reference samples in adjacent blocks in the same picture may be used for encoding. For video blocks in an inter-frame codec (P or B) slice of a picture, spatial prediction relative to reference samples in adjacent blocks in the same picture or temporal prediction relative to reference samples in other reference pictures may be used. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the invention

[0005] Generally, this disclosure describes techniques for reducing inter prediction complexity, such as by simplifying storage of geometric segmentation (GEO) patterns.Example techniques may provide technical solutions to technical problems with practical applications to improve operation of a video codec (eg, a video encoder or a video decoder).

[0006] For example, in GEO mode, a video codec (e.g., a video encoder or a video decoder) partitions a current block into a first partition and a second partition. The video codec determines a first motion vector for the first partition and a second motion vector for the second partition, and determines a first prediction block and a second prediction block based on the corresponding motion vectors. The video codec mixes the first prediction block and the second prediction block as part of generating a final prediction block for the current block.

[0007] In GEO, for the storage of motion vectors of a current block (e.g., motion vectors stored and utilized later), a video codec divides the current block into sub-blocks and stores one or more motion vectors for each sub-block. The present disclosure describes example techniques for determining which one or more motion vectors to store for each sub-block. For example, the motion vectors stored for each sub-block may affect deblocking filtering of the current block, or may affect the construction of motion vector candidate lists for encoding and decoding subsequent blocks. By utilizing the example techniques described in the present disclosure, a video codec may store motion vectors and provide video codec gains when using the stored motion vectors (e.g., in deblocking or motion vector candidate list construction).

[0008] In one example, the present disclosure describes a method for processing video data, the method comprising: determining a first partition of a current block of video data encoded and decoded in a geometric partitioning mode and a second partition of the current block of the video data; determining a first prediction block of the video data based on a first motion vector of the first partition and determining a second prediction block of the video data based on a second motion vector of the second partition; mixing the first prediction block and the second prediction block based on a weight indicating an amount by which samples in the first prediction block are scaled and an amount by which samples in the second prediction block are scaled to generate a final prediction block of the current block; dividing the current block into a plurality of sub-blocks; determining a subset of sub-blocks, wherein each sub-block includes at least one sample corresponding to a prediction sample in a final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; and storing a corresponding bidirectional prediction motion vector for each sub-block in the determined subset of sub-blocks.

[0009] In one example, the present disclosure describes a device for processing video data, the device including a memory configured to store the video data and a processing circuit coupled to the memory, the processing circuit configured to: determine a first partition of a current block of video data encoded and decoded in a geometric partitioning mode and a second partition of the current block of the video data; determine a first prediction block from the stored video data based on a first motion vector of the first partition and determine a second prediction block from the stored video data based on a second motion vector of the second partition; blend the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block of the current block; divide the current block into a plurality of sub-blocks; determine a subset of sub-blocks, wherein each sub-block includes at least one sample corresponding to a prediction sample in a final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; and store a corresponding bidirectional prediction motion vector for each sub-block in the determined subset of sub-blocks.

[0010] In one example, the present disclosure describes a computer-readable storage medium having instructions stored thereon, which when executed cause one or more processors of a device for processing video data to: determine a first partition of a current block of video data encoded and decoded in a geometric partitioning mode and a second partition of the current block of the video data; determine a first prediction block of the video data based on a first motion vector of the first partition and determine a second prediction block of the video data based on a second motion vector of the second partition; blend the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block of the current block; divide the current block into a plurality of sub-blocks; determine a subset of sub-blocks, wherein each sub-block includes at least one sample corresponding to a prediction sample in a final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; and store a corresponding bidirectional prediction motion vector for each sub-block in the determined subset of sub-blocks.

[0011] In one example, the present disclosure describes an apparatus for processing video data, the apparatus comprising: a component for determining a first partition of a current block of video data encoded and decoded in a geometric partitioning mode and a second partition of the current block of the video data; a component for determining a first prediction block of the video data based on a first motion vector of the first partition and a second prediction block of the video data based on a second motion vector of the second partition; a component for mixing the first prediction block and the second prediction block to generate a final prediction block of the current block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block; a component for dividing the current block into a plurality of sub-blocks; a component for determining a subset of sub-blocks, wherein each sub-block includes at least one sample corresponding to a prediction sample in a final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; and a component for storing a corresponding bidirectional prediction motion vector for each sub-block in the determined subset of sub-blocks.

[0012] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may perform the techniques of this disclosure.

[0014] Figure 2A and 2B is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and a corresponding codec tree unit (CTU).

[0015] Figure 3 is a block diagram illustrating an example video encoder that may perform the techniques of this disclosure.

[0016] Figure 4 is a block diagram illustrating an example video decoder that may perform the techniques of this disclosure.

[0017] Figure 5A and 5B are conceptual diagrams respectively illustrating examples of diagonal partitioning and anti-diagonal partitioning triangulation based on inter prediction.

[0018] Figure 6 is a conceptual diagram illustrating example spatial and temporal neighboring blocks for constructing a candidate list.

[0019] Figure 7 is a table showing the motion vector prediction selections for triangular partitioning mode.

[0020] Fig. 8A and8B 2 are conceptual diagrams showing weights for mixing processes of luminance components and chrominance components, respectively.

[0021] Fig. 9 is a conceptual diagram showing an example of a triangular partition mode (TPM) as applied to a geometric partition mode (GEO).

[0022] Fig.10 is a conceptual diagram illustrating an example of GEO segmentation signaling.

[0023] Fig.11 is a conceptual diagram illustrating an example of partial transformation of an inter-prediction block for GEO prediction.

[0024] Fig.12 is a conceptual diagram illustrating an example of a partially transformed deblocking-like process.

[0025] Fig.13 is a conceptual diagram showing examples of GEO angles including 0° and 90° angles on a TPM angle machine.

[0026] Fig.14 is a conceptual diagram illustrating an example of a GEO mode with a bounding box shown in dotted lines.

[0027] Fig.15A and Fig. 15B is a conceptual diagram showing an example of weights applied to a TPM.

[0028] Fig.16 is a conceptual diagram illustrating an example of a GEO edge in a bounding box of a virtual bounding box having a width and a height that are powers of 2.

[0029] Fig.17A and Fig. 17B is a conceptual diagram illustrating a current coding unit (CU) in a hypothetical CU, where the current CU forms its mask by sampling part of the weight values ​​of the mask of the hypothetical CU.

[0030] Figures 18A-18D : is a conceptual diagram showing the support angle of the current CU in the virtual CU.

[0031] Fig.19A and Fig.19B is a conceptual diagram showing a support angle when a CU is split from one angle at a starting point.

[0032] Fig. 20 is a conceptual diagram illustrating an example of sampling weight values ​​in a hypothetical CU with an offset.

[0033] Fig.21 is a flow chart illustrating an example method for processing a current block. DETAILED DESCRIPTION

[0034] In video coding and decoding, the video encoder encodes the current block using inter-frame prediction or intra-frame prediction. In inter-frame prediction and intra-frame prediction, the video encoder generates a prediction block for the current block, determines the difference between the current block and the prediction block (e.g., a residual block), and signals information indicating the residual block. The video encoder may also signal a prediction mode that indicates the manner in which the prediction block is generated. The video decoder receives the prediction mode information and the residual block information, generates a prediction block based on the signaled prediction mode information, and adds the prediction block to the residual block to reconstruct the current block. In inter-frame prediction, the prediction block is generated from samples identified by a motion vector and may be samples in a picture different from the picture including the current block. In intra-frame prediction, the prediction block is generated from samples in the same picture as the current block (such as samples adjacent to the current block).

[0035] One example of inter prediction is geometric partitioning (GEO) mode. In GEO mode, the video encoder partitions the current block into two partitions. The partition line that partitions the current block into two partitions can be a diagonal line, and although possible, does not need to start or end at a corner of the current block (for example, although possible, the angle of the partition does not need to be 45 degrees). The video encoder can determine a first motion vector for the first partition and a second motion vector for the second partition. In some examples, the video encoder can identify a first prediction block based on the first motion vector, and identify a second prediction block based on the second motion vector.

[0036] The video encoder 200 may then use weighting to blend the first prediction block and the second prediction block to generate a final prediction block. For example, for samples close to the dividing line, the video encoder may scale the samples from the first prediction block and the second prediction block equally and add the resulting values ​​to generate the samples in the final prediction block. However, for samples close to the boundary of the final prediction block, the video encoder may only use samples from the first prediction block or the second prediction block to generate the samples in the final prediction block. For samples between the dividing line and the boundary, the video encoder may weight the samples of one of the first prediction block or the second prediction block greater than the weight of the samples of the other of the first prediction block or the second prediction block for blending.

[0037] In GEO mode, the video encoder may determine a residual between a current block and a final predicted block, and signal information indicating the residual. The video encoder may also signal information indicating that the current block is encoded in GEO mode, information indicating a partition line, and information for determining a first motion vector and a second motion vector and reference pictures to which the first motion vector and the second motion vector point.

[0038] The video decoder may segment the current block based on the received information indicating the segmentation line, determine two prediction blocks based on the received information to determine the first motion vector and the second motion vector, and generate a final prediction block (e.g., using weighted mixing). The video decoder may then add the final prediction block to the residual to reconstruct the current block.

[0039] Then, the video encoder and the video decoder can store the motion vector information of the current block. The stored motion vector information of the current block can be used to perform deblocking filtering on the current block (for example, to determine the boundary strength value) or to construct a motion vector candidate list for encoding or decoding a subsequent block.

[0040] The present disclosure describes example techniques for determining motion vector information for a current block to be stored. The stored motion vector information may affect the quality of deblocking filtering or the robustness of motion vector information in a motion vector candidate list used to encode or decode subsequent blocks. Quality deblocking filtering may refer to the reduction in visual artifacts along the boundaries of blocks from deblocking filtering. The robustness of motion vector information refers to candidate motion vectors that tend to be similar to motion vectors of subsequent blocks. By utilizing the example techniques described in the present disclosure, a video encoder and a video decoder may store motion vector information for a current block in a motion vector candidate list that potentially provides higher quality deblocking filtering and more robust motion vector information.

[0041] In one or more examples, the video encoder and the video decoder may divide the current block into a plurality of sub-blocks (e.g., 4×4 sub-blocks). The video encoder and the video decoder may store motion vector information for each sub-block. According to one or more examples described in the present disclosure, the video encoder and the video decoder may determine whether any sample in the sub-block has a corresponding sample in a final prediction block generated by equally weighting samples in a first prediction block and samples in a second prediction block.

[0042] If the sub-block includes samples having corresponding samples in a final prediction block generated by equally weighting samples in a first prediction block and samples in a second prediction block, the video encoder and video decoder may determine a bidirectionally predicted motion vector. It should be understood that, although possible, "bidirectionally predicted motion vector" does not necessarily mean that there are two motion vectors. Instead, there may be example operations that the video encoder and video decoder perform to determine the bidirectionally predicted motion vector. For example, if the first motion vector or the second motion vector used to identify the first prediction block and the second prediction block references reference pictures in different reference picture lists, the video encoder and video decoder may store the first motion vector and the second motion vector for the sub-block as the bidirectionally predicted motion vector for the sub-block. If the first motion vector and the second motion vector reference a reference picture in the same reference picture, the video encoder and video decoder may select one of the first motion vector or the second motion vector as the bidirectionally predicted motion vector.

[0043] If the sub-block does not include samples that have corresponding samples in the final prediction block generated by equally weighting the samples in the first prediction block and the samples in the second prediction block, the video encoder and the video decoder may store the first motion vector or the second motion vector as the motion vector of the sub-block. Whether the video encoder and the video decoder store the first motion vector or the second motion vector may be based on the position of the sub-block in the current block. For example, if a majority of the sub-block resides in the first partition, the video encoder and the video decoder may store the first motion vector for the sub-block. If a majority of the sub-block resides in the second partition, the video encoder and the video decoder may store the second motion vector for the sub-block.

[0044] Figure 1 1 is a block diagram illustrating an example video encoding and decoding system 100 that may perform the techniques of the present disclosure. The techniques of the present disclosure are generally directed to encoding and decoding (encoding and / or decoding) video data. Generally, video data includes any data used to process video. Thus, video data may include original unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0045] like Figure 1As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a target device 116. Specifically, source device 102 provides the video data to target device 116 via computer-readable medium 110. Source device 102 and target device 116 may include any of a variety of devices, including desktop computers, notebooks (i.e., laptop computers), tablet computers, set-top boxes, handheld phones (such as smart phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and target device 116 may be equipped for wireless communication and, therefore, may be referred to as wireless communication devices.

[0046] exist Figure 1 In the example of, source device 102 includes video source 104, memory 106, video encoder 200 and output interface 108. Target device 116 includes input interface 122, video decoder 300, memory 120 and display device 118. According to the present disclosure, the video encoder 200 of source device 102 and the video decoder 300 of target device 116 can be configured to apply the technology of motion field storage and motion weight derivation for unified triangular prediction mode (TPM) and geometric segmentation mode (GEO). Thus, source device 102 represents an example of video encoding device, and target device 116 represents an example of video decoding device. In other examples, source device and target device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, target device 116 may be connected to an external display device through an interface without including an integrated display device.

[0047] like Figure 1 The system 100 shown is only an example. Generally, any digital video encoding and / or decoding device can perform the technology of motion field storage and motion weight derivation for unified triangular prediction mode (TPM) and geometric segmentation mode (GEO). The source device 102 and the target device 116 are only examples of such codec devices, wherein the source device 102 generates codec video data for transmission to the target device 116. The present disclosure represents the "codec" device as a device that performs data coding (coding and / or decoding). Thus, the video encoder 200 and the video decoder 300 represent examples of codec devices, specifically, video encoders and video decoders, respectively. In some examples, the source device 102 and the target device 116 can operate in a substantially symmetrical manner, so that each of the source device 102 and the target device 116 includes a video encoding and decoding component. Thus, the system 100 can support one-way or two-way video transmission between the source device 102 and the target device 116, such as for video streaming, video playback, video broadcasting or video telephony.

[0048] Generally, the video source 104 represents a source of video data (i.e., original, unencoded video data), and provides a sequence of continuous pictures (also referred to as "frames") of the video data to the video encoder 200, which encodes the data of the pictures. The video source 104 of the source device 102 may include a video capture device, such as a camera, a video archive including previously captured original video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 104 may generate computer graphics-based data as a source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may rearrange the pictures from the receiving order (sometimes referred to as "display order") to the encoding and decoding order for encoding and decoding. The video encoder 200 may generate a bitstream including the encoded video data. Then, the source device 102 may output the encoded video data to a computer-readable medium 110 via an output interface 108, for example, for reception and / or retrieval by an input interface 122 of a target device 116.

[0049] The memory 106 of the source device 102 and the memory 120 of the target device 116 represent general purpose memory. In some examples, the memory 106, 120 can store original video data, such as the original video from the video source 104 and the original decoded video data from the video decoder 300. Additionally or alternatively, the memory 106, 120 can store software instructions that can be executed by, for example, the video encoder 200 and the video decoder 300, respectively. Although the memory 106 and the memory 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 can also include internal memory that achieves functionally similar or equivalent purposes. Further, the memory 106, 120 can store, for example, encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, some portions of the memory 106, 120 can be allocated as one or more video buffers, for example, to store original decoded and / or encoded video data.

[0050] The computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded video data from the source device 102 to the target device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to transmit the encoded video data directly to the target device 116 in real time, for example, via a radio network or a computer-based network. According to a communication standard such as a wireless communication protocol, the output interface 108 may modulate a transmission signal including the encoded video data, and the input interface 122 may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other equipment that facilitates communication from the source device 102 to the target device 116.

[0051] In some examples, source device 102 may output the encoded video data to file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. Target device 116 may access the stored video data from file server 114 via streaming or downloading.

[0052] The file server 114 may be any type of server device capable of storing encoded video data and sending the encoded video data to the target device 116. The file server 114 may represent a web server (e.g., for a website), a server configured to provide a file transfer protocol service (e.g., a file transfer protocol (FTP) or a file delivery over unidirectional transport (FLUTE) protocol), a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a multimedia broadcast multicast service (MBMS) or an enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as dynamic adaptive streaming over HTTP (DASH), HTTP real-time streaming (HLS), real-time streaming protocol (RTSP), HTTP dynamic streaming, etc.

[0053] Target device 116 may access the encoded video data from file server 114 through any standard data connection including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on file server 114. Input interface 122 may be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.

[0054] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data, such as encoded video data, according to a cellular communication standard such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, or the like. In certain examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data, such as encoded video data, according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee 5G). TM ),Bluetooth TM Standards, etc. to transmit data such as encoded video data. In some examples, source device 102 and / or target device 116 may include respective system-on-chip (SoC) devices. For example, source device 102 may include a SoC device to perform functions attributed to video encoder 200 and / or output interface 108, and target device 116 may include a SoC device to perform functions attributed to video decoder 300 and / or input interface 122.

[0055] The technology of the present disclosure can be applied to video encoding and decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded to a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0056] The input interface 122 of the target device 116 receives an encoded video bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements having values ​​describing characteristics and / or processing of video blocks or other codec units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays decoded pictures of the decoded video data to a user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0057] Although not in Figure 1 , but in some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or an audio decoder and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams including audio and video in a common data stream. If applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).

[0058] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store instructions for the software in a suitable non-temporary computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in each device. The device including the video encoder 200 and / or the video decoder 300 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.

[0059] The video encoder 200 and the video decoder 300 may operate according to a video codec standard such as ITU-T H.265, also known as High Efficiency Video Codec (HEVC), or an extension thereof such as a multi-view and / or scalable video codec extension. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Versatile Video Codec (VVC)). A recent draft of the VVC standard is described in: Bross et al., “Versatile Video Coding (Draft 6)”, ITU-T SG 16WP3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 15th Meeting: Gothenburg, Sweden, July 3-12, 2019, JVET-O2001-vE (hereinafter referred to as “VVC Draft 6”). A more recent draft of the VVC standard is described in Bross et al., "Versatile Video Coding (Draft 10)", ITU-T SG 16 WP3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 18th Meeting: Teleconference, June 22-July 1, 2020, JVET-S2001-vA (hereinafter referred to as "VVC Draft 10"). However, the techniques of this disclosure are not limited to any particular codec standard.

[0060] Generally, the video encoder 200 and the video decoder 300 can perform block-based encoding and decoding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., to be encoded, to be decoded, or otherwise used in the encoding and / or decoding process). For example, a block may include a two-dimensional sample matrix of luminance and / or chrominance data. Generally, the video encoder 200 and the video decoder 300 can encode and decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, the video encoder 200 and the video decoder 300 can encode and decode luminance and chrominance components, where the chrominance components may include both red hue and blue hue chrominance components, rather than encoding and decoding red, green, and blue (RGB) data of the sample points of the picture. In some examples, the video encoder 200 converts the received RGB format data into a YUV representation before encoding, and the video decoder 300 converts the YUV representation into an RGB format. Alternatively, a pre-processing and post-processing unit (not shown) can perform these conversions.

[0061] The present disclosure generally relates to the encoding and decoding of pictures (e.g., encoding and decoding) to include the process of encoding or decoding picture data. Similarly, the present disclosure may relate to the encoding and decoding of blocks of pictures to include the process of encoding or decoding the data of the blocks, such as prediction and / or residual encoding and decoding. The encoded video bitstream generally includes a series of values ​​of syntax elements used to represent codec decisions (e.g., codec mode) and the partitioning of pictures into blocks. Thus, references to encoding and decoding pictures or blocks should generally be understood as encoding and decoding the values ​​of syntax elements that form pictures or blocks.

[0062] HEVC defines various blocks, including codec units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video codec (such as video encoder 200) partitions a codec tree unit (CTU) into CUs according to a quadtree structure. That is, the video codec partitions the CTU and CU into four equal non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without a child node may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video codec may further partition the PU and TU. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of a TU. In HEVC, a PU represents inter-frame prediction data, and a TU represents residual data. An intra-predicted CU includes intra-frame prediction information, such as an intra-frame mode indication.

[0063] As another example, the video encoder 200 and the video decoder 300 can be configured to operate according to VVC. According to VVC, a video codec (such as the video encoder 200) partitions a picture into multiple codec tree units (CTUs). The video encoder 200 can partition the CTU according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the separation between CU, PU, ​​and TU of HEVC. The QTBT structure includes two levels: a first level partitioned according to quadtree partitioning, and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to codec units (CUs).

[0064] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also known as ternary tree (TT)) partitioning. Triple tree or ternary tree partitioning is a partitioning method that divides a block into three sub-blocks. In some examples, the ternary tree or ternary tree partitioning divides the block into three sub-blocks without dividing the original block by the center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0065] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the corresponding chroma components).

[0066] The video encoder 200 and the video decoder 300 may be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures in accordance with HEVC. For illustrative purposes, a description of the technology of the present disclosure is presented with respect to QTBT segmentation. However, it should be understood that the technology of the present disclosure may also be applied to video codecs configured to use quadtree segmentation or other types of segmentation.

[0067] Blocks (e.g., CTUs or CUs) may be grouped in a picture in various ways. As an example, a brick may refer to a rectangular area of ​​a CTU row in a particular tile in a picture. A tile may be a rectangular area of ​​a CTU in a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular area of ​​a CTU having a height equal to the picture height and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular area of ​​a CTU having a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to the picture width.

[0068] In some examples, a tile may be partitioned into multiple bricks, each of which may include one or more CTU rows in the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick that is a proper subset of a tile may not be referred to as a tile.

[0069] Tiles in a picture can also be arranged in slices. A slice can be an integer number of tiles of a picture that can be contained exclusively in a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes multiple complete tiles or only a contiguous sequence of complete tiles of a tile.

[0070] This disclosure may use "N×N" and "N by N" interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU will have 16 samples in the vertical direction (y=16) and 16 samples in the horizontal direction (x=16). Similarly, an N×N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU may be arranged in rows and columns. In addition, a CU does not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may contain N×M samples, where M is not necessarily equal to N.

[0071] The video encoder 200 encodes the video data of the CU representing prediction and / or residual information and other information. The prediction information indicates how the CU will be predicted in order to form a prediction block for the CU. The residual information generally represents the sample-by-sample difference between the samples of the CU before encoding and the prediction block.

[0072] In order to predict a CU, the video encoder 200 can generally form a prediction block of the CU by inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting a CU from data of a previously coded picture, while intra-frame prediction generally refers to predicting a CU from previously coded data of the same picture. In order to perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate a prediction block. The video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, for example, in terms of the difference between the CU and the reference block. The video encoder 200 can use the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean square difference (MSD), or other such difference calculations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional prediction or bidirectional prediction to predict the current CU.

[0073] VVC can also provide an affine motion compensation mode, which can be considered an inter-prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion, such as enlargement or reduction, rotation, perspective motion, or other irregular motion types.

[0074] In order to perform intra prediction, the video encoder 200 can select an intra prediction mode to generate a prediction block. VVC can provide sixty-seven intra prediction modes, including modes of various directions as well as a plane mode and a DC mode. Generally, the video encoder 200 selects an intra prediction mode that describes the neighboring samples of the current block (e.g., a block of a CU) from which the predicted samples of the current block are predicted. Assuming that the video encoder 200 encodes and decodes the CTU and CU in a raster scan order (from left to right, from top to bottom), such samples can typically be above, to the upper left, or to the left of the current block in the same picture as the current block.

[0075] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter-frame prediction mode, the video encoder 200 can encode data indicating which of various available inter-frame prediction modes is used and motion information of the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 can encode the motion vector using advanced motion vector prediction (AMVP) or merge mode. The video encoder 200 can use a similar mode to encode the motion vector of the affine motion compensation mode.

[0076] After prediction (such as intra-frame prediction or inter-frame prediction of a block), the video encoder 200 can calculate residual data for the block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and a prediction block for the block, which is formed using a corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to generate transform data in a transform domain rather than a sample domain. For example, the video encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. In addition, the video encoder 200 can apply a secondary transform after a primary transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.

[0077] As described above, after performing any transforms to produce transform coefficients, the video encoder 200 can perform quantization on the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round down an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bitwise right shift of the value to be quantized.

[0078] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place the transform coefficients with higher energy (and therefore lower frequency) in front of the vector and the transform coefficients with lower energy (and therefore higher frequency) in the back of the vector. In some examples, the video encoder 200 can scan the quantized transform coefficients using a predefined scan order to generate a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform adaptive scanning. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy encode the one-dimensional vector, for example, according to context adaptive binary arithmetic coding (CABAC). The video encoder 200 can also entropy encode the values ​​of syntax elements, which describe metadata associated with the encoded video data used by the video decoder 300 in decoding the video data.

[0079] To perform CABAC, the video encoder 200 may assign context within a context model to a symbol to be transmitted. For example, the context may relate to whether the neighboring values ​​of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.

[0080] The video encoder 200 may further generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, to the video decoder 300, for example, in a picture header, a block header, a slice header, or generate other syntax data, such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS). The video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.

[0081] In this way, the video encoder 200 can generate a bitstream including the encoded video data, such as syntax elements describing the partitioning of a picture into blocks (eg, CUs) and prediction and / or residual information of the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0082] In general, the video decoder 300 performs the reverse process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may use CABAC to decode the values ​​of the syntax elements of the bitstream in a manner substantially similar to (although reversed from) the CABAC encoding process of the video encoder 200. The syntax elements may define partitioning information for partitioning a picture into CTUs and partitioning each CTU according to a corresponding partitioning structure such as a QTBT structure to define CUs of the CTU. The syntax elements may further define prediction and residual information for a block (e.g., a CU) of video data.

[0083] For example, the residual information may be represented by quantized transform coefficients. The video decoder 300 may inverse quantize and inverse transform the quantized transform coefficients of the block to reproduce a residual block for the block. The video decoder 300 uses the signaled prediction mode (intra-frame or inter-frame prediction) and related prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for the block. The video decoder 300 may then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The video decoder 300 may perform additional processing (such as performing a deblocking process) to reduce visual artifacts along block boundaries.

[0084] According to the technology of the present disclosure, a video codec (e.g., video encoder 200 or video decoder 300) may be configured to perform example techniques. For example, video encoder 200 and video decoder 300 may be configured to perform inter-frame prediction on a current block in GEO mode (also referred to as geometric partitioning mode (GPM)). In GEO mode, video encoder 200 and video decoder 300 may partition the current block into a first partition and a second partition (e.g., based on a partition line that divides the current block into the first partition and the second partition). For each partition, video encoder 200 and video decoder 300 may determine a motion vector (i.e., a first motion vector for the first partition and a second motion vector for the second partition). Video encoder 200 and video decoder 300 may identify a first prediction block based on the first motion vector, and identify a second prediction block based on the second motion vector.

[0085] The video encoder 200 and the video decoder 300 may mix the first prediction block and the second prediction block to generate a final prediction block. As an example of mixing, the video encoder 200 and the video decoder 300 may scale the samples in the first prediction block by a first weight, and scale the co-located samples in the second prediction block by a second weight, and add the resulting values ​​together. The samples in the first prediction block and the samples in the second prediction may be co-located with two samples in the same position in the corresponding first prediction block and the second prediction block.

[0086] The weight applied to the blend may indicate the amount that a sample from the first partition or the second partition contributes to the final prediction block. For example, as described in more detail below, the weight of a sample in the first prediction block may be 1 / 8, and the weight of a co-located sample in the second prediction block may be 7 / 8. In this example, the sample in the second prediction block contributes more to the value of the sample in the final prediction block than the sample in the first prediction block. As another example, the weight of a sample in the first prediction block and the weight of the co-located sample in the second prediction block may be 4 / 8. In this example, the sample in the first prediction block and the co-located sample in the second prediction block contribute the same to the value of the sample in the final prediction block.

[0087] Through this mixing, the video encoder 200 and the video decoder 300 can generate a final prediction block. There may be a one-to-one correspondence between the samples in the final prediction block and the current block. For example, if the current block is 8×8, then the prediction block may also be 8×8. The video encoder 200 and the video decoder 300 can encode or decode the current block using inter-frame prediction based on the prediction block. For example, the video encoder 200 can determine the residual between the current block and the prediction block and signal the residual. The video decoder 300 can receive the residual and add the residual to the prediction block to reconstruct the current block.

[0088] According to one or more examples described in the present disclosure, the video encoder 200 and the video decoder 300 may be configured to store motion vector information of the current block. To store the motion vector information, the video encoder 200 and the video decoder 300 may divide the current block into sub-blocks (eg, 4×4 sub-blocks).

[0089] Each sample in each sub-block may correspond to a sample in a prediction block. As an example, if a sample in a sub-block of a current block is co-located with a sample in a prediction block, the two samples may correspond. The video encoder 200 and the video decoder 300 may determine whether the samples in the sub-block correspond to predicted samples in a final prediction block generated based on equal weighting of samples in a first prediction block and samples in a second prediction block. For example, if the predicted samples in the final prediction block are generated by scaling the samples in the first prediction block by 4 / 8 and scaling the samples in the second prediction block by 4 / 8, the predicted samples in the final prediction block may be considered to be generated based on equal weighting of the samples in the first prediction block and the samples in the second prediction block.

[0090] If the subblock includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in the first prediction block and samples in the second prediction block, the video encoder 200 and the video decoder 300 may store a bidirectional prediction motion vector for the subblock. If the subblock does not include any sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in the first prediction block and samples in the second prediction block, the video encoder 200 and the video decoder 300 may store a unidirectional prediction motion vector.

[0091] The term "bidirectional prediction motion vector" should not be considered limited to requiring the presence of two motion vectors. Instead, the video encoder 200 and the video decoder 300 may perform one or more operations to determine the bidirectional prediction motion vector. For example, the first prediction block may be in a first reference picture, and the second prediction block may be in a second reference picture. The first reference picture and the second reference picture may be in different reference picture lists or in the same reference picture list. If the first prediction block and the second prediction block are in different reference picture lists (i.e., the first motion vector and the second motion vector are from different reference picture lists), the bidirectional prediction motion vector is the first motion vector and the second motion vector. If the first prediction block and the second prediction block are in the same reference picture list (i.e., the first motion vector and the second motion vector are from the same reference picture list), the bidirectional prediction motion vector is one of the first motion vector or the second motion vector.

[0092] The term "unidirectional motion vector" means that there is only one motion vector and may be based on the position of the subblock in the current block. For example, if the video encoder 200 and the video decoder 300 are to store a unidirectional motion vector for a subblock, the video encoder 200 and the video decoder 300 may store a first motion vector if most of the samples of the subblock are in the first partition, and store a second motion vector if most of the samples of the subblock are in the second partition. If the subblock has equal samples in the first partition and the second partition, the subblock includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in the first prediction block and samples in the second prediction block, and thus the video encoder 200 and the video decoder 300 may store a bidirectional prediction motion vector.

[0093] For example, the video codec may be configured to determine one or more angles from a set of angles for segmenting a current block using a geometric segmentation mode (GEO), wherein the set of angles from which the one or more angles are determined is the same as a set of angles that can be used for a triangular segmentation mode (TPM); segment the current block based on the determined one or more angles; and encode and decode the current block based on the segmentation of the current block. As another example, the video codec may be configured to determine one or more weights from a set of weights for blending the current block using GEO, wherein the set of weights from which the one or more weights are determined is the same as a set of weights that can be used for TPM; blend a previous block based on the determined one or more weights; and encode and decode the current block based on the blending of the current block. As another example, the video codec may be configured to determine a motion field storage for using GEO, wherein the motion field storage is the same as the motion field storage that can be used for TPM; and encode and decode the current block based on the determined motion field storage.

[0094] In general, the present disclosure may involve "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, the video encoder 200 may signal the values ​​of syntax elements in a bitstream. In general, signaling refers to generating values ​​in a bitstream. As described above, the source device 102 may transmit the bitstream to the target device 116 in substantially real time (or non-real time, such as may occur when storing syntax elements to the storage device 112 for later retrieval by the target device 116).

[0095] Figure 2A and 2B is a conceptual diagram showing an example quadtree binary tree (QTBT) structure 130 and a corresponding codec tree unit (CTU) 132. Solid lines represent quadtree partitions and dashed lines indicate binary tree partitions. In each partition (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used, where in this example, 0 indicates horizontal partitioning and 1 indicates vertical partitioning. For quadtree partitioning, since the quadtree node divides the block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the partition type. Accordingly, the video encoder 200 can encode syntax elements (e.g., partition information) of the region tree level (i.e., solid line) of the QTBT structure 130 and syntax elements (e.g., partition information) of the prediction tree level (i.e., dashed line) of the QTBT structure 130, and the video decoder 300 can decode the above. For a CU represented by a terminal leaf node of the QTBT structure 130 , the video encoder 200 may encode video data (such as prediction and transform data), and the video decoder 300 may decode the same.

[0096] Generally, Figure 2B The CTU 132 of the first and second levels may be associated with parameters that define the size of blocks corresponding to nodes of the first and second level QTBT structures 130. These parameters may include a CTU size (representing the size of the CTU 132 in a sample), a minimum quadtree size (MinQTSize, representing the minimum allowed quadtree leaf node size), a maximum binary tree size (MaxBTSize, representing the maximum allowed binary tree root node size), a maximum binary tree depth (MaxBTDepth, representing the maximum allowed binary tree depth), and a minimum binary tree size (MinBTSize, representing the minimum allowed binary tree leaf node size).

[0097] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each child node can be segmented according to the quadtree segmentation. That is, the node of the first level is a leaf node (no child node) or has four child nodes. The example of the QTBT structure 130 represents such a node, which includes a child node and a parent node with a solid line branch. If the node of the first level is not larger than the maximum allowed binary tree root node size (MaxBTSize), the node can be further segmented by the respective binary trees. The binary tree partitioning of a node can be iterated until the node generated by the partition reaches the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). The example of the QTBT structure 130 represents such a node as having a dotted branch. The binary tree leaf node is represented as a coding unit (CU), which is used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any further segmentation. As described above, a CU can also be represented as a "video block" or "block".

[0098] In one example of a QTBT partitioning structure, the CTU size is set to 128×128 (luminance sample and two corresponding 64×64 chroma samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quadtree node is 128×128, the leaf quadtree node will not be further divided by the binary tree because its size exceeds MaxBTSize (64×64 in this example). Otherwise, the leaf quadtree node will be further divided by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further division is not allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), the binary tree node means that no further horizontal partitioning is allowed. Similarly, a binary tree node with a height equal to MinBTSize indicates that no further vertical partitioning is allowed for the binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further segmentation.

[0099] Figure 3 is a block diagram illustrating an example video encoder 200 that may perform the techniques of this disclosure. Figure 3 For purposes of explanation and should not be considered a constraint on the techniques extensively exemplified and described in this disclosure. For purposes of illustration, this disclosure describes the video encoder 200 in the context of video codec standards such as the ITU-U H.265 (HEVC) video codec standard and the developing ITU-T H.266 (VVC) video codec standard. However, the techniques of this disclosure are not limited to these video codec standards and are generally applicable to video encoding and decoding.

[0100] exist Figure 3In the example of , the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or in processing circuits. In addition, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.

[0101] The video data memory 230 may store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive video data from, for example, the video source 104 ( Figure 1 ) receives video data stored in video data memory 230. DPB 218 can be used as a reference picture memory, which stores reference video data for use by video encoder 200 when predicting subsequent video data. Video data memory 230 and DPB 218 can be formed by any of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM) or other types of storage devices. Video data memory 230 and DPB 218 can be provided by the same storage device or a separate storage device. In various examples, video data memory 230 can be placed on-chip with other components of video encoder 200, as shown, or placed off-chip relative to those components.

[0102] In the present disclosure, references to the video data memory 230 should not be interpreted as limited to a memory inside the video encoder 200 (unless specifically stated otherwise) or a memory outside the video encoder 200 (unless specifically stated otherwise). Of course, references to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data of a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage for outputs from the various units of the video encoder 200 .

[0103] Shown Figure 3The various units are used to help understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are preset on the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (for example, to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more units may be different circuit blocks (fixed function or programmable), and in some examples, one or more units may be integrated circuits.

[0104] The video encoder 200 may include an arithmetic logic unit (ALU), an elementary function unit (EFU), a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video encoder 200 is performed using software executed by a programmable circuit, the memory 106 ( Figure 1 ) may store instructions (eg, object code) for software received and executed by the video encoder 200, or another memory within the video encoder 200 (not shown) may store such instructions.

[0105] The video data memory 230 is configured to store received video data. The video encoder 200 may retrieve a picture of video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be original video data to be encoded.

[0106] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0107] Generally, the mode selection unit 202 coordinates multiple encoding passes to test combinations of encoding parameters and derive rate-distortion values ​​of such combinations. The encoding parameters may include the partitioning of CTUs into CUs, the prediction mode of the CUs, the transform type of the residual data of the CUs, the quantization parameter of the residual data of the CUs, etc. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0108] The video encoder 200 may partition a picture retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs in a slice. The mode selection unit 202 may partition the CTUs of the picture according to a tree structure (such as the QTBT structure or quadtree structure of HEVC described above). As described above, the video encoder 200 may form one or more CUs by partitioning the CTUs according to the tree structure. Such CUs may also be generally referred to as "video blocks" or "blocks".

[0109] In general, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for the current block (e.g., the overlapping portion of the current CU or PU and TU in HEVC). In order to inter-predict the current block, the motion estimation unit 222 can perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 can calculate a value indicating how similar the potential reference block is to the current block based on, for example, the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 can generally perform these calculations using the sample-by-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 can identify the reference block with the lowest value generated by these calculations, which indicates the reference block that most closely matches the current block.

[0110] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to a current block in a current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, for unidirectional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bidirectional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve data for the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate the prediction block according to one or more interpolation filters. In addition, for bidirectional inter prediction, the motion compensation unit 224 may retrieve data for two reference blocks identified by respective motion vectors and combine the retrieved data (e.g., by sample-by-sample averaging or weighted averaging).

[0111] As another example, the mode selection unit 202 may configure the motion estimation unit 222 and the motion compensation unit 224 to perform techniques related to the GEO mode. For example, the motion estimation unit 222 may divide the current block into two partitions (e.g., divide the current block into two partitions along a partition line). The motion estimation unit 222 may determine a first motion vector for the first partition and a second motion vector for the second partition. The motion compensation unit 224 may determine a first prediction block based on the first motion vector, and determine a second prediction block based on the second motion vector.

[0112] The motion compensation unit 224 may blend the first prediction block and the second prediction block based on a weight indicating the amount of scaled samples in the first prediction block and the amount of scaled samples in the second prediction block to generate a final prediction block for the current block. For example, the motion compensation unit 224 may determine a first scaling factor (e.g., 0, 1 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 7 / 8, or 8 / 8) for samples in the first prediction block and a second scaling factor for co-located samples in the second prediction block. In some examples, the first scaling factor and the second scaling factor add up to one.

[0113] The scaling factor may be a weight indicating an amount by which the samples in the first prediction block are scaled and an amount by which the samples in the second prediction block are scaled, so as to generate a final prediction block of the current block by summing the scaled samples of the first prediction block and the second prediction block. The scaled samples may be luma values ​​of a luma block and chroma values ​​of a chroma block. For example, in order to generate the predicted samples in the final prediction block, the motion compensation unit 224 may multiply the sample values ​​of the samples in the first prediction block by 3 / 8 and multiply the sample values ​​of the co-located samples in the second prediction block by 5 / 8 to generate the predicted sample values ​​of the co-located samples in the final prediction block. In this example, 3 / 8 and 5 / 8 are examples of the amount by which the samples in the first prediction block are scaled and the amount by which the samples in the second prediction block are scaled, respectively. The scaled results are then added together to generate the predicted samples in the final prediction block.

[0114] In one or more examples, the motion estimation unit 222 may be configured to store motion vector information of the current block (e.g., storing the motion vector information in the decoded picture buffer 218, the video data memory 230, or the memory 106, as a few examples). To determine the motion vector information of the current block, the motion estimation unit 222 may divide the current block into a plurality of sub-blocks (e.g., each sub-block is 4×4 in size). The motion estimation unit 222 may determine a set of sub-blocks, each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. For example, the motion estimation unit 222 may determine which sub-blocks include at least one sample corresponding to a predicted sample in a final prediction block generated by scaling the samples in the first prediction block by 4 / 8 and scaling the samples in the second prediction block by 4 / 8 and summing the scaled samples. In this example, since the 4 / 8 scaling of the samples in the first prediction block and the second prediction block is the same, the samples in the first prediction block and the samples in the second prediction block have equal weights.

[0115] The motion estimation unit 222 may store a corresponding bidirectional prediction motion vector for each sub-block in the determined sub-block set. For example, the motion estimation unit 222 may determine whether the first motion vector and the second motion vector are from different reference picture lists (e.g., reference pictures in different reference picture lists). The motion estimation unit 222 may be configured to do one of the following: (1) store both the first motion vector and the second motion vector for each sub-block in the sub-block set based on the first motion vector and the second motion vector being from different reference picture lists, or (2) select one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and store the selected one of the first motion vector or the second motion vector for each sub-block in the sub-block set.

[0116] The set of subblocks may be considered as a first set of subblocks for the current block. The motion estimation unit 222 may be configured to determine a second set of subblocks that does not include any samples corresponding to prediction samples in a final prediction block that is generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. The motion estimation unit 222 may determine, for each subblock in the second set of subblocks, whether a majority of the subblock is in the first partition or in the second partition, and, for each subblock in the second set of subblocks, store a first motion vector based on the majority of the subblock being in the first partition or store a second motion vector based on the majority of the subblock being in the second partition.

[0117] As an example, filter unit 216 may utilize the motion vector stored for the sub-block for deblocking filtering. As another example, motion estimation unit 222 and motion compensation unit 224 may utilize the stored motion vector to construct a candidate list for merge mode or AMVP mode when encoding and decoding subsequent blocks.

[0118] For intra prediction or intra prediction codec, the intra prediction unit 226 may generate a prediction block based on samples adjacent to the current block. For example, for directional mode, the intra prediction unit 226 may generally mathematically combine the values ​​of adjacent samples and fill these calculated values ​​along a defined direction on the current block to generate a prediction block. As another example, for DC mode, the intra prediction unit 226 may calculate an average value of adjacent samples of the current block and generate a prediction block to include the obtained average value for each sample of the prediction block.

[0119] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original uncoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines a residual block for the current block. In some examples, the residual generation unit 204 may also use residual differential pulse code modulation (RDPCM) to determine the difference between the sample values ​​in the residual block to generate the residual block. In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0120] In the example where the mode selection unit 202 partitions the CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of the CU may refer to the size of the luma codec block of the CU, and the size of the PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra-frame prediction, and symmetric PU sizes of 2N×2N, 2N×N, N×2N, N×N or similar for inter-frame prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning of PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.

[0121] In an example where the mode selection unit 202 does not further split the CU into PUs, each CU may be associated with a luma codec block and a corresponding chroma codec block. As described above, the size of a CU may refer to the size of the luma codec block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0122] For other video codec techniques, such as intra block copy mode codec, affine mode codec, and linear model (LM) mode codec as some examples, the mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the codec technique. In some examples, such as palette mode codec, the mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how to reconstruct the block based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.

[0123] As described above, the residual generation unit 204 receives video data of a current block and a corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0124] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form a transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, for example, a primary transform and a secondary transform such as a rotation transform. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0125] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0126] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct a residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add samples of the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit 202 to generate a reconstructed block.

[0127] Filter unit 216 may perform one or more filter operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In some examples, the operations of filter unit 216 may be skipped.

[0128] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in examples where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks to the DPB 218. In examples where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed blocks to the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference picture from the DPB 218, which is formed by the reconstructed (and potentially filtered) blocks, to perform inter-frame prediction on blocks of subsequently encoded pictures. In addition, the intra-frame prediction unit 226 can use the reconstructed blocks in the DPB 218 of the current picture to perform intra-frame prediction on other blocks in the current picture.

[0129] In general, the entropy coding unit 220 may entropy encode syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 may entropy encode a block of quantized transform coefficients from the quantization unit 208. As another example, the entropy coding unit 220 may entropy encode a prediction syntax element (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from the mode selection unit 202. The entropy coding unit 220 may perform one or more entropy coding operations on syntax elements of another example of video data to generate entropy-encoded data. For example, the entropy coding unit 220 may perform a context-adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable to variable (V2V) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an exponential-Golomb coding operation, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 may operate in a bypass mode in which the syntax elements are not entropy encoded.

[0130] The video encoder 200 may output a bitstream including syntax elements of entropy coding required to reconstruct a block of a slice or a picture. Specifically, the entropy coding unit 220 may output the bitstream.

[0131] The above operations are described for blocks. Such descriptions should be understood as operations for luma codec blocks and / or chroma codec blocks. As described above, in some examples, the luma codec block and the chroma codec block are the luma and chroma components of the CU. In some examples, the luma codec block and the chroma codec block are the luma and chroma components of the PU.

[0132] In some examples, operations performed with respect to luma codec blocks do not need to be repeated for chroma codec blocks. As an example, operations for identifying motion vectors (MVs) and reference pictures for luma codec blocks do not need to be repeated to identify MVs and reference pictures for chroma blocks. Instead, the MVs of luma codec blocks may be scaled to determine the MVs of chroma blocks, and the reference pictures may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.

[0133] In one or more examples, for the GEO mode, the mode selection unit 202 may enable the entropy coding unit 220 to signal the video decoder 300 with information for reconstructing the current block. For example, the mode selection unit 202 may signal information indicating that the current block is encoded and decoded in the GEO mode, and may signal information indicating a manner in which the current block is segmented (e.g., based on an angle of a segmentation line). The mode selection unit 202 may also signal information for determining a first motion vector for the first segmentation and a second motion vector for the second segmentation.

[0134] The video encoder 200 represents an example of a device configured to encode video data, the device including a memory configured to store the video data, and one or more processing units implemented in a circuit, the processing unit configured to: determine one or more angles from a set of angles for segmenting a current block using a geometric segmentation mode (GEO), wherein the set of angles from which the one or more angles are determined is the same as a set of angles that can be used for a triangular segmentation mode (TPM); segment the current block based on the determined one or more angles; and encode the current block based on the segmentation of the current block. As another example, the video encoder 200 can be configured to determine one or more weights from a set of weights for blending the current block using GEO, wherein the set of weights from which the one or more weights are determined is the same as a set of weights that can be used for a triangular segmentation mode (TPM); blend a previous block based on the determined one or more weights; and encode the current block based on the blending of the current block. As another example, the video encoder 200 can be configured to determine a motion field storage for using GEO, wherein the motion field storage is the same as the motion field storage that can be used for TPM; and encode the current block based on the determined motion field storage.

[0135] Figure 4 is a block diagram illustrating an example video decoder 300 that may perform the techniques of this disclosure. Figure 4 The techniques widely exemplified and described in this disclosure are for the purpose of explanation and not limitation. For the purpose of illustration, this disclosure describes a video decoder 300 according to the technical description of VCC and HEVC. However, the techniques of this disclosure can be performed by video codec devices configured for other video codec standards.

[0136] exist Figure 4 In the example of , the video decoder 300 includes a codec picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or processing circuits. In addition, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.

[0137] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units to perform predictions according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0138] CPB memory 320 may store video data, such as an encoded video bitstream, to be decoded by components of video decoder 300. For example, the video bitstream may be obtained from computer readable medium 110 ( Figure 1 ) obtains video data stored in CPB memory 320. CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from an encoded video bitstream. Moreover, CPB memory 320 may store video data other than syntax elements of encoded and decoded pictures, such as temporary data representing outputs from various units of video decoder 300. Generally, DPB 314 stores decoded pictures, and video decoder 300 may output the decoded pictures and / or use them as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory 320 and DPB 314 may be formed by any of a variety of storage devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of storage devices. CPB memory 320 and DPB 314 may be provided by the same storage device or an independent storage device. In various examples, CPB memory 320 may be placed on-chip with other components of video decoder 300, or placed off-chip relative to those components.

[0139] Additionally or alternatively, in some examples, video decoder 300 may retrieve the video from memory 120 ( Figure 1 ) to retrieve the encoded and decoded video data. That is, the memory 120 can store data together with the CPB memory 320 as discussed above. Similarly, when some or all of the functions of the video decoder 300 are implemented in software to be executed by the processing circuitry of the video decoder 300, the memory 120 can store instructions to be executed by the video decoder 300.

[0140] Show Figure 4 Various units are shown to aid in understanding the operations performed by the video decoder 300. These units may be implemented as fixed function circuits, programmable circuits, or a combination thereof. Figure 3, fixed-function circuits refer to circuits that provide specific functions and are preset in the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (for example, to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more units may be different circuit blocks (fixed function or programmable), and in some examples, one or more units may be integrated circuits.

[0141] The video decoder 300 may include an ALU, an EFU, a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video decoder 300 is performed by software executed on a programmable circuit, an on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0142] The entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate decoded video data based on the syntax elements extracted from the bitstream.

[0143] Generally, the video decoder 300 reconstructs a picture on a block-by-block basis. The video decoder 300 may perform a reconstruction operation on each block individually (where the block currently being reconstructed (ie, decoded) may be referred to as a "current block").

[0144] The entropy decoding unit 302 may entropy decode syntax elements defining the quantized transform coefficients of the quantized transform coefficient block and transform information such as a quantization parameter (QP) and / or (one or more) transform mode indications. The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine a degree of quantization, and likewise, determine a degree of inverse quantization for the inverse quantization unit 306 to apply. The inverse quantization unit 306 may inverse quantize the quantized transform coefficients (e.g., performing a bitwise left shift operation). The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0145] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the transform coefficient block.

[0146] Further, prediction processing unit 304 generates a prediction block based on the prediction information syntax element entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-predicted, motion compensation unit 316 may generate the prediction block. In this case, the prediction information syntax element may indicate a reference picture in DPB 314 from which to retrieve the reference block, and a motion vector that identifies the position of the reference block in the reference picture relative to the current block in the current picture. Motion compensation unit 316 may generally generate a prediction block in the same manner as for motion compensation unit 224 ( Figure 3 ) is performed in a manner substantially similar to that described herein.

[0147] For example, the prediction processing unit 304 may determine that the current block is encoded and decoded in GEO mode. The prediction processing unit 304 may receive information indicating a manner of segmenting the current block (e.g., based on an angle of a segmentation line) and segment the current block into a first segmentation and a second segmentation. The prediction processing unit 304 may also determine a first motion vector for the first segmentation and a second motion vector for the second segmentation based on information signaled by the video encoder 200.

[0148] The motion compensation unit 316 may determine a first prediction block based on the first motion vector and determine a second prediction block based on the second motion vector. The motion compensation unit 316 may blend the first prediction block and the second prediction block based on a weight indicating an amount of scaled samples in the first prediction block and an amount of scaled samples in the second prediction block to generate a final prediction block for the current block. For example, the motion compensation unit 316 may determine a first scaling factor (e.g., 0, 1 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 7 / 8, or 8 / 8) for samples in the first prediction block and a second scaling factor for co-located samples in the second prediction block. In some examples, the first scaling factor and the second scaling factor add up to one.

[0149] As described above, the scaling factor may be a weight indicating the amount by which the samples in the first prediction block are scaled and the amount by which the samples in the second prediction block are scaled in order to generate the final prediction block of the current block. For example, in order to generate the prediction samples in the final prediction block, the motion compensation unit 316 may multiply the sample values ​​of the samples in the first prediction block by 3 / 8 and multiply the sample values ​​of the co-located samples in the second prediction block by 5 / 8 to generate the prediction sample values ​​of the co-located samples in the final prediction block. In this example, 3 / 8 and 5 / 8 are examples of the amount by which the samples in the first prediction block are scaled and the amount by which the samples in the second prediction block are scaled, respectively.

[0150] In one or more examples, the motion compensation unit 316 can be configured to store motion vector information of the current block (e.g., storing the motion vector information in the decoded picture buffer 314, the CPB memory 320, or the memory 120, as a few examples). To determine the motion vector information of the current block, the motion compensation unit 316 can divide the current block into a plurality of sub-blocks (e.g., each sub-block has a size of 4×4). The motion compensation unit 316 can determine a set of sub-blocks, each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of the samples in the first prediction block and the samples in the second prediction block. For example, the motion compensation unit 316 can determine which sub-blocks include at least one sample corresponding to a predicted sample in a final prediction block generated by scaling the samples in the first prediction block by 4 / 8 and scaling the samples in the second prediction block by 4 / 8. In this example, since the 4 / 8 scaling of the samples in the first prediction block and the second prediction block is the same, the samples in the first prediction block and the samples in the second prediction block have equal weights.

[0151] The motion compensation unit 316 may store a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks. For example, the motion compensation unit 316 may determine whether the first motion vector and the second motion vector are from different reference picture lists (e.g., reference pictures in different reference picture lists). The motion compensation unit 316 may be configured to do one of the following: (1) store both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks based on the first motion vector and the second motion vector being from different reference picture lists, or (2) select one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and store the selected one of the first motion vector or the second motion vector for each sub-block in the set of sub-blocks.

[0152] The set of subblocks may be considered as a first set of subblocks for the current block. The motion compensation unit 316 may be configured to determine a second set of subblocks that does not include any samples corresponding to prediction samples in a final prediction block that is generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. The motion compensation unit 316 may determine, for each subblock in the second set of subblocks, whether a majority of the subblock (e.g., a majority of the samples of the subblock) is within the first partition or within the second partition, and, for each subblock in the second set of subblocks, store a first motion vector based on the majority of the subblock being in the first partition or store a second motion vector based on the majority of the subblock being in the second partition.

[0153] As an example, the filter unit 312 may utilize the motion vector stored for the sub-block to perform deblocking filtering. As another example, the motion compensation unit 316 may utilize the stored motion vector to construct a candidate list for merge mode or AMVP mode when encoding and decoding subsequent blocks.

[0154] If the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 may generally generate a prediction block in the same manner as for the intra-prediction unit 226 ( Figure 3 The intra prediction process is performed in a manner substantially similar to that described in the foregoing description. The intra prediction unit 318 may retrieve data of neighboring samples from the DPB 314 to form the current block.

[0155] The reconstruction unit 310 may reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.

[0156] The filter unit 312 may perform one or more filter operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 may not necessarily be performed in all examples.

[0157] The video decoder 300 may store the reconstructed block in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 may store the reconstructed block to the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 may store the filtered reconstructed block to the DPB 314. As described above, the DPB 314 may provide reference information to the prediction processing unit 304, such as samples of the current picture for intra-frame prediction and previously decoded pictures for subsequent motion compensation. In addition, the video decoder 300 may output a decoded picture (e.g., a decoded video) from the DPB 314 for subsequent presentation on a display such as a video processor. Figure 1 on a display device 118 of the display device.

[0158] In this way, the video encoder 200 represents an example of a video decoding device, the device including a memory configured to store video data, and one or more processing units implemented in a circuit, the processing unit configured to: determine one or more angles from a set of angles for segmenting a current block using a geometric segmentation mode (GEO), wherein the set of angles from which the one or more angles are determined is the same as a set of angles that can be used for a triangular segmentation mode (TPM); segment the current block based on the determined one or more angles; and decode the current block based on the segmentation of the current block. As another example, the video codec 300 can be configured to determine one or more weights from a set of weights for blending the current block using GEO, wherein the set of weights from which the one or more weights are determined is the same as a set of weights that can be used for TPM; blend a previous block based on the determined one or more weights; and decode the current block based on the blending of the current block. As another example, the video decoder 300 can be configured to determine a motion field storage for using GEO, wherein the motion field storage is the same as the motion field storage that can be used for TPM; and decode the current block based on the determined motion field storage.

[0159] Some video codec standards are introduced below, including the geometric partitioning pattern (GEO) storage related technology of video codec standards. Video codec standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IECMPEG-4AVC), including its Scalable Video Codec (SVC) and Multiview Video Codec (MVC) extensions. In March 2010, ITU-T proposed H.264, "Advanced Video Codecs for Generic Audiovisual Services", which describes the latest MVC joint draft.

[0160] In addition, there is a High Efficiency Video Codec (HEVC) standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). As mentioned above, the specification text of the generic video coding and test model 6 (VTM6) can be referred to as VVC draft 6.

[0161] As introduced in JVET-N1002 (Sullivan et al., "Algorithmic Description of Generic Video Codec and Test Model 5", ITU-T SG 16WP3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 14th Meeting: Geneva, Switzerland, March 19-27, 2019), the triangular partitioning mode (TPM) is only applied to CUs encoded or decoded in skip or merge mode, but not in MMVD (merge with motion vector differencing) or CPU (combined inter and intra prediction) mode. For CUs that meet those conditions, a flag is signaled to indicate whether the triangular partitioning mode is applied.

[0162] When using the triangle partition mode, using diagonal partition or anti-diagonal partition, the CU is evenly divided into two triangle-shaped partitions, such as Figure 5A and Figure 5B For example, Figure 5A and Figure 5B Blocks 500A and 500B are shown, respectively. Block 500A is segmented into a first segment 502A and a second segment 504A, and block 500B is segmented into a first segment 502B and a second segment 504B.

[0163] Each triangular partition in a CU uses its own motion for inter prediction. For each partition, only unidirectional prediction may be allowed. That is, each partition (e.g., partition 502A, 504A, 502B, or 504B) has one motion vector and one reference index to the reference picture list. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, only two motion compensated predictions are required for each CU. For example, partition 502A and partition 504A may each have only one motion vector, meaning that block 500A is limited to only two motion vectors.

[0164] The unidirectional prediction motion for each partition (e.g., the first motion vector of the first partition 502A and the second motion vector of the second partition 504A, and similarly for partitions 502B and 504B) is derived from the unidirectional prediction candidate list constructed using the process described for unidirectional prediction candidate list construction. If the CU level flag indicates that the current CU is coded using triangular partitioning mode, and if triangular partitioning mode is used, a flag indicating the direction of triangular partitioning (diagonal or anti-diagonal) and two merge indices (one for each partition) are further signaled. Figure 5A and Figure 5B Examples of diagonal and anti-diagonal segmentations with corresponding dashed segmentation lines are shown.

[0165] After predicting each triangular partition (e.g., determining a first prediction block for the first partition 502A or 502B and a second prediction block for the second partition 504A or 504B based on the corresponding motion vector), a blending process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal edges (e.g., the partition lines). For example, the video encoder 200 and the video decoder 300 can generate a final prediction block by blending the first partition block and the second partition block with adaptive weights. This is the prediction signal for the entire CU (e.g., the final prediction block), and the transform and quantization process can be applied to the entire CU as in other prediction modes. As described with respect to blending along the triangular partition edges, the motion field of a CU predicted using the triangular partition mode is stored in 4×4 units.

[0166] The following describes the construction of a unidirectional prediction candidate list. The unidirectional prediction candidate list may include five unidirectional prediction motion vector candidates. The motion vector candidate is selected from five spatial neighboring blocks ( Figure 6 601-605) and two time-colocated blocks ( Figure 6 The motion vectors of the seven neighboring blocks are collected and put into the unidirectional prediction candidate list in the following order: first, the motion vectors of the unidirectionally predicted neighboring blocks; then, for the bidirectionally predicted neighboring blocks, the L0 (list 0) motion vector (i.e., the L0 motion vector part of the bidirectionally predicted MV), the L1 (list 1) motion vector (i.e., the L1 motion vector part of the bidirectionally predicted MV), and the average motion vector of the L0 and L1 motion vectors of the bidirectionally predicted MVS. If the number of candidates is less than five, a zero motion vector is added to the end of the list.

[0167] L0 or List 0 and L1 or List 1 refer to reference picture lists. For example, for inter-frame prediction, the video encoder 200 and the video decoder 300 each construct one or two reference picture lists (e.g., List 0 and / or List 1). The reference picture list includes multiple reference pictures, and the index in the reference picture list is used to identify one or more reference pictures for inter-frame prediction. List 0 motion vector or List 1 motion vector refers to a motion vector pointing to a reference picture identified in List 0 or List 1, respectively. For example, the video encoder 200 or the video decoder 300 can determine whether the motion vector is from a reference picture list, which can mean that the video encoder 200 and the video decoder 300 determine whether the motion vector points to a reference picture stored in a first reference picture list (e.g., List 0) or stored in a second reference picture list (e.g., List 1).

[0168] The following describes the inference of triangular prediction mode (TPM) motion from the merge list. The following describes the TPM candidate list construction. Given a merge candidate index, derive a unidirectional prediction motion vector from the merge candidate list. For a candidate in the merge list, the LX MV of that candidate (where X equals the parity of the merge candidate index value) is used as the unidirectional prediction motion vector for the triangular partitioning mode. These motion vectors are in Figure 7 In the example above, the L(1-X) motion vector of the same candidate in the extended merge prediction candidate list is marked with an "x". In the absence of a corresponding LX motion vector, the L(1-X) motion vector of the same candidate in the extended merge prediction candidate list is used as the unidirectional prediction motion vector for the triangular partitioning mode. For example, assuming that the merge list consists of 5 groups of bidirectional prediction motions, the TPM candidate list consists of the L0 / L1 / L0 / L1 / L0MVs of the 0th / 1st / 2nd / 3rd / 4th merge candidates from the first to the last. The TPM mode then includes signaling two different merge indices, one for each triangular partitioning, to indicate the use of the candidates in the TPM candidate list.

[0169] The following describes the blending along the edges of a triangle partition. After each triangle partition is predicted using the motion of the triangle partition itself, blending is applied to the two prediction signals to derive the samples around a diagonal or anti-diagonal edge. The following weights are used in this blending process: Fig. 8A As shown in , the values ​​used for brightness are {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8}, as Figure 8B As shown in , the ones used for chrominance are {6 / 8, 4 / 8, 2 / 8}.

[0170] in other words, Fig. 8A and Figure 8BThe final prediction blocks of luminance and chrominance generated for the current block are shown respectively, and the current block is divided by a diagonal line from the upper left corner to the lower right corner to form a first division and a second division. The video encoder 200 and the video decoder 300 can determine the first prediction block based on the first motion vector, and determine the second prediction block based on the second motion vector. The video encoder 200 and the video decoder 300 can determine the first prediction block based on the first motion vector, and determine the second prediction block based on the second motion vector. Fig. 8A and Figure 8B The weighting shown mixes the first prediction block and the second prediction block.

[0171] For example, to generate the upper left prediction sample in the final prediction block, the video encoder 200 and the video decoder 300 may scale the upper left prediction sample in the first prediction block by 4 / 8, and scale the upper left prediction sample in the second prediction block by 4 / 8, and add the results together (e.g., 4 / 8*P1+4 / 8*P2). To generate the prediction sample to the right of the upper left prediction sample in the final prediction block, the video encoder 200 and the video decoder 300 may scale the prediction sample to the right of the upper left prediction sample in the first prediction block by 5 / 8, and scale the prediction sample to the right of the upper left prediction sample in the second prediction block by 3 / 8, and add the results together (e.g., 5 / 8*P1+3 / 8*P2). The video encoder 200 and the video decoder 300 may repeat such operations to generate the final prediction block.

[0172] Some samples in the final prediction block may be equal to the co-located samples in the first prediction block or the second prediction block. For example, the upper right sample in the final prediction block 800A is equal to the upper right sample in the first prediction block, which is why the upper right sample in the final prediction block 800A is shown as P1. The lower right sample in the final prediction block 800A is equal to the lower right sample in the second prediction block, which is why the lower right sample in the final prediction block 800A is shown as P2.

[0173] therefore, Fig. 8A and Figure 8B It can be considered as showing weights indicating the amount by which the samples in the first prediction block are scaled and the amount by which the samples in the second prediction block are scaled in order to generate the final prediction block for the current block. For example, a weight of 4 means that the samples in the first prediction block are scaled at a ratio of 4 / 8 and the samples in the second prediction block are scaled at a ratio of 4 / 8. A weight of 2 means that the samples in the first prediction block are scaled at a ratio of 2 / 8 and the samples in the second prediction block are scaled at a ratio of 6 / 8.

[0174] Fig. 8A and Figure 8BThe weights shown in are an example. For example, the weights may be different for blocks of different sizes. Moreover, the dividing line may not be from one corner of a block to another corner of the block. For example, in GEO mode, the dividing line may be at different angles. That is, the TPM mode can be considered an example of GEO mode, where the dividing line is a diagonal or anti-diagonal line of the block. However, as illustrated and described in more detail below, in GEO mode, there may be other angles for the dividing line. In one or more examples, there may be different weights for different angles of segmentation. The video encoder 200 and the video decoder 300 may store weights for different angles of segmentation and use the stored weights to determine the amount to scale the samples in the first prediction block and the second prediction block.

[0175] The motion field storage is described below. The motion vectors of a CU encoded and decoded in a triangular partitioning mode (TPM) or a GEO mode are stored in 4×4 units. That is, the video encoder 200 and the video decoder 300 may divide the current block into sub-blocks (e.g., of size 4×4) and store motion vector information for each sub-block. Depending on the position of each 4×4 sub-block, in some techniques, the video encoder 200 and the video decoder 300 may store a unidirectional prediction or a bidirectional prediction motion vector, represented as Mv1 and Mv2, respectively, as a unidirectional prediction motion vector for partition 1 and partition 2. As described above, the position of each 4×4 sub-block may determine whether a unidirectional prediction motion vector or a bidirectional prediction motion vector is stored. In some examples, the position of each 4×4 sub-block may also indicate whether most of the 4×4 sub-block is in the first partition or the second partition (e.g., whether most of the samples of the 4×4 sub-block are in the first partition or the second partition).

[0176] When the CU is split from the upper left corner to the lower right corner (i.e., 45° split), split 1 and split 2 are triangular blocks located at the upper right corner and lower left corner, respectively; and when the CU is split from the upper right corner to the lower left corner (i.e., 135° split), split 1 and split 2 are triangular blocks located at the upper left corner and lower right corner, respectively. Figure 5A As shown in FIG. 1 , when the segmentation line is 45°, segmentation 502A is segmentation 1, and segmentation 504A is segmentation 2, and as shown in FIG. Figure 5B As shown in , when the segmentation line is 135°, segmentation 502B is segmentation 1 and segmentation 504B is segmentation 2. Although the examples are described with respect to triangular segmentation with a segmentation line of 45° or 135°, the technology is not limited thereto. The example technology can be applied to examples where triangular segmentation does not exist, such as in various examples of GEO mode. Even in these examples of GEO mode, there can be a first segmentation (e.g., segmentation 1) and a second segmentation (e.g., segmentation 2).

[0177] If a 4×4 unit (e.g., a sub-block) is located Fig. 8A and Figure 8B In the unweighted region shown in the example of , Mv1 or Mv2 is stored for the 4×4 cell. Fig. 8A and Figure 8B In the example, where the unweighted samples from the first prediction block or the second prediction block are samples in the final prediction block, the video encoder 200 and the video decoder 300 may select Mv1 (e.g., the first motion vector) or Mv2 (e.g., the second motion vector). Otherwise, if the 4×4 unit is located in the weighted region, the bidirectional prediction motion vector is stored. The bidirectional prediction motion vector is derived from Mv1 and Mv2 according to the following process:

[0178] a. If Mv1 and Mv2 come from different reference picture lists (one from L0 and the other from L1), simply combine Mv1 and Mv2 to form a bidirectional prediction motion vector. That is, the motion vector information of the sub-block includes both Mv1 and Mv2 (i.e., both the first motion vector and the second motion vector).

[0179] b. Otherwise, if Mv1 and Mv2 are from the same list, only Mv2 is stored (ie, only the second motion vector is stored).

[0180] Geometric partitioning is described below. Geometric partitioning was introduced in JVET-O0489 (Esenlik et al., Joint Video Experts Group (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC 29AVG 11, 15th Meeting: Gothenburg, Sweden, July 3-12, 2019) as a proposed extension to the non-rectangular partitioning introduced by TPM. As introduced in JVET-O0489, the geometric partitioning mode (GEO) is only applied to CUs encoded or decoded in skip or merge mode, not to CUs encoded or decoded in MMVD or CIIP mode. For CUs that meet those conditions, a flag is signaled to indicate whether GEO is applied. Fig. 9 The TPM in VVC-6 (VVC Draft 6) and the additional shapes proposed for non-rectangular inter blocks are shown.

[0181] For example, Fig. 9 Blocks 900A and 900B are shown partitioned in TPM mode. TPM mode can be viewed as a subset of GEO mode. In TPM, the partition line extends from one corner to the opposite corner of the block, as shown in blocks 900A and 900B. However, in GEO mode, in general, the partition line is a diagonal line and does not necessarily start from or end at a corner of the block. Fig. 9 Blocks 900C-900I in FIG. 9A show different examples of segmentation lines that do not necessarily start or end at the corners of blocks 900C-900I.

[0182] The total number of GPM divisions can be 140. For example, Fig. 9 Blocks 900A-900I are shown, but there may be 140 possible ways to partition the blocks, which is desirable for flexibility in determining how to partition, but can increase signaling. There may be additional signaling for GEO (e.g., Fig.10 ), such as the angle α and the displacement of the separation line relative to the center of the block ρ. In this example, α and ρ together define Fig.10 The position and slope of the dividing line of block 1000 in FIG. 1000. In some examples, α represents a quantized angle between 0 and 360 degrees with a separation of 11.25 degrees, and ρ represents a distance with 5 different values ​​(e.g., 5 displacements). The value α and ρ pairs are stored in a table of size 140×(3+5) / 8=140 bytes. For example, for a separation of 11.25 degrees, there can be 32 angles (e.g., 11.25*32=360). For 5 displacement values, there can be 160 modes, where one mode is a combination of an angle and a displacement (e.g., 32*5=160). It is possible that certain redundant modes can be removed, such as an angle of 0 degrees with 0 displacement and an angle of 180 degrees with 0 displacement, which give the same result. By removing redundant modes, the number of modes is 140. If 3 bits are needed to store 5 displacement values, and 5 bits are needed to store 32 angle values, then there are a total of 140*(3+5) bits, divided by 8 bits per byte, we get 140 bytes.

[0183] In some examples, video encoder 200 may signal an index into the table, and video decoder 300 may determine the α and ρ values. Based on the α and ρ values, video decoder 300 may determine the number of blocks (such as Fig.10 The dividing line of block 1000).

[0184] Similar to TPM, GEO partitioning for inter prediction is allowed for unidirectional prediction blocks no smaller than 8×8 to have the same memory bandwidth as bidirectional prediction blocks at the decoder side (eg, video decoder 300). Motion vector prediction for GEO partitioning can be aligned with TPM.

[0185] Table 1 below describes the mode signaling. According to the technique described in JVET-O0489, GEO mode is signaled as an additional merge mode.

[0186]

[0187] Table 1 Suggested grammatical elements

[0188] geo_merge_idx0 and geo_merge_idx1 are encoded and decoded using the same CABAC context and binarization as the TPM merge index. geo_partition_idx indicates the partition mode (140 possibilities) and is encoded and decoded using truncated binary binarization and bypass codec. For example, Fig.10 As shown in , geo_partition_idx is the index described above to determine the α and ρ values ​​used to determine the partition line.

[0189] When GEO mode is not selected, TPM may be selected. The segmentation of GEO mode does not include the segmentation that can be obtained by TPM with binary partitioning. To some extent, the signaling scheme proposed by JVET-O049 is similar to intra-frame mode signaling, where the TPM segmentation corresponds to the most likely segmentation and the GEO mode corresponds to the remaining segmentation. geo_partition_idx is used as an index into a lookup table that stores α and ρ pairs. As mentioned above, 140 bytes are used to store this table.

[0190] The following describes the blending operation for the luminance block. As in the case of TPM, the final prediction for the codec block is obtained by weighted averaging the first unidirectional prediction and the second unidirectional prediction according to the sample weights. sampleWeightL[x][y] = GeoFilter[distScaled] If distFromLine <= 0

[0191] sampleWeightL[x][y]=8-GeoFilter[distScaled] if distFromLine>0 where the sample weights are implemented as a lookup table, as shown in Table 2 below:

[0192]

[0193] Table 2 Hybrid filter weights

[0194] The number of operations required to calculate the sample weights is approximately one addition operation per sample, and its computational complexity is similar to TPM. In more detail, for each sample, distScaled is calculated according to the following two equations:

[0195] distFromLine=((x<<1)+1)*Dis[displacementX]+

[0196] ((y<<1)+1))*Dis[displacementY]-rho

[0197] distScaled=min((abs(distFromLine)+8)>>4,14)

[0198] where the variables rho, displacementX, and displacementY are calculated once per codec block, and Dis[] is a lookup table with 32 entries (8-bit resolution) that stores the cosine values. distFromLine can be calculated by incrementing it for each sample, with the value in a sample row being 2^Dis[displacementX], and the value from one sample row to the next being 2^Dis[displacementX]. The distFromLine values ​​can be obtained with slightly more than 1 addition per sample. Additionally, minimum, absolute value, and downshift operations can be used without introducing any significant complexity.

[0199] All operations of GEO can be implemented using integer arithmetic. The computational complexity of GEO is likely to be very similar to that of TPM. There are additional details about blending operations in the draft specification revision document, for example in the section "8.5.7.5 Sample weight derivation process for geometric partitioning merge mode" provided by JVET-O0489.

[0200] The following describes the blending operation of the chroma block. The sample weights calculated for the luma samples are subsampled and used for chroma blending without any calculation. With respect to the top left corner sample of the luma block, the chroma sample weight at coordinate (x, y) is set equal to the luma sample weight at coordinate (2x, 2y).

[0201] Motion vector derivation is described below. The same merge list derivation process used for TPM is used to derive the motion vectors for each partition of the GEO block. Each partition can only be predicted by unidirectional prediction.

[0202] An example technique for motion vector storage is described below. The luminance sample weights at the four corners of a 4×4 motion storage unit (e.g. Fig. 8A and 8B As shown in , the sum is calculated based on the description of the blend along the triangular partition edge. The sum is compared with two thresholds to determine whether to store one of the two unidirectional prediction motion information or the bidirectional prediction motion information. The bidirectional prediction motion information is derived using the same process as TPM.

[0203] In other words, in some techniques, the video encoder 200 and the video decoder 300 may divide the current block into sub-blocks (e.g., 4×4 sub-blocks). For each sub-block, the video encoder 200 and the video decoder 300 may determine a sample weight for scaling the samples in the first prediction block and the second prediction block. For example, referring back to Fig. 8A , there may be a 4×4 sub-block at the upper left corner of block 800A. The 4×4 sub-block includes a sample located at the upper left corner, and the sample weight is 4, indicating that the sample in the first prediction block is co-located with the position of the upper left corner sample in block 800A, and the samples in the second prediction block are similarly scaled (e.g., 4 / 8*P1+4 / 8*P2, as shown in FIG. Fig. 8A ). The 4×4 sub-block includes a sample located at the lower left corner, and the sample weight is 1, indicating that the sample of the first prediction block co-located with the sample at the lower left corner of the 4×4 sub-block is scaled by 1 / 8, and the sample in the second prediction block similarly co-located is scaled by 7 / 8. The 4×4 sub-block includes a sample located at the upper right corner, and the sample weight is 7, indicating that the sample of the first prediction block co-located with the sample at the upper right corner of the 4×4 sub-block is scaled by 7 / 8, and the sample in the second prediction block similarly co-located is scaled by 1 / 8, as shown in Fig. 8A The 4×4 sub-block includes a sample located at the lower right corner, and the sample weight is 4, indicating that the sample of the first prediction block co-located with the sample at the lower right corner of the 4×4 sub-block is scaled by 4 / 8, and similarly the sample in the co-located second prediction block is scaled by 4 / 8, meaning that the samples in the first prediction block and the second prediction block are scaled by the same amount.

[0204] In this example, the video encoder 200 and the video decoder 300 may sum the weights of the four corners of the 4×4 sub-block, which is 4+7+1+4=16. The video encoder 200 and the video decoder 300 may compare the resulting value (e.g., 16) with two thresholds to determine whether the video encoder 200 and the video decoder 300 store a first motion vector identifying the first prediction block, a second motion vector identifying the second prediction block, or both the first motion vector and the second motion vector as motion vector information for the 4×4 sub-block.

[0205] According to the present technology, instead of or in addition to summing the luma sample weights as described above, the video encoder 200 and the video decoder 300 may determine a set of subblocks, each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in a first prediction block and samples in a second prediction block. As an example, the video encoder 200 and the video decoder 300 may determine a set of subblocks (e.g., 4×4 subblocks), each of which includes at least one sample with a weight of 4. As described above, if the samples have a weight of 4, the video encoder 200 may scale the co-located samples in the first prediction block by 4 / 8 (1 / 2) and scale the co-located samples in the second prediction block by 4 / 8 (1 / 2) to generate the predicted samples in the final prediction block.

[0206] In this manner, it may not be necessary to sum the weights at the corners of the sub-block and compare to a threshold, as is done in some other techniques described above. Instead, the video encoder 200 and the video decoder 300 may determine whether the sub-block includes a sample for which a prediction sample in the final prediction block is generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. If the sub-block includes such a sample, the video encoder 200 and the video decoder 300 may store a bidirectional prediction motion vector, and if the sub-block does not include such a sample, the video encoder 200 and the video decoder 300 may store a unidirectional prediction motion vector.

[0207] As described above, the bidirectional prediction motion vector is not necessarily two motion vectors. Instead, the video encoder 200 and the video decoder 300 may perform certain operations to determine the bidirectional prediction motion vector. As an example, the video encoder 200 and the video decoder 300 may determine whether the first motion vector identifying the first prediction block and the second motion vector identifying the second prediction block are from different reference picture lists, and one of the following: based on the first motion vector and the second motion vector being from different reference picture lists, storing both the first motion vector and the second motion vector for the sub-block, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and storing the selected one of the first motion vector or the second motion vector for the sub-block.

[0208] The unidirectional prediction motion vector may be one of the first motion vector or the second motion vector. For example, the video encoder 200 and the video decoder 300 may determine a subblock that does not include any samples corresponding to the prediction samples in the final prediction block generated based on equal weighting of the samples in the first prediction block and the samples in the second prediction block (for example, the sample weight of none of the samples in the subblock is 4). In this example, the video encoder 200 and the video decoder 300 may determine, for the subblock, whether most of the subblock is in the first partition or in the second partition, and, for the subblock, store the first motion vector based on the fact that most of the subblock is in the first partition or store the second motion vector based on the fact that most of the subblock is in the second partition.

[0209] By storing motion vector information using the example techniques described in the present disclosure, the video encoder 200 and the video decoder 300 can store motion vector information that provides overall encoding and decoding and visual gain. For example, the motion vector information stored for each sub-block may affect the strength of the deblocking filter. Using the example techniques described in the present disclosure, the motion vector information of the sub-block may lead to the determination of the strength of the deblocking filter to remove artifacts. Moreover, the motion vector information stored for each sub-block may affect the candidate list generated for the merge mode or the AMVP mode for encoding and decoding subsequent blocks. Using the example techniques described in the present disclosure, the motion vector information of the sub-block can be a better candidate for the candidate list than other techniques for generating the candidate list.

[0210] Some of the transformations for GEO are described below. Since GEO segmentation provides flexibility for inter prediction, even objects with complex shapes can be predicted well; hence, the residuals are smaller compared to rectangular or triangular blocks. In some cases, non-zero residuals for GEO blocks are only observed around internal boundaries, such as Fig.11 Some techniques to reduce the residual include setting the size or W×H block to W×(H / n) or (W / n)×H, keeping non-zero residual only around internal boundaries, such as Fig.11 1100A. Fig.11 As shown in FIG. 1 , the area captured by reference numeral 1100A can be reoriented to form a rectangular block 1100B. With this change, the residual size to be processed and the number of coefficients to be signaled can be made smaller by a factor of n. The value of n is signaled in the bitstream and can be equal to 1 (when no partial transform is applied), 2, or 4. In some examples, a partial transform, an example of which is shown in FIG. Fig.11 As shown in , it may only be applicable to GEO and TPM blocks.

[0211] The residual propagation of the partial transform is described below. On top of the partial transform, a process similar to the deblocking form is applied.

[0212] The deblocking operation can be described as follows:

[0213] - Get sample point p on the block boundary marked by the black circle

[0214] - If k <= 3, assign the value of (p>>k) to the sample at position (x, yk)

[0215] - Get sample point p on the block boundary marked by the white circle

[0216] - If k<=3, assign the value of (p>>k) to the sample at position (x, y+k).

[0217] Fig.12 Examples of k values ​​are provided. The value of k is the number of samples between the boundary samples in block 1200 (represented by circles) and the fill samples in block 1200 (under the corresponding arrows). The black and white arrows show the deblocking direction of the upper and lower boundaries.

[0218] There may be certain issues with the GEO design. For example, the current GEO design can be considered an extension of the TPM. However, there are some differences in the design that need to be coordinated for implementation.

[0219] The following describes the coordination issue of TPM and GEO motion field storage. The TPM algorithm only uses the location of the 4×4 unit in the CU to determine the motion vector that needs to be stored, while the GEO method uses the weights used for motion compensation for storage. In addition, if the existing GEO motion field storage algorithm is applied to TPM, the TPM storage result will change. For these two methods, it is better to have a unified storage method. The present disclosure describes example techniques for storing motion fields (e.g., motion vector information).

[0220] The following describes the coordination problem of TPM and GEO weight derivation. The GEO algorithm for weight derivation described above for the blending operation of the luma block is different from the GEO algorithm used for TPM weight derivation. It would be desirable to have a unified weight derivation method for both methods.

[0221] The problem of partial transforms is described below. Current GEO technology has a partial transform option that applies residual filtering after a partial inverse transform. Current GEO technology utilizes a transform module that may also add additional stages in the implementation. In one example, vertical residual filtering may not be performed "on the fly" (i.e., right after the inverse transform) because the read in the vertical direction may require reading more data from the buffer and cannot be simply done while copying the residual from the transform output block into the residual block. In this case, an additional stage in the implementation may be required.

[0222] This disclosure describes example techniques for unifying motion field storage and motion weight derivation for TPM and GEO. As an example, the following describes a change in GEO angle. In some examples, the angle used for GEO can be changed to coordinate with the existing TPM angle. The available angle for GEO will become Fig.13 , which can be described by an approximate angle value expressed in degrees or radians, or by an exact integer ratio of width / height when each angle is from corner pixel to corner pixel of each block. 0° and 90° angles can also be added to provide more variety.

[0223] The following describes changes to the derivation of GEO motion weights. In some examples, the weights used in the GEO for blending can be changed so that they blend along the edges of the triangulated segments using the TPM weight processing described above, and so that both GEO and TPM have a coordinated weight derivation process. Given a starting point with coordinates (sx; sy) and an edge endpoint with coordinates (ex; ey), the edge ( Fig.14 Identify the bounding box (the dashed rectangle in the figure), that is, the bounding box with upper left corner coordinates (bx;by) = (min(sx,ex); (min(sy,ey)) and size abs(sx-ex) multiplied by abs(sy-ey).

[0224] For example, Fig.14 Block 1400 is shown partitioned by partition line 1402. Video encoder 200 and video decoder 300 may determine weights for blending the two partitions of block 1400. Again, the weights may indicate how much to scale samples in a first prediction block identified by a first motion vector from a first partition, and how much to scale samples in a second prediction block identified by a second motion vector from a second partition (e.g., as Fig. 8A and 8B shown).

[0225] exist Fig.14 In the example of , to determine the weight, the video encoder 200 and the video decoder 300 may determine a bounding box 1404, wherein one corner of the bounding box 1404 is one end of the dividing line 1402, and the diagonal of the bounding box 1404 is the other end of the dividing line 1402, such as Fig.14As shown, segmentation line 1402 then divides bounding box 1404 into segmentation 1406A and segmentation 1406B. Video encoder 200 and video decoder 300 can then determine motion vectors for segmentation 1406A and 1406B, determine prediction blocks for segmentation 1406A and 1406B, and blend the first prediction block and the second prediction block based on the weight of bounding box 1404. In some examples, the weight of bounding box 1404 can be similar to Fig. 8A and Figure 8B Again, the weights indicate the amount by which the samples in the first prediction block are scaled and the amount by which the samples in the second prediction block are scaled in order to generate the final prediction block.

[0226] By using TPM angles for GEO, the TPM weight calculation can be applied directly to the bounding box (e.g. similar to Fig. 8A and 8B ). One difference may be that the video encoder 200 and the video decoder 300 add the offset (bx; by) to the start position of the weighted calculation. In some examples, the mixed area may extend outside the bounding box.

[0227] In the TPM weight calculation, the weights for the 0° and 90° angles can be calculated by using an offset of 0 or "infinity". In some examples, the weight calculation is done by using a width / height ratio of 0 for the 0° angle and a width / height ratio of 2*MAX_CU_SIZE for the 90° angle. The weight calculation of the existing TPM can also be changed. For example, in some examples, the mixing area for the TPM weights can be reduced to have sharper weights, such as Fig. 15B middle. Fig.15A shows the existing TPM weights with a width / height ratio equal to 4, and Fig. 15B An example of different weights for the same angle is shown.

[0228] In some examples, weights can be modified in the following ways:

[0229] a. The TPM weight derivation can be changed to use the GEO weight derivation described in the previous section;

[0230] b.GEO weight derivation can use TPM's chrominance weight derivation, while TPM still uses its luminance weight derivation;

[0231] c. TPM and GEO weights can subsample existing weights for TPM to have a smaller mixing area; and

[0232] d. The TPM and GEO weights may be made to subsample the existing weights for the TPM to have smaller mixed areas for certain angles or combinations of angles and block sizes (e.g., if angles assumed for higher aspect ratios are used for blocks with smaller aspect ratios).

[0233] For example, Fig.15A Block 1500 is shown as a block of 32×8 size, and is partitioned into partition 1503A and partition 1503B by partition line 1502. In one or more examples, video encoder 200 and video decoder 300 may partition block 1500 including sub-blocks (e.g., 4×4 sub-blocks), as shown by sub-blocks 1504A-1504P. As shown, each sample in each sub-block may be associated with a weight indicating an amount by which the samples in the first prediction block are scaled and an amount by which the samples in the second prediction block are scaled in order to generate a final prediction block.

[0234] For example, for sub-block 1504A, the video encoder 200 and the video decoder 300 may determine that for the first top left corner sample of sub-block 1504A, the video encoder 200 and the video decoder 300 scale the co-located samples in the first prediction block by 7 / 8 and scale the co-located samples in the second prediction block by 1 / 8, and sum the resulting values ​​as part of determining the prediction samples in the final prediction block. The video encoder 200 and the video decoder 300 may determine the prediction sample for each sample in block 1500 based on the weight associated with the sample. In one or more examples, if a sample in block 1500 has a weight of 4 (e.g., the top left corner sample in sub-block 1504P), the video encoder 200 and the video decoder 300 may scale the co-located samples in the first prediction block and the second prediction block equally (e.g., scaling each co-located sample by 4 / 8 or 1 / 2).

[0235] The same technique can be applied to Fig. 15B 1504P. As shown, each sample in each sub-block may be associated with a weight indicating an amount by which the samples in the first prediction block are scaled and an amount by which the samples in the second prediction block are scaled in order to generate the final prediction block.

[0236] For example, for sub-block 1510A, the video encoder 200 and the video decoder 300 may determine that for the first upper left corner sample of sub-block 1510A, the video encoder 200 and the video decoder 300 scale the co-located samples in the first prediction block by 8 / 8 and scale the co-located samples in the second prediction block by 0 / 8, and sum the resulting values ​​as part of determining the prediction samples in the final prediction block. In this example, the second prediction block does not contribute to the generation of the final prediction block. The video encoder 200 and the video decoder 300 may determine the prediction sample for each sample in block 1506 based on the weight associated with the sample. In one or more examples, if the sample in block 1506 has a weight of 4 (e.g., the upper left corner sample in sub-block 1510P), the video encoder 200 and the video decoder 300 may scale the co-located samples in the first prediction block and the second prediction block equally (e.g., scaling each co-located sample by 4 / 8 or 1 / 2).

[0237] The following describes changes in GEO motion field storage. Specifically, the video encoder 200 and the video decoder 300 can be configured to store motion vector information for each sub-block in the sub-blocks 1504A-1504P and 1510A-1510P. For example, the video encoder 200 and the video decoder 300 can determine a first motion vector for the first partition 1503A and a second motion vector for the second partition 1503B. The video encoder 200 and the video decoder 300 can determine a first prediction block based on the first motion vector and determine a second prediction block based on the second motion vector. The video encoder 200 and the video decoder 300 can ... Fig.15A The weights shown are used to mix the first prediction block and the second prediction block to generate the final prediction block. The video encoder 200 and the video decoder 300 can perform similar operations to determine Fig. 15B The final predicted block of block 1506.

[0238] In addition, the video encoder 200 and the video decoder can determine the motion vector information (e.g., motion field) stored for each of the sub-blocks 1504A-1504P and 1510A-1510P and use it for future encoding and decoding operations (such as deblocking filtering or candidate lists for merge mode or AMVP mode for encoding and decoding subsequent blocks). An example of a manner of determining the motion vector information stored for the sub-blocks 1504A-1504P and 1510A-1510P is described below.

[0239] In some examples, the GEO's kinematic storage can be modified in the following way so that both the TPM and the GEO use the same kinematic storage:

[0240] a. In some examples, a search may be performed within each 4x4 cell of the weight map to use the biMv motion vector if and only if a '4' is found within the block.

[0241] b. In some examples, a similar process for weights as described above with respect to GEO angle variation may be used, using TPM storage for blocks with bounding box dimensions offset by (bx; by).

[0242] c. In some examples, (bx; by) can be used to find the starting point of the weighted region to store biMv to the corresponding 4×4 cell, and the offset (ox; oy) is determined by the slope of the angle used from the width / height ratio of the angle. For example, Fig.13 The angles presented in are described by the following ratios:

[0243] {0:1; 1:8; 1:4; 1:2; 1:1; 2:1; 4:1; 8:1;

[0244] MAX_CU_SIZE<<1:1;-8:1;-4:1;-2:1;-1:1;-1:2;-1:4;-1:8}

[0245] This will result in an offset (ox;oy) equal to:

[0246] {(0;1);(1;8);(1;4);(1;2);(1;1);(2;1);(4;1);(8;1);

[0247] (MAX_CU_SIZE<<1;1);(-8;1);(-4;1);(-2;1);(-1;1);(-1;2);(-1;4);(-1;8)}

[0248] Or, from a more general perspective:

[0249] ox=(a!=90)? max(1,abs(tan(a))):MAX_CU_SIZE<<1;

[0250] ox=(a==0)? 0:(a>90)? -ox:ox;

[0251] oy=(a%90!=0)? max(1,1 / tan(a)):1;

[0252] where a is equal to the angle used, expressed in degrees, modulo 180°.

[0253] These points (b.x+i*ox; b.y+i*oy) are identified, and the video encoder 200 and the video decoder 300 store biMv to a 4×4 unit containing at least one of these points.

[0254] As described above, the video encoder 200 and the video decoder 300 may search each 4×4 subblock of the weight map to use the biMv motion vector if and only if a '4' is found within the block. Also as described above, a weight of '4' for a sample may mean that the video encoder 200 and the video decoder 300 scale the co-located samples in the first prediction block and the second prediction block by equal weighting (e.g., 4 / 8*Pl+4 / 8*P2, where Pl refers to a sample in the first prediction block and P2 refers to a sample in the second prediction block). In other words, the video encoder 200 and the video decoder 300 may determine a set of subblocks (e.g., in subblocks 1504A-1504P and 1510A-1510P), each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. Examples of these sub-blocks include any one of sub-blocks 1504A-1504P and 1510A-1510P, each of which includes at least one sample with a weight of 4.

[0255] In one or more examples, the video encoder 200 and the video decoder 300 can store a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks. For example, the video encoder 200 and the video decoder 300 can determine whether the first motion vector and the second motion vector are from different reference picture lists, and one of the following: based on the first motion vector and the second motion vector being from different reference picture lists, storing both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and storing the selected one of the first motion vector or the second motion vector for each sub-block in the set of sub-blocks.

[0256] As an example, it is assumed that a first motion vector (Mv1) refers to a first picture in reference picture list 1. In this example, motion vector information of the first motion vector is L0: refIdx=-1, Mv=(0,0), L1: refIdx=1, Mv1=(x1, y1), which means that the first motion vector identifies a picture with an index of '1' in reference picture list 1. It is assumed that a second motion vector (Mv2) refers to a second picture in reference picture list 0. In this example, motion vector information of the second motion vector is L0: refIdx=2, Mv2=(x2, y2), L1: refIdx=0, Mv=(x', y'), which means that the second motion vector identifies a picture with an index of '2' in reference picture list 0. In this example, because Mvl and Mv2 reference pictures in different reference picture lists, the video encoder 200 and the video decoder 300 can combine the motion vectors for storage (for example, the stored motion vector information can be L0: refIdx=2, Mv2=(x2, y2), L1: refIdx=1, Mv1=(x1, y1).

[0257] However, if the first motion vector and the second motion vector refer to a picture in the same reference picture list (e.g., reference picture list 0 (L0) or reference picture list 1 (L1)), the video encoder 200 and the video decoder 300 may select one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and store the selected one of the first motion vector or the second motion vector for each sub-block in the sub-block set. As an example, the video encoder 200 and the video decoder 300 may be configured to always store Mv2.

[0258] An example of a manner in which the video encoder 200 and the video decoder 300 may store a corresponding bidirectional prediction vector for each subblock in a determined set of subblocks is described above. The determined set of subblocks may be a first set of subblocks, each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in a first prediction block and samples in a second prediction block.

[0259] In one or more examples, the video encoder 200 and the video decoder 300 may be configured to determine a second set of subblocks that do not include any samples corresponding to prediction samples in a final prediction block that is generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. Examples of such subblocks include any of subblocks 1504A-1504P and 1510A-1510P, each of which does not include any samples with a weight of 4. The video encoder 200 and the video decoder 300 may determine, for each subblock in the second set of subblocks, whether a majority of the subblock is in the first partition or in the second partition, and, for each subblock in the second set of subblocks, store a first motion vector based on the majority of the subblock being in the first partition or store a second motion vector based on the majority of the subblock being in the second partition.

[0260] As described above, in some examples, (bx; by) can be used to find the starting point of the weighted region to store biMv to the corresponding 4×4 cell, and the offset (ox; oy) is determined by the slope of the angle used from the width / height ratio of the angle. For example, Fig.15A and 15B are two examples of weights applied to samples. In some examples, for another example of weights applied to samples and another block, the partition line of the other block may have the same slope as the slope of partition line 1502 or 1508, but the starting and ending points of the partition line may be different. In such an example, the weights applied to samples in the sub-blocks of the other block may be the same as Fig.15A and Fig. 15B Same example as shown in , but with an offset.

[0261] For example, Fig.15A The third sub-block in block 1500 includes samples in the first row with a weight of 5, samples in the second row with a weight of 6, samples in the third row with a weight of 7, and samples in the fourth row with a weight of 8. In some examples, if the partition line 1502 is moved to the left by one sub-block, the third sub-block may have the same weight as the fourth sub-block in block 1500.

[0262] exist Fig.15AIn the example of block 1500, in the top row of block 1500, the motion vectors stored for the sub-blocks may be as follows: the first sub-block (e.g., sub-block 1504A), storing a unidirectional prediction motion vector because there is no weight '4', the second sub-block, storing a unidirectional prediction motion vector because there is no weight '4', the third sub-block, storing a unidirectional prediction motion vector because there is no weight '4', the fourth sub-block, storing a bidirectional prediction motion vector because at least one sample has a weight '4', the fifth sub-block, storing a bidirectional prediction motion vector because at least one sample has a weight '4', the sixth sub-block, storing a bidirectional prediction motion vector because at least one sample has a weight '4', the seventh sub-block, storing a bidirectional prediction motion vector because at least one sample has a weight '4', and the eighth sub-block, storing a unidirectional prediction motion vector because there is no weight '4'. Therefore, the motion vector storage of the sub-blocks in the first row of block 1500 may be as follows: unidirectional prediction, unidirectional prediction, unidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, and unidirectional prediction.

[0263] If the partition line 1502 is moved one sub-block to the left and the weights are changed as described below, the motion vector storage of the sub-blocks in the first row can be as follows: unidirectional prediction, unidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, unidirectional prediction, and unidirectional prediction. This pattern of which sub-blocks have unidirectional prediction motion vectors and which sub-blocks have bidirectional prediction motion vectors (e.g., unidirectional prediction, unidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, unidirectional prediction) is the same as the pattern of which sub-blocks have unidirectional prediction motion vectors and which sub-blocks have bidirectional prediction motion vectors (e.g., unidirectional prediction, unidirectional prediction, unidirectional prediction, bidirectional prediction, bidirectional prediction, bidirectional prediction, unidirectional prediction) described above.

[0264] For the partition line 1502, the video encoder 200 and the video decoder 300 may store a weight table for the slope of the partition line 1502, the weight table indicating which sub-blocks have a weight of '4' and which sub-blocks do not have a weight of '4' for the slope. Then, in order to determine which sub-block set has samples with a weight of '4', the video encoder 200 and the video decoder 300 may access the table to determine which sub-block set has samples with a weight of '4'. For example, the video encoder 200 and the video decoder 300 may store a weight table for the slope of the partition line used to partition the current block into a first partition and a second partition. In order to determine the sub-block set, the video decoder 300 may determine the sub-block set based on the stored table and the offset into the table.

[0265] The following describes some of the transformations and residual propagation. In some examples, Fig.14 The bounding box of the description can also be made so that the width and height are powers of 2, such as Fig.16 For example, Fig.16Block 1600 is shown with a bounding box 1602 and GEO edges (eg, dividing lines) 1604 . Figure 6 Also shown is a virtual bounding box 1606 , which may have a width and height that are powers of two.

[0266] In some examples, this can be achieved by creating an extended virtual bounding box with a power-of-two width and height that has the same upper-left corner coordinates as the original bounding box, but with a width / height that is a power of 2 and greater than or equal to the width / height of the original bounding box. The residual in the virtual bounding box, but not the residual in the original bounding box, is set to zero. Fig.16 The gray area in . This part of the transformation can be applied to a virtual bounding box.

[0267] In some examples, by restricting the combination of displacement and angle, the width and height values ​​can be achieved to be powers of 2, so that the bounding box always uses powers of 2 for the width and height. This partial transformation can be applied to the bounding box. When this restriction is applied (any size is a power of 2), the signaling of the angle is modified to exclude other unused angles. For example, only angles that result in a size that is a power of 2 for the bounding box can be signaled.

[0268] The following techniques can be applied to the boundary block based partial transform or the one proposed in the original GER contribution, or any other partial transform.

[0269] The size or number of samples involved in the partial transform may be limited, where the limit may be expressed as a fractional threshold applied to the CU width and / or height or CU area, for example. In one example, the fractional threshold may be set to 1 / 4 of the CU size.

[0270] In some examples, residual filtering can be applied if and only if the size of the partial transform is less than or equal to a certain fractional threshold of the CU size. In one example, the fractional threshold can be set to 1 / 4 of the CU size. In some examples, partial transform and residual propagation are applied only if the size of the partial transform is less than or equal to a fractional threshold of the CU size.

[0271] In some examples, residual filtering is applied only in a particular direction (e.g., only in the horizontal direction). This can allow a video codec (e.g., video encoder 200 or video decoder 300) to perform residual filtering "on the fly," e.g., right after performing an inverse transform and immediately while filling a residual buffer with the inverse transformed residual. In this case, residual filtering is applied without accessing any residual samples in the vertical direction and avoiding writing to the residual block in the vertical direction. To accomplish this, padding the residual block from the inverse transform output block is performed in some manner, e.g., only in the horizontal direction.

[0272] For some implementations, limiting the partial transform size and also limiting the transform size by applying residual filtering may be desirable, because if the limiting threshold is small enough, the partial transform and residual filtering may require the same processing time as doing a full transform. In such examples, no additional stages in the implementation are introduced.

[0273] Additionally, the signaling of a flag or indicator of whether a partial transform or a full transform is applied in a block may also be limited, taking into account the limitations of the partial transform. For example, when it is not possible to apply a partial transform, i.e., it is limited, the flag or indicator is not signaled, and a full transform is applied. Similarly, when the partial transform may not fully cover the edge area, i.e., there are samples in the edge area that are not transformed, the partial transform may be disabled, and the signaling of the use of the partial transform or the full transform is not performed, and a full transform is applied. This may happen, in which case the number of samples is not the same as the available size (width * height) of the transform block. In some cases, only transforms of lengths that are powers of 2 are generally supported.

[0274] In some examples, the partial transformation is only applied to GEOs with certain angles, such as angles in the set {0°, 90°, 180°, 270°, 360°}. That is, the partial transformation is applied only when the partition direction is horizontal or vertical.

[0275] The following describes the extension of the triangle merge mode angles. In some examples, the weights used in the TPM blending process can be extended to cover an additional 40 angles per CU, where the associated weight value for each angle can be represented by using the same derivation function as the end-to-end diagonal partitioning TPM mode (e.g., as described above with respect to triangle partitioning).

[0276] The mask-based extension is described below. The process of generating the weight values ​​of the random CU is the same as using some weight values ​​from the predefined weight value mask. Fig.17A and Fig. 17B As described above, the weight values ​​of a CU can be sampled anywhere from the underlying mask (called a hypothetical CU) corresponding to the desired angle (e.g., 45°). The underlying mask of the hypothetical CU can be changed according to the desired angle. According to the TPM design, the supported angles cover m*π+arctan(2 n ), where m∈{0,1 / 2,1,3 / 2} and n∈{-5,-4,…,4,5}. Figures 18A-18D Examples of additional supported angles are shown in .

[0277] The generation of the hypothetical CU weights is a two-step process: (a) taking weights from the regular TPM CU to fill a portion of the mask for the hypothetical CU, and (b) filling the remainder of the mask with weights of 0 or 8, depending on whether the remainder is spatially closer to the top or bottom triangulation. Fig.18A The mask of the hypothetical CU with arctan(128 / 64) depicted in FIG is formed as follows: First, it fills its upper half with the weight values ​​of the 128×64 TPM CU. Then, the rest of the hypothetical CU is filled by using the same weight values ​​assigned to the bottom triangle of the 128x64 TPM CU. In another example, Fig.18C The mask for the hypothetical CU with arctan(128 / 64) depicted in is formed as follows: first fill the left half using the weights of the 64x128 TPM CU; then fill the remainder of the hypothetical CU by using the same weight values ​​assigned to the bottom triangle of the 64x128 TPM CU.

[0278] The computation-based extension is described below. The weight values ​​for a random CU are generated by taking some weight values ​​from a larger hypothetical CU with a diagonal split edge. Given a CU of size (w)x(h), TPM already supports diagonal splits of angle arctan(w / h). This extension is intended to extend the range of TPM angles to cover Figures 18A-18D These angles can be expressed as m*π+arctan(s*w / h), where s is an integer and 2 -5 ≤arctan(s*w / h)≤2 5 .

[0279] For example, in Fig.19A In the example, the weight value of the (w)x(h) CU with an angle of arctan(2w / h) is the same as the left half weight value of the (2w)x(h) hypothetical CU. Fig.19B In the example, the weight value of the (w)x(h) CU with an angle of arctan(4w / h) is the same as the left half weight value of the (4w)x(h) hypothetical CU. Therefore, TPM can support more angles for each CU by using the same weight value derivation process as the larger CU (i.e., the hypothetical CU).

[0280] Similarly, when m is non-zero, the weight values ​​referenced from the imaginary CU may change. Assuming the size of the imaginary CU is (W)x(H):

[0281] a. When m=0 and W>H, the weight value of the left half of the imaginary CU is referenced;

[0282] b. When m = 0 and W < H, the weight value of the upper left half of the imaginary CU is referenced;

[0283] c When m = 1 / 2 and W > H, the weight value of the right half of the imaginary CU is referenced;

[0284] d. When m = 1 / 2 and W < H, the weight value of the upper half of the imaginary CU is referenced;

[0285] e. When m = 1 and W > H, the weight value of the right part of the imaginary CU is referenced;

[0286] f. When m = 1 and W < H, the weight value of the lower part of the imaginary CU is referenced;

[0287] g. When m = 3 / 2 and W > H, the weight value of the left part of the imaginary CU is referenced;

[0288] h. When m = 3 / 2 and W < H, the weight value of the lower part of the imaginary CU is referenced;

[0289] The two-dimensional offset (dx, dy) can be added to the coordinates of each CU corner to accommodate more weight options. Therefore, sampling the weight values from the imaginary CU does not need to start from any CU corner. For example, Fig. 20 in, the sampling position can be moved from the upper left CU corner to a certain position within the imaginary CU, and then the weight value is sampled at the position pointed to by the moving offset.

[0290] Fig.21 is a flowchart showing an example method for processing the current block. The video encoder 200 and the video decoder 300 can determine a first segmentation of the current block of video data to be encoded and decoded in a geometric partitioning mode and a second segmentation of the current block of the video data (2100). Examples of the current block can be Fig.15A and Fig. 15B blocks 1500 or 1506 of. Examples of the first segmentation include segmentation 1503A or 1509A, and examples of the second segmentation include segmentation 1503B or 1509B.

[0291] The video encoder 200 and the video decoder 300 can determine a first prediction block of the video data based on the first motion vector of the first segmentation and a second prediction block of the video data based on the second motion vector of the second segmentation (2102). For example, the video encoder 200 and the video decoder 300 can determine the first prediction block of segmentation 1503A or 1509A and the second prediction block of segmentation 1503B or 1509B based on the corresponding motion vectors.

[0292] The video encoder 200 and the video decoder 300 may mix the first prediction block and the second prediction block based on a weight indicating an amount by which samples in the first prediction block are scaled and an amount by which samples in the second prediction block are scaled to generate a final prediction block of the current block (2104). For example, Fig.15A and Fig. 15B An example of weights associated with samples in block 1500 or block 1506 is shown. The weights indicate the amount by which to scale the samples in the first prediction block and the co-located samples in the second prediction block. The video encoder 200 and the video decoder 300 may perform the mixing of the samples in the first prediction block and the second prediction block based on the weights (e.g., Fig.15A and 15B As shown) to generate the final prediction block.

[0293] The video encoder 200 and the video decoder 300 may divide the current block into a plurality of sub-blocks 2106. For example, the video encoder 200 and the video decoder 300 may divide the block 1500 into sub-blocks 1504A-1504P of 4×4 size, and divide the block 1506 into sub-blocks 1510A-1510P of 4×4 size.

[0294] The video encoder 200 and the video decoder 300 may determine a set of subblocks, each of which includes at least one sample corresponding to a predicted sample in a final prediction block generated based on equal weighting of samples in a first prediction block and samples in a second prediction block (2108). Fig.15A , the set of subblocks may include the third, fourth, fifth, sixth subblocks in the top row of block 1500 and the last subblock in block 1500 (eg, subblock 1504P) because each of these subblocks includes at least one sample associated with weight '4'.

[0295] The video encoder 200 may store a corresponding bidirectional prediction motion vector for each sub-block in the determined sub-block set (2110). For example, the video encoder 200 and the video decoder 300 may determine whether the first motion vector and the second motion vector are from different reference picture lists, and one of the following: based on the first motion vector and the second motion vector being from different reference picture lists, storing both the first motion vector and the second motion vector for each sub-block in the sub-block set, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and storing the selected one of the first motion vector or the second motion vector for each sub-block in the sub-block set.

[0296] In some examples, the set of subblocks may be a first set of subblocks. The video encoder 200 and the video decoder 300 may be configured to determine a second set of subblocks that does not include any samples corresponding to prediction samples in a final prediction block that is generated based on equal weighting of samples in the first prediction block and samples in the second prediction block. Examples of such subblocks include the first, second, third, and last subblocks in the top row of block 1500 and the first, second, third, fourth, fifth, sixth, and seventh subblocks in the bottom row of block 1500, and since none of these subblocks include samples associated with the weights of the video encoder 200, the video decoder 300 may determine for each subblock in the second set of subblocks whether a majority of the subblock is in the first partition or in the second partition, and for each subblock in the second set of subblocks, store a first motion vector based on the majority of the subblock being in the first partition, or store a second motion vector based on the majority of the subblock being in the second partition.

[0297] The video encoder 200 and the video decoder 300 may have various reasons to store motion vector information. As an example, the video encoder 200 and the video decoder 300 may construct a candidate list for merge mode or advanced motion vector prediction (AMVP) mode for subsequent blocks based on the stored corresponding bi-directionally predicted motion vectors.

[0298] In addition to storing motion vector information for the sub-block, the video encoder 200 and the video decoder 300 may encode or decode the current block. For example, the video encoder 200 may determine a residual block based on the difference between the current block and the final prediction block and signaling information indicating the residual block. The video decoder 300 may receive information of the residual block indicating the difference between the final prediction block and the current block, and reconstruct the current block based on the final prediction block and the residual block.

[0299] The following describes example techniques that may be used alone or in combination.

[0300] Example 1. A method for encoding and decoding video data, the method comprising: determining one or more angles from a set of angles for segmenting a current block using a geometric segmentation mode (GEO), wherein the set of angles from which the one or more angles are determined is the same as a set of angles that can be used for a triangular segmentation mode (TPM); segmenting the current block based on the determined one or more angles; and encoding and decoding the current block based on the segmentation of the current block.

[0301] Example 2. A method for encoding and decoding video data, the method comprising: determining one or more weights from a set of weights for blending a current block using a geometric partitioning mode (GEO), wherein the set of weights from which the one or more weights are determined is the same as a set of weights that can be used for a triangular partitioning mode (TPM); blending the current block based on the determined one or more weights; and encoding and decoding the current block based on the blending of the current block.

[0302] Example 3. A method for encoding and decoding video data, the method comprising: determining a motion field storage for using a geometric partitioning mode (GEO), wherein the motion field storage is the same as the motion field storage that can be used for a triangular partitioning mode (TPM); and encoding and decoding a current block based on the determined motion field storage.

[0303] Example 4. A method comprising any combination of Example 1-Example 3.

[0304] Example 5. A method according to any one or any combination of Examples 1 to 3, wherein encoding and decoding includes decoding.

[0305] Example 6. A method according to any one or any combination of Examples 1-3, wherein encoding and decoding includes encoding.

[0306] Example 7. A device for encoding and decoding video data, the device comprising a memory configured to store the video data and a processing circuit coupled to the memory, the processing circuit being configured to perform any one or any combination of Examples 1-3.

[0307] Example 8. The apparatus of Example 7, further comprising a display configured to display the decoded video data.

[0308] Example 9. The device of any one of Examples 7 and 9, wherein the device comprises one or more of: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0309] Example 10. The device of any one of Examples 7-9, wherein the device comprises a video decoder.

[0310] Example 11. A device according to any one of Examples 7-9, wherein the device includes a video encoder.

[0311] Example 12. A computer-readable storage medium having instructions stored thereon, which instructions, when executed, cause one or more processors to perform the method of any one or any combination of Examples 1-3.

[0312] Example 13. A device for encoding and decoding video data, the device comprising components for performing the method of any one or any combination of Examples 1-3.

[0313] It should be appreciated that, according to examples, certain actions or events of any technology described herein may be performed in a different sequence, may be added, merged, or omitted together (e.g., not all described actions or events are necessary for the practice of the technology). In addition, in some examples, actions or events may be performed concurrently rather than sequentially, such as by multithreading, interrupt handling, or multiple processors.

[0314] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media, or communication media, including, for example, any media that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) non-temporary tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures to implement the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0315] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory or any other medium that can be used to store the required program code in the form of instructions or data structures and can be accessed by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave) is used to send instructions from a website, server or other remote source, the definition of the medium includes coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave). However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals or other temporary media, but are directed to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser optical disks, optical optical disks, digital versatile disks (DVDs), floppy disks and blue-ray disks, wherein disks usually reproduce data magnetically, and optical disks reproduce data optically with lasers. The above combination should also be included in the scope of computer-readable media.

[0316] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used in this application may refer to any of the aforementioned structures or any other structure suitable for implementing the technology described in this application. In addition, in some aspects, the functions described in this application may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Similarly, the technology may be fully implemented in one or more circuits or logic elements.

[0317] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units, including one or more processors as described above in combination with appropriate software and / or firmware.

[0318] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for processing video data, the method comprising: Determining a first partition and a second partition of a current block of video data encoded and decoded in a geometric partitioning mode; determining a first prediction block of the video data based on a first motion vector of the first partition, and determining a second prediction block of the video data based on a second motion vector of the second partition; blending the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block of the current block; Dividing the current block into a plurality of sub-blocks; determining a set of subblocks, each subblock comprising at least one sample corresponding to a corresponding prediction sample in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; as well as storing a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks; The method further comprises storing a weight table of slopes of a segmentation line for segmenting the current block into the first segmentation and the second segmentation, and Wherein, determining the sub-block set comprises determining the sub-block set based on a stored table and an offset into the table.

2. The method according to claim 1, wherein: The sub-block set includes a first sub-block set, and the method further includes: determining a second set of subblocks, the second set of subblocks not including any samples corresponding to prediction samples in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; For each subblock in the second set of subblocks, determining whether a majority of the subblock is in the first partition or in the second partition; and For each subblock in the second set of subblocks, the first motion vector is stored based on the majority of the subblock being in the first partition or the second motion vector is stored based on the majority of the subblock being in the second partition.

3. The method according to claim 1, wherein: Storing the corresponding bidirectional prediction motion vector for each sub-block includes: determining whether the first motion vector and the second motion vector are from different reference picture lists; and One of the following: storing both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks based on that the first motion vector and the second motion vector are from different reference picture lists, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and The selected one of the first motion vector or the second motion vector is stored for each subblock in the set of subblocks.

4. The method according to claim 1, wherein: The size of each sub-block is 4×4.

5. The method according to claim 1, further comprising: A candidate list for merge mode or advanced motion vector prediction (AMVP) mode is constructed for subsequent blocks based on the stored corresponding bi-directional predictive motion vectors.

6. The method according to claim 1, further comprising: receiving information of a residual block indicating a difference between the final prediction block and the current block; as well as The current block is reconstructed based on the final prediction block and the residual block.

7. The method according to claim 1, further comprising: Determine a residual block based on a difference between the current block and the final prediction block; as well as The signaling indicates information of the residual block.

8. A device for processing video data, the device comprising: A memory configured to store the video data; as well as a processing circuit coupled to the memory and configured to: Determining a first partition and a second partition of a current block of video data encoded and decoded in a geometric partitioning mode; determining a first prediction block from the stored video data based on a first motion vector of the first partition, and determining a second prediction block from the stored video data based on a second motion vector of the second partition; blending the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block of the current block; Dividing the current block into a plurality of sub-blocks; determining a set of subblocks, each subblock comprising at least one sample corresponding to a corresponding prediction sample in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; as well as storing a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks; Wherein, the processing circuit is further configured to: storing a weight table for storing the slopes of the segmentation lines used to segment the current block into the first segmentation and the second segmentation, Therein, in order to determine the sub-block set, the processing circuit is configured to determine the sub-block set based on a stored table and an offset into the table.

9. The device according to claim 8, wherein: The sub-block set includes a first sub-block set, and the processing circuit is further configured to: determining a second set of subblocks, the second set of subblocks not including any samples corresponding to prediction samples in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; For each sub-block in the second set of sub-blocks, determining whether a majority of the sub-block is in the first partition or in the second partition; as well as For each subblock in the second set of subblocks, the first motion vector is stored based on the majority of the subblock being in the first partition or the second motion vector is stored based on the majority of the subblock being in the second partition.

10. The device according to claim 8, wherein: In order to store the corresponding bidirectional prediction motion vector for each sub-block, the processing circuit is configured to: determining whether the first motion vector and the second motion vector are from different reference picture lists; as well as One of the following: storing both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks based on that the first motion vector and the second motion vector are from different reference picture lists, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and The selected one of the first motion vector or the second motion vector is stored for each subblock in the set of subblocks.

11. The device according to claim 8, wherein: The processing circuit is configured to construct a candidate list for a merge mode or an advanced motion vector prediction (AMVP) mode for a subsequent block based on the stored corresponding bi-directional predictive motion vector.

12. The device according to claim 8, wherein: The processing circuit is configured to: receiving information of a residual block indicating a difference between the final prediction block and the current block; and The current block is reconstructed based on the final prediction block and the residual block.

13. The apparatus according to claim 8, wherein: The processing circuit is configured to: determining a residual block based on a difference between the current block and the final prediction block; and The signaling indicates information of the residual block.

14. The apparatus according to claim 8, wherein: The device includes one or more of: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

15. A computer-readable storage medium having stored thereon instructions which, when executed, cause one or more processors of a device for processing video data to: Determining a first partition and a second partition of a current block of video data encoded and decoded in a geometric partitioning mode; determining a first prediction block of the video data based on a first motion vector of the first partition, and determining a second prediction block of the video data based on a second motion vector of the second partition; blending the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block of the current block; Dividing the current block into a plurality of sub-blocks; determining a set of subblocks, each subblock comprising at least one sample corresponding to a corresponding prediction sample in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; as well as storing a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks; The computer-readable storage medium further comprises instructions, wherein the instructions cause the one or more processors to: storing a weight table for storing the slopes of the segmentation lines used to segment the current block into the first segmentation and the second segmentation, The instructions causing the one or more processors to determine the set of sub-blocks include instructions causing the one or more processors to determine the set of sub-blocks based on a stored table and an offset into the table.

16. The computer-readable storage medium of claim 15, wherein: The sub-block set includes a first sub-block set, and the instructions further include instructions for causing the one or more processors to perform the following operations: determining a second set of subblocks, the second set of subblocks not including any samples corresponding to prediction samples in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; For each sub-block in the second set of sub-blocks, determining whether a majority of the sub-block is in the first partition or in the second partition; as well as For each subblock in the second set of subblocks, the first motion vector is stored based on the majority of the subblock being in the first partition or the second motion vector is stored based on the majority of the subblock being in the second partition.

17. The computer-readable storage medium of claim 15, wherein: The instructions causing the one or more processors to store a corresponding bidirectional predictive motion vector for each sub-block include instructions causing the one or more processors to: determining whether the first motion vector and the second motion vector are from different reference picture lists; and One of the following: storing both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks based on that the first motion vector and the second motion vector are from different reference picture lists, or selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and The selected one of the first motion vector or the second motion vector is stored for each subblock in the set of subblocks.

18. The computer-readable storage medium of claim 15, wherein: The size of each sub-block is 4×4.

19. A device for processing video data, the device comprising: means for determining a first partition and a second partition of a current block of video data encoded and decoded in a geometric partitioning mode; means for determining a first prediction block of the video data based on a first motion vector of the first partition, and determining a second prediction block of the video data based on a second motion vector of the second partition; means for blending the first prediction block and the second prediction block based on a weight indicating an amount to scale samples in the first prediction block and an amount to scale samples in the second prediction block to generate a final prediction block for the current block; A means for dividing the current block into a plurality of sub-blocks; means for determining a set of subblocks, each subblock comprising at least one sample corresponding to a corresponding prediction sample in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; as well as means for storing a corresponding bidirectional prediction motion vector for each sub-block in the determined set of sub-blocks; Wherein, the device also includes: means for storing a weight table of slopes of a segmentation line for segmenting the current block into the first segmentation and the second segmentation, Wherein, the means for determining the set of sub-blocks comprises means for determining the set of sub-blocks based on a stored table and an offset into the table.

20. The apparatus of claim 19, wherein: The sub-block set includes a first sub-block set, and the device further includes: means for determining a second set of subblocks, the second set of subblocks not including any samples corresponding to prediction samples in the final prediction block, the final prediction block being generated based on equal weighting of samples in the first prediction block and samples in the second prediction block; means for determining, for each sub-block in the second set of sub-blocks, whether a majority of the sub-block is in the first partition or in the second partition; and Means for storing, for each subblock in the second set of subblocks, the first motion vector based on the majority of the subblocks being in the first partition or storing the second motion vector based on the majority of the subblocks being in the second partition.

21. The apparatus of claim 19, wherein: The means for storing a corresponding bidirectional prediction motion vector for each sub-block comprises: means for determining whether the first motion vector and the second motion vector are from different reference picture lists; and means for storing both the first motion vector and the second motion vector for each sub-block in the set of sub-blocks based on the first motion vector and the second motion vector being from different reference picture lists, or means for selecting one of the first motion vector or the second motion vector based on the first motion vector and the second motion vector being from the same reference picture list, and storing the selected one of the first motion vector or the second motion vector for each subblock in the set of subblocks.