VIDEO DECODING METHOD, VIDEO DECODING APPARATUS, COMPUTER PROGRAM, AND VIDEO ENCODING METHOD
The video encoding method employs a triangular prediction mode to split coding blocks into triangular units, determining merge indexes to efficiently encode motion vectors, thereby improving compression efficiency and reducing storage requirements.
Patent Information
- Application Number
- JP2023191919
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2039-12-04
AI Technical Summary
Existing video encoding technologies face challenges in efficiently reducing redundancy in video signals, particularly in encoding motion vectors, which affects compression ratio and storage requirements.
The method involves using a triangular prediction mode for video encoding, where a coding block is split into two triangular prediction units based on a split direction indicated by a syntax element. Merge indexes for each unit are determined from index syntax elements, allowing for efficient reconstruction of the coding block.
This approach enhances compression efficiency by reducing the data required to encode motion vectors, thereby improving the compression ratio and reducing storage needs while maintaining video quality.
Smart Images

Figure 0007682977000003 
Figure 0007682977000004 
Figure 0007682977000005
Abstract
Description
[Background technology]
[0001] This disclosure claims priority to U.S. Application No. 16 / 425,404, entitled "Method and Apparatus for Video Encoding," filed May 29, 2019, which claims priority to U.S. Provisional Application No. 62 / 778,832, entitled "Signaling and Derivation for Triangular Prediction Parameters," filed December 12, 2018.
[0002] Technical Field This disclosure describes embodiments generally relating to video encoding.
[0003] background The background discussion provided in this application is intended to generally present the context of the present disclosure. The work of the presently identified inventors is not admitted, expressly or impliedly, as prior art to the present disclosure, to the extent that that work is described in this background section and in a manner of description that may not have been granted prior art status at the time of filing.
[0004] Video encoding and decoding can be performed using inter-picture prediction along with motion compensation. Uncompressed digital video can include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luminance samples and associated chrominance samples. The sequence of pictures can have a fixed or variable picture rate (informally known as the frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.
[0005] One of the goals of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce the bandwidth or storage space requirements, in some cases by more than one order of magnitude. Both lossless and non-lossless compression, as well as a combination of both, can be used. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using non-lossless compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. For video, non-lossless compression is widely used. The amount of distortion that is tolerated depends on the application, for example, a user of a consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio can reflect that a higher tolerable / bearable distortion can result in a higher compression ratio.
[0006] Motion compensation can be a non-lossless compression technique and can also refer to a technique in which blocks of sample data from a previously reconstructed picture or part of it (reference picture) are used to predict a newly reconstructed picture or part of a picture, after a spatial shift in the direction indicated by a motion vector (hereafter MV). In some cases, the reference picture can be the same as the picture currently being reconstructed. MV can have two dimensions X and Y, or three dimensions, the third being an indication of the reference picture in use (the latter can indirectly be a temporal dimension).
[0007] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, e.g., from those associated with another area of sample data that is spatially adjacent to the area being reconstructed and that precedes the MV in decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thereby removing redundancy and improving compression. MV prediction can work effectively, for example, when encoding an input video signal derived from a camera (known as natural video), because areas larger than the area to which a single MV is applicable are statistically likely to move in similar directions and can therefore, in some cases, be predicted using similar motion vectors derived from the MVs of neighboring areas. As a result, the MV found for a given area is similar or identical to the MV predicted from the surrounding MVs, which can be represented, after entropy coding, with fewer bits than would be used if the MV were encoded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself may not be lossless, for example due to rounding errors when computing the predictor from several surrounding MVs.
[0008] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, a technique called "spatial merging" is described in the following of this application.
[0009] Referring to Figure 1, a current block (101) contains samples that the encoder found during the motion search process to be predictable from a spatially shifted previous block of the same size. Instead of directly encoding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the most recent reference picture (in decoding order) using MVs associated with any of the five surrounding samples denoted A0, A1, and B0, B1, B2 (102 to 106, respectively). In H.265, MV prediction can use predictors from the same reference picture as the neighboring blocks use. Summary of the Invention
[0010] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit can be configured to receive a split direction syntax element, a first index syntax element, and a second index syntax element associated with a coding block of a picture. The coding block can be coded by a triangular prediction mode. The coding block can be partitioned into a first triangular prediction unit and a second triangular prediction unit according to a split direction indicated by the split direction syntax element. The first and second index syntax elements can indicate a first merge index and a second merge index for merge candidate lists configured for the first and second triangular prediction units, respectively. The split direction, the first merge index, and the second merge index can be determined based on the split direction syntax element, the first index syntax element, and the second index syntax element. The coding block can be reconstructed according to the determined split direction, the determined first merge index, and the determined second merge index.
[0011] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a video decoding method. [Brief description of the drawings]
[0012] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0013] [Figure 1] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merging candidates in one example.
[0014] [Diagram 2] FIG. 2 is a schematic diagram of a simplified block diagram of a communication system (200) according to an embodiment.
[0015] [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to an embodiment.
[0016] [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment;
[0017] [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment;
[0018] [Figure 6] 4 shows a block diagram of an encoder according to another embodiment;
[0019] [Figure 7] 4 shows a block diagram of a decoder according to another embodiment;
[0020] [Figure 8] 1 illustrates examples of candidate locations for constructing a merge candidate list according to an embodiment.
[0021] [Figure 9] 1 illustrates an example of partitioning a coding unit into two triangular prediction units according to an embodiment.
[0022] [Figure 10] 1 illustrates an example of spatial and temporal neighboring blocks used to construct a merge candidate list according to an embodiment.
[0023] [Figure 11] 13 illustrates an example lookup table used to derive partition direction and partition motion information based on triangle partition index according to an embodiment.
[0024] [Figure 12] 1 illustrates an example of a coding unit that applies a set of weighting factors in an adaptive blending process according to an embodiment.
[0025] [Figure 13] 13 illustrates an example of motion vector storage in triangular prediction mode according to an embodiment.
[0026] [Figure 14A] 1 illustrates an example of deriving a bi-directional predictive motion vector based on the motion vectors of two triangular prediction units according to an embodiment. [Figure 14B] 1 illustrates an example of deriving a bi-directional predictive motion vector based on the motion vectors of two triangular prediction units according to an embodiment. [Figure 14C] 1 illustrates an example of deriving a bi-directional predictive motion vector based on the motion vectors of two triangular prediction units according to an embodiment. [Figure 14D] 1 illustrates an example of deriving a bi-directional predictive motion vector based on the motion vectors of two triangular prediction units according to an embodiment.
[0027] [Figure 15] 1 illustrates an example of a triangular prediction process according to some embodiments.
[0028] [Figure 16] 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0029] I. Video Encoding Encoders and Decoders FIG. 2 illustrates a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via the network (250). In the example of FIG. 2, the first pair of terminal devices (210) and (220) perform a unidirectional transmission of data. For example, the terminal device (210) may encode video data (e.g., a stream of video pictures captured by the terminal device (210)) for transmission to the other terminal device (220) via the network (250). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (220) can receive the encoded video data from the network (250), decode the encoded video data to reconstruct a video picture, and display the video picture according to the reconstructed video data. One-way data transmission may be common in media serving applications, etc.
[0030] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) for bidirectional transmission of encoded video data, such as may occur during a video conference. For bidirectional transmission of data, for example, each of the terminal devices (230) and (240) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (230) and (240) via the network (250) of the terminal devices (230) and (240). Each of the terminal devices (230) and (240) may receive encoded video data transmitted by the other of the terminal devices (230) and (240), decode the encoded video data to reconstruct the video pictures, and display the video pictures on an accessible display device in accordance with the reconstructed video data.
[0031] In the example of FIG. 2, the terminal devices (210), (220), (230), (240) may be described as servers, personal computers, and smartphones, although the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find application to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks that carry encoded video data between the terminal devices (210), (220), (230), and (240), including, for example, wired (wired) and / or wireless communication networks. The communication network (250) may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of the present disclosure, the architecture and topology of the network (250) may not be important to the operation of the present disclosure, unless otherwise described below.
[0032] 3 illustrates the placement of a video encoder and video decoder in a streaming environment as an example application of the disclosed subject matter. The disclosed subject matter can be applied to other video-enabled applications as well, including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0033] The streaming system may include a capture subsystem (313), which may include a video source (301), e.g., a digital camera, for generating a stream of, e.g., uncompressed video pictures (302). In one example, the stream of video pictures (302) includes samples taken by the digital camera. The stream of video pictures (302), depicted as a thick line emphasizing the amount of data when compared to the encoded video data (304) (or encoded video bitstream), may be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof, which may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (304) (or encoded video bitstream (304)), depicted as thin lines to emphasize the smaller amount of data compared to the stream of video pictures (302), can be stored on the streaming server (305) for future use. One or more streaming client subsystems, such as the client subsystems (306) and (308) of FIG. 3, can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) can include a video decoder (310), for example within an electronic device (330). The video decoder (310) decodes an input copy of the encoded video data (307) and generates an output stream of video pictures (311) that can be rendered on a display (312) (e.g., a display screen) or other rendering device (not shown).In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) may be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. For example, one evolving video encoding standard is informally known as Universal Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0034] It should be noted that the electronic devices (320) and (330) may include other components (not shown). For example, the electronic device (320) may include a video decoder (not shown), and the electronic device (330) may also include a video encoder (not shown).
[0035] 4 shows a block diagram of a video decoder (410) according to an embodiment of the present disclosure. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., receiving circuitry). The video decoder (410) may be used in place of the video decoder (310) in the example of FIG. 3.
[0036] The receiver (431) may receive one or more coded video sequences to be decoded by the video decoder (410), one coded video sequence at a time, in the same or another embodiment, and the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from a channel (401), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (431) may receive the coded video data together with other data, such as coded audio data and / or auxiliary data streams, which may be transferred using respective entities (not shown). The receiver (431) may separate the coded video sequences from the other data. To address network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420), hereafter referred to as the parser (420). In certain applications, the buffer memory (415) is part of the video decoder (410). In other cases, it can be external to the video decoder (410) (not shown). In yet other cases, there can be a buffer memory (not shown) external to the video decoder (410), for example to deal with network jitter, and another buffer memory (415) internal to the video decoder (410), for example to handle playback timing. If the receiver (431) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from a synchronous network, the buffer memory (415) may not be necessary or can be small.For use in a best effort packet network such as the Internet, the buffer memory (415) may be as required and may be relatively large, and may advantageously be of an adaptive size, and may be implemented at least in part within an operating system or similar element (not shown) outside the video decoder (410).
[0037] The video decoder (410) may include a parser (420) to reconstruct symbols (421) from the encoded video sequence. These symbol categories include information used to manage the operation of the video decoder (410) and potential information for controlling a rendering device such as a rendering device (412) (e.g., a display screen) that is not an integral part of the electronic device (430) but may be coupled to the electronic device (430) as shown in FIG. 4. The rendering device control information may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (420) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, context-dependent or not arithmetic coding, etc. The parser (420) can extract, at the video decoder, a set of subgroup parameters for at least one of the subgroups of pixels from the coded video sequence based on the at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (420) can also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0038] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (415) to generate symbols (421).
[0039] The reconstruction of the symbols (421) may involve several different units depending on the type of coded video picture or part thereof (inter picture, intra picture, inter block, intra block) and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (420). The flow of such subgroup control information between the parser (420) and the following units is not depicted for the sake of simplicity.
[0040] Beyond the functional blocks already mentioned, the video decoder (410) may be conceptually divided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into multiple functional units is reasonable.
[0041] The first unit is a scalar / inverse transform unit (451), which receives quantized transform coefficients and control information as symbols (421) from the parser (420), including the transform to use, block size, quantization factor, quantization scaling matrix, etc. The scalar / inverse transform unit (451) can output a block containing sample values that can be input to an aggregator (455).
[0042] In some cases, the output samples of the scalar / inverse transform (451) may be associated with intra-coded blocks, i.e. blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed part of the current image. Such prediction information may be provided by an intra picture prediction unit (452). In some cases, the intra picture prediction unit (452) uses already reconstructed surrounding information retrieved from a current picture buffer (458) to generate blocks of the same size and shape of the block being reconstructed. The current picture buffer (458) may, for example, buffer a partially reconstructed and / or a fully reconstructed current picture. The aggregator (455) may add the prediction information generated by the intra prediction unit (452) to the output sample information as provided by the scalar / inverse transform unit (451) on a sample-by-sample basis.
[0043] In other cases, the output samples of the scalar / inverse transform unit (451) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (453) may access the reference picture memory (457) to retrieve samples used for prediction. After motion compensating the retrieved samples according to the symbols (421) associated with the block, these samples may be added by the aggregator (455) to the output of the scalar / inverse transform unit (451) (in this case referred to as residual samples or residual signal) to generate output sample information. The addresses in the reference picture memory (457) from which the motion compensated prediction unit (453) retrieves the prediction samples may be controlled by a motion vector, which is available to the motion compensated prediction unit (453), for example in the form of a symbol (421) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values taken from a reference picture memory (457), motion vector prediction mechanisms, etc., where sub-sample accurate motion vectors are used.
[0044] The output samples of the aggregator (455) can be subjected to various loop filtering techniques in a loop filter unit (456). Video compression techniques can include in-loop filter techniques, which are controlled by parameters contained in the coded video sequence (also called coded video bitstream) and made available to the loop filter unit (456) as symbols (421) from the parser (420), but which can also be responsive to previously reconstructed loop filtered sample values in addition to being responsive to meta information obtained during the decoding of previous parts of the coded picture or coded video sequence (in decoding order).
[0045] The output of the loop filter unit (456) may be a sample stream that may be stored in a reference picture memory (457) for use in future inter-picture prediction, in addition to being output to a rendering device (412).
[0046] Once a particular coded picture has been fully reconstructed, it can be used as a reference picture for future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified (e.g., by the parser (420)) as a reference picture, the current picture buffer (458) can become part of the reference picture memory (457) and a new current picture buffer can be reassigned before starting the reconstruction of the next coded picture.
[0047] The video decoder (410) may perform decoding operations according to a given video compression technique in a standard such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technique or standard being used in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile as documented in the video compression technique or standard. In particular, a profile may select a particular tool from all tools available in the video compression technique or standard as the only tool available for use under that profile. Also, a requirement for compliance may be that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further restricted through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management, possibly signaled in the coded video sequence.
[0048] In an embodiment, the receiver (431) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0049] 5 shows a block diagram of a video encoder (503) according to an embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) can be used in place of the video encoder (303) in the example of FIG. 3.
[0050] The video encoder (503) can receive video samples from a video source (501) (which in the example of FIG. 5 is not part of the electronic device (520)) capable of capturing video images to be encoded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).
[0051] The video source (501) may provide a source video sequence to be encoded by the video encoder (503) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (501) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (501) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that convey motion when viewed in succession. The picture itself may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0052] According to an embodiment, the video encoder (503) can encode and compress pictures of a source video sequence into an encoded video sequence (543) in real-time or under any other time constraint required by the application. Imposing an appropriate encoding rate is one of the functions of the controller (550). In some embodiments, the controller (550) controls and is operatively coupled to other functional units, as described below, whose coupling is not depicted for simplicity. Parameters set by the controller (550) can include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures layout, maximum motion vector search range, etc. The controller (550) can be configured to have other suitable functions associated with the video encoder (503) optimized for a particular system design.
[0053] In some implementations, the video encoder (503) is configured to operate in an encoding loop. As an oversimplified explanation, in one example, the encoding loop can include a source encoder (530) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs the symbols to generate sample data in a similar manner that the (remote) decoder also generates (such that any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (534). Since the decoding of the symbol stream produces bit-perfect results independent of the location of the decoder (local or remote), the contents in the reference picture memory (534) are also bit-perfect between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values as the decoder would "see" if it were to use the prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift when synchronization cannot be maintained, e.g. due to channel errors) is used in several related techniques as well.
[0054] The operation of the "local" decoder (533) may be the same as a "remote" decoder, such as the video decoder (410) described in detail above in connection with Figure 4. However, with brief reference also to Figure 4, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy encoder (545) and parser (420) may be lossless, the entropy decoding portion of the video decoder (410), including the buffer memory (415) and parser (420), may not be fully implemented in the local decoder (533).
[0055] An insight that can be made at this point is that any decoder techniques other than parser processing / entropy decoding that are present in the decoder must also be present in the corresponding encoder, in substantially the same functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques that are described generically. Only in certain areas are more detailed descriptions required, and are provided below.
[0056] In some examples, during operation, the source encoder (530) can perform motion-compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence, designated as “reference pictures.” In this manner, the encoding engine (532) encodes differences between pixel blocks of the input picture and pixel blocks of the reference pictures, which can be selected as prediction references for the input picture.
[0057] The local video decoder (533) can decode the coded video data of the pictures that can be designated as reference pictures based on the symbols generated by the source encoder (530). The operation of the coding engine (532) can advantageously be a non-lossless process. If the coded video data can be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence can be a replica of the source video sequence, typically with some errors. The local video decoder (533) can repeat the decoding process performed by the video decoder with respect to the reference pictures and store the reconstructed reference pictures in the reference picture cache (534). In this way, the video encoder (503) can locally store copies of reconstructed reference pictures that have a common content with the reconstructed reference pictures obtained by the far-end video decoder (in the absence of transmission errors).
[0058] The predictor (535) can perform a prediction search for the coding engine (532). That is, for a new picture to be coded, the predictor (535) can search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (535) can operate on a sample block pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (535), the input picture can have prediction references drawn from multiple reference pictures stored in the reference picture memory (534).
[0059] The controller (550) can manage the encoding operations of the source encoder (530), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0060] The output of all the aforementioned functional units may be subjected to entropy coding in an entropy coder (545), which converts the symbols as produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0061] The transmitter (540) may buffer the coded video sequence created by the entropy coder (545) and prepare it for transmission over a communication channel (560), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (540) may combine the coded video data from the video coder (503) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0062] The controller (550) can manage the operation of the video encoder (503). During encoding, the controller (550) can assign a particular coding picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, pictures are often designated as one of the following picture types:
[0063] An intra picture (I-picture) is a possible picture that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.
[0064] A predictive picture (P-picture) can be one that can be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values of each block.
[0065] Bidirectionally predicted pictures (B-pictures) may be those that can be coded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a block.
[0066] A source picture is usually spatially divided into several sample blocks (e.g., blocks of 4x4, 8x8, 4x8, 16x16 samples, respectively) and can be coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture of the block. For example, blocks of I pictures can be non-predictively coded or they can be predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of P pictures can be predictively coded with spatial or temporal prediction with reference to one previously coded reference picture. Blocks of B pictures can be predictively coded with spatial or temporal prediction with reference to one or two previously coded reference pictures.
[0067] The video encoder (503) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (503) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.
[0068] In an embodiment, the transmitter (540) may transmit additional data along with the encoded video. The source encoder (530) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0069] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as "intra prediction") exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture under encoding / decoding, referred to as the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and can have a third dimension that identifies the reference picture if multiple reference pictures are in use.
[0070] In some embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture, both preceding in decoding order (but potentially past and future, respectively, in display order) a current picture in a video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first and second reference blocks.
[0071] Furthermore, to improve coding efficiency, it is possible to use merge mode techniques in inter-picture prediction.
[0072] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. In general, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in encoding / decoding are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luminance values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0073] 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processed block of sample values (e.g., a predictive block) in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (603) is used in place of the video encoder (303) in the example of FIG. 3.
[0074] In an HEVC example, a video encoder (603) receives a matrix of sample values for a processing block, such as a prediction block of 8 by 8 samples. The video encoder (603) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-directional prediction mode, e.g., using rate-distortion optimization. If the processing block is to be coded in an intra mode, the video encoder (603) may use intra prediction techniques to code the processing block into a coded picture; if the processing block is to be coded in an inter mode or a bi-directional prediction mode, the video encoder (603) may use inter prediction techniques or bi-directional prediction techniques, respectively, to code the processing block into a coded picture. In certain video encoding techniques, the merge mode can be an inter picture prediction sub-mode, where motion vectors are derived from one or more motion vector predictors without benefit of motion vector components coded outside the predictors. In certain other video encoding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.
[0075] In the example of FIG. 6, the video encoder (603) includes an inter-encoder (630), an intra-encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general controller (621), and an entropy encoder (625), which are coupled together as shown in FIG. 6.
[0076] The inter encoder (630) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks of previous and subsequent pictures), generate inter prediction information (e.g., a description of redundant information due to inter coding techniques, motion vectors, merge mode information), and calculate an inter prediction result (e.g., a prediction block) using any suitable technique based on the inter prediction information. In some examples, the reference picture is a decoded reference picture that is decoded based on the coded video information.
[0077] The intra encoder (622) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block to previously encoded blocks in the same picture, and generate transformed and quantized coefficients and optionally intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (622) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and reference blocks in the same picture.
[0078] The general controller (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In one example, the general controller (621) determines a mode of the block and provides a control signal to the switch (626) based on the mode. For example, if the mode is an intra mode, the general controller (621) controls the switch (626) to select an intra mode result for use by the residual calculator (623) and controls the entropy encoder (625) to select intra prediction information and include the intra prediction information in the bitstream; and if the mode is an inter mode, the general controller (621) controls the switch (626) to select an inter prediction result for use by the residual calculator (623) and controls the entropy encoder (625) to select inter prediction information and include the inter prediction information in the bitstream.
[0079] The residual calculator (623) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (622) or the inter-encoder (630). The residual encoder (624) is configured to operate on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (624) is configured to transform the residual data from a spatial domain to a frequency domain to generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used by the intra-encoder (622) and the inter-encoder (630) as appropriate. For example, the inter-encoder (630) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (622) may generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are appropriately processed to generate decoded pictures, which may be buffered in a memory circuit (not shown) and, in some examples, used as reference pictures.
[0080] The entropy encoder (625) is configured to format a bitstream to include the encoded block. The entropy encoder (625) is configured to include various information in accordance with an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, in accordance with the disclosed subject matter, when encoding a block in a merged sub-mode of either the inter mode or the bi-prediction mode, the residual information is not present.
[0081] 7 shows a diagram of a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive coded pictures that are part of a coded video sequence and to decode the coded pictures to generate reconstructed pictures. In an embodiment, the video decoder (710) is used in place of the example video decoder (310) of FIG. 3.
[0082] In the example of FIG. 7, the video decoder (710) includes an entropy decoder (771), an inter decoder (780), a residual decoder (773), a reconstruction module (774), and an intra decoder (772), which are coupled together as shown in FIG. 7.
[0083] The entropy decoder (771) may be configured to reconstruct from the coded picture certain symbols that represent syntax elements that make up the coded picture. Such symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predictive mode, the latter two in merged or separate submodes), prediction information (e.g., intra prediction information or inter prediction information) that may identify certain samples or metadata used for prediction by the intra decoder (772) or the inter decoder (780), respectively, residual information (e.g., in the form of quantized transform coefficients), etc. As an example, if the prediction mode is an inter or bi-predictive mode, the inter prediction information is provided to the inter decoder (780); if the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (772). The residual information may undergo inverse quantization and is provided to the residual decoder (773).
[0084] The inter decoder (780) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.
[0085] The intra decoder (772) is configured to receive intra prediction information and to generate intra prediction results based on the intra prediction information.
[0086] The residual decoder (773) is configured to perform inverse quantization to extract unquantized transform coefficients, and process the unquantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (771) (datapath not depicted as this may only be a small volume of control information).
[0087] The reconstruction module (774) is configured to combine, in the spatial domain, the residual as output by the residual decoder (773) and the prediction result (as output by the optional inter or intra prediction module) to form a reconstructed block that may become part of the reconstructed picture, which may then become part of the reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve the visual quality.
[0088] It should be noted that the video encoders (303), (503), (603) and the video decoders (310), (410), (710) may be implemented using any suitable technology. In some embodiments, the video encoders (303), (503), (603) and the video decoders (310), (410), (710) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (303), (503), (503) and the video decoders (310), (410), (710) may be implemented using one or more processors executing software instructions.
[0089] II. Triangular Prediction 1. Encoding in merge mode A picture can be partitioned into a number of blocks, for example using a partitioning scheme based on a tree structure. The resulting blocks can be processed in various processing modes, such as intra prediction mode, inter prediction mode (e.g. merge mode, skip mode, advanced motion vector prediction (AVMP) mode). If a currently processed block, referred to as a current block, is processed in a merge mode, neighboring blocks can be selected from the spatial or temporal neighborhood of the current block. The current block can be merged with selected neighboring blocks by sharing the same motion data set (or called motion information) from the selected neighboring blocks. This merge mode operation can be performed on a group of neighboring blocks, such that regions of the neighboring blocks can be merged together and share the same motion data set. During transmission from the encoder to the decoder, instead of transmitting the entire set of motion data, an index indicating the motion data of the selected neighboring blocks can be transmitted for the current block. In this way, the amount of data (bits) used for transmitting the motion information can be reduced, and the coding efficiency can be improved.
[0090] In the above example, the neighboring blocks providing the motion data can be selected from a set of candidate positions. The candidate positions can be predefined relative to the current block. For example, the candidate positions can include spatial candidate positions and temporal candidate positions. Each spatial candidate position is associated with a spatial neighboring block in the vicinity of the current block. Each temporal candidate position is associated with a temporal neighboring block located in another coded picture (e.g., a previously coded picture). The neighboring blocks (called candidate blocks) that the candidate positions overlap are a subset of the spatial or temporal neighboring blocks of the current block. In this way, the candidate blocks can be evaluated for the selection of the block to be merged instead of the entire set of neighboring blocks.
[0091] FIG. 8 shows an example of candidate positions. From these candidate positions, a set of merge candidates can be selected to build a merge candidate list. For example, the candidate positions defined in FIG. 8 can be used in the HEVC standard. As shown, a current block (810) is to be processed in merge mode. A set of candidate positions {A1, B1, B0, A0, B2, C0, C1} is defined for merge mode processing. Specifically, the candidate positions {A1, B1, B0, A0, B2} are spatial candidate positions that represent the positions of candidate blocks in the same picture as the current block (810). In contrast, the candidate positions {C0, C1} are temporal candidate positions that represent the positions of candidate blocks in another neighboring coded picture or that overlap with a co-located block of the current block (810). As shown, the candidate position C1 can be located near (e.g., adjacent to) the center of the current block (810).
[0092] The candidate positions can be represented by a block of samples or samples in various examples. In FIG. 8, each candidate position is represented by a block of samples having a size of, for example, 4×4 samples. The size of such a block of samples corresponding to a candidate position is equal to or smaller than the minimum allowed size of a PB (for example, 4×4 samples) defined for the tree-based partitioning scheme used to generate the current block (810). In such a configuration, a block corresponding to a candidate position can always be covered within a single neighboring PB. In another example, a sample position (for example, the bottom right sample in block A1, or the top right sample in block A0) can be used to represent the candidate position. Such a sample is called a representative sample, while such a position is called a representative position.
[0093] In one example, based on the candidate positions {A1, B1, B0, A0, B2, C0, C1} defined in Figure 8, a merge mode process may be performed to select merge candidates from the candidate positions {A1, B1, B0, A0, B2, C0, C1} to construct a candidate list. The candidate list may have a predetermined maximum number of merge candidates Cm. Each merge candidate in the candidate list may include a set of motion data that may be used for motion compensated prediction.
[0094] The merge candidates may be listed in the candidate list according to a particular order. For example, depending on how the merge candidates are derived, different merge candidates may have different probabilities of being selected. Merge candidates with a higher probability of being selected are placed before merge candidates with a lower probability of being selected. Based on such an order, each merge candidate is associated with an index (called a merge index). In one embodiment, merge candidates with a higher probability of being selected will have a smaller index value, and as a result, fewer bits are required to encode each f.
[0095] In one example, the motion data of a merge candidate may include horizontal and vertical motion vector displacement values of one or more motion vectors, one or two picture indices associated with the one or two motion vectors, and optionally, an identifier of a reference picture list associated with each index.
[0096] In one example, according to a predetermined order, a first number of merge candidates Ca is derived from spatial candidate positions according to the order {A1, B1, B0, A0, B2}, and a second number of merge candidates Cb=Cm-Ca is derived from temporal candidate positions in the order {C0, C1}. Reference numbers A1, B1, B0, A0, B2, C0, C1 for representing candidate positions can also be used to refer to merge candidates. For example, a merge candidate obtained from candidate position A1 is called merge candidate A1.
[0097] In some scenarios, a merge candidate may not be available at a candidate position. For example, a candidate block at a candidate position may be intra predicted outside the slice or tile containing the current block (810) or in a row of the same coding tree block (CTB) as the current block (810). In some scenarios, a merge candidate at a candidate position may be redundant. For example, one neighboring block of the current block (810) may overlap two candidate positions. A redundant merge candidate may be removed from the candidate list (e.g., by performing a pruning process). If the total number of available merge candidates in the candidate list (after the redundant candidates have been removed) is less than the maximum number of merge candidates Cm, additional merge candidates may be generated (according to a pre-configured rule) to fill the candidate list, so that the candidate list may be maintained to have a fixed length. For example, the additional merge candidates may include a combined bi-prediction candidate and a zero motion vector candidate.
[0098] After the candidate list is created, an evaluation process can be performed in the encoder to select a merging candidate from the candidate list. For example, a rate-distortion (RD) performance corresponding to each merging candidate can be calculated and the one with the best RD performance can be selected. Then, a merging index associated with the selected merging candidate is determined for the current block (810) and signaled to the decoder.
[0099] At the decoder, a merge index for the current block (810) may be received. A similar candidate list construction process as described above may be performed to generate a candidate list that is the same as the candidate list generated at the encoder side. After the candidate list is constructed, in some instances, a merge candidate may be selected from the candidate list based on the received merge index without any further evaluation. The motion data of the selected merge candidate may be used for subsequent motion compensated prediction of the current block (810).
[0100] In some examples, a skip mode is also introduced. For example, in the skip mode, the current block can be predicted using the merge mode as described above to determine a set of motion data, but no residual is generated and no transform coefficients are signaled. A skip flag can be associated with the current block. The skip flag and the merge index, which indicate the relevant motion information of the current block, can be signaled to the video decoder. For example, at the beginning of a CU in an inter-picture prediction slice, the skip flag can signal the following: the CU contains only one PU (2Nx2N); the merge mode is used to derive the motion data; and there is no residual data in the bitstream. At the decoder side, based on the skip flag, a prediction block can be determined based on the merge index for decoding the respective current block without adding residual information. Thus, the various methods for video coding using the merge mode disclosed herein can be utilized in combination with the skip mode.
[0101] 2. Triangle prediction mode In some embodiments, it is possible to use triangular prediction mode for inter prediction. In an embodiment, the triangular prediction mode is applied to CUs with a size of 8×8 samples or more and coded in skip or merge mode. In an embodiment, a CU level flag is signaled to indicate whether the triangular prediction mode is applied to CUs that meet these conditions (sample size of 8×8 or more and coded in skip or merge mode).
[0102] When a triangular prediction mode is used, in some implementations, the CU is evenly divided into two triangular partitions using either diagonal or anti-diagonal partitioning, as shown in FIG. 9. In FIG. 9, the first CU (910) is divided from the upper left corner to the lower right corner, resulting in two triangular prediction units, PU1 and PU2. The second CU (920) is divided from the upper right corner to the lower left corner, resulting in two triangular prediction units, PU1 and PU2. Each triangular prediction unit PU1 or PU2 of the CU (910) or (920) is inter-predicted using its own motion information. In some implementations, only uni-prediction is allowed for each triangular prediction unit. Thus, each triangular prediction unit has one motion vector and one reference picture index. A uni-prediction motion constraint is applied to ensure that at most two motion-compensated predictions are performed for each CU, similar to conventional bidirectional prediction methods. In this way, processing complexity can be reduced. The uni-predictive motion information for each triangular prediction unit may be derived from the uni-predictive merge candidate list. In some other embodiments, bi-prediction is allowed for each triangular prediction unit. Thus, the bi-predictive motion information for each triangular prediction unit may be derived from the bi-predictive merge candidate list.
[0103] In some embodiments, if the CU-level flag indicates that the current CU is coded using a triangular partition mode, an index called a triangular partition index is further signaled. For example, the triangular partition index may have a value in the range of [0, 39]. Using this triangular partition index, the direction of the triangular partition (diagonal or anti-diagonal) and the motion information of each partition (e.g., merge index for each uni-prediction candidate list) may be obtained by a look-up table at the decoder side. After predicting each of the triangular prediction units based on the obtained motion information, in an embodiment, the sample values along the diagonal or anti-diagonal edges of the current CU are adjusted by performing a blending process with adaptive weights. As a result of the blending process, a prediction signal for the entire CU may be obtained. Thereafter, the transformation and quantization process may be applied to the entire CU in the same way as for other prediction modes. Finally, a motion field of a CU predicted using a triangular partition mode may be created, for example, by storing the motion information in a set of 4×4 units partitioned from the CU. The motion field may be used to build a merge candidate list, for example, in a subsequent motion vector prediction process.
[0104] 3. Building a list of uni-prediction candidates In some embodiments, a merge candidate list for prediction of two triangular prediction units of a coding block processed in a triangular prediction mode can be constructed based on a set of spatial and temporal neighboring blocks of the coding block. In one embodiment, the merge candidate list is a uni-prediction candidate list. The uni-prediction candidate list includes, in an embodiment, five uni-prediction motion vector candidates. For example, the five uni-prediction motion vector candidates are derived from seven neighboring blocks, including five spatial neighboring blocks (labeled with numbers 1-5 in FIG. 10) and two co-located temporal blocks (labeled with numbers 6-7 in FIG. 10).
[0105] In one example, the motion vectors of seven neighboring blocks are collected and put into the uni-prediction candidate list according to the following order: first, the motion vector of the uni-prediction neighboring block, then the L0 motion vector (i.e., the L0 motion vector part of the bi-prediction MV), the L1 motion vector (i.e., the L1 motion vector part of the bi-prediction MV), and the averaged motion vector of the L0 and L1 motion vectors of the bi-prediction MV for the bi-prediction neighboring block. In one example, if the number of candidates is less than five, a zero motion vector is added to the end of the list. In some other embodiments, the merge candidate list may include less than five or more than five uni-prediction or bi-prediction merge candidates selected from the same or different candidate positions as shown in FIG. 10.
[0106] 4. Lookup Tables and Table Indexes In an embodiment, the CU is coded in a triangular partition mode with a merge candidate list including five candidates. Thus, if five merge candidates are used for each triangular PU, there are 40 possible ways to predict the CU. In other words, there can be 40 different combinations of split directions and merge indexes: 2 (possible split directions) x (5 (possible merge indexes for the first triangular prediction unit) x (5 (possible merge indexes for the second triangular prediction unit) - 5 (possible number when a pair of the first and second prediction units share the same merge index)). For example, if the same merge index is determined for two triangular prediction units, the CU can be processed using, for example, the regular merge mode described in section II.1 instead of the triangular prediction mode.
[0107] Thus, in an embodiment, a triangular partition index in the range of [0,39] may be used to represent that one of 40 combinations is used based on the lookup table. FIG. 11 illustrates an exemplary lookup table (1100) used to derive split directions and merge indexes based on the triangular partition index. As shown in the lookup table (1100), a first row (1101) includes triangular partition indexes in the range of 0 to 39, a second row (1102) includes possible split directions represented by 0 or 1, a third row (1103) corresponds to a first triangular prediction unit and includes possible first merge indexes in the range of 0 to 4, and a fourth row (1104) corresponds to a second triangular prediction unit and includes possible second merge indexes in the range of 0 to 4.
[0108] For example, based on the column (1120) of the lookup table (1100), when a triangular partition index having a value of 1 is received at the decoder, it can be determined that the split direction is the split direction represented by the value of 1, and the first and second merge indices are 0 and 1, respectively. Because the triangular partition index is associated with the lookup table, the triangular partition index is also referred to as a table index in this disclosure.
[0109] 5. Adaptive Blending Along Triangular Partition Edges In an embodiment, after predicting each triangular prediction unit with its respective motion information, a blending process is applied to the two prediction signals of the two triangular prediction units to derive samples near the diagonal or anti-diagonal edges. The blending process adaptively selects between two groups of weighting factors depending on the difference of the motion vectors between the two triangular prediction units. In an embodiment, the two weighting factor groups are as follows: (1) First weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} for luma component samples and {7 / 8, 4 / 8, 1 / 8} for chroma component samples. (2) Second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} for luma component samples and {6 / 8, 4 / 8, 2 / 8} for chroma component samples. The second weighting factor group has more luma weighting factors and mixes more luma samples along the partition edges.
[0110] In an embodiment, the following condition is used to select one of the two weighting factor groups: if the reference pictures of the two triangular partitions are different from each other or the difference in the motion vectors between the two triangular partitions is greater than a threshold (e.g., 16 luma samples), the second weighting factor group is selected; otherwise, the first weighting factor group is selected.
[0111] FIG. 12 shows an example of a CU applying a first weighting factor group. As shown, a first coding block (1201) contains luma samples and a second coding block (1202) contains chroma samples. A set of pixels along a diagonal edge in coding block (1201) or (1202) are labeled with numbers 1, 2, 4, 6, and 7, corresponding to weighting factors 7 / 8, 6 / 8, 4 / 8, 2 / 8, and 1 / 8, respectively. For example, for a pixel labeled with the number 2, the sample value of the pixel after blending can be obtained as follows: Mixed sample value = 2 / 8 x P1 + 6 / 8 x P2 Here, P1 and P2 represent the sample values at each pixel, which belong to the prediction of the first triangular prediction unit and the second triangular prediction unit, respectively.
[0112] 6. Motion Vector Storage in the Motion Field FIG. 13 shows an example of how the motion vectors of two triangular prediction units in a CU coded in triangular prediction mode are combined and stored to form a motion field useful for subsequent motion vector prediction. As shown, a first coding block (1301) is divided into two triangular prediction units from the upper left corner to the lower right corner along a first diagonal edge (1303), while a second coding block (1302) is divided into two triangular prediction units from the upper right corner to the lower left corner along a second diagonal edge (1304). The first motion vector corresponding to the first triangular prediction unit of the coding block (1301) or (1302) is represented as Mv1, and the second motion vector corresponding to the second triangular prediction unit of the coding block (1301) or (1302) is represented as Mv2. Taking a coding block (1301) as an example, at the decoder side, two merge indexes corresponding to the first and second triangular prediction units in the coding block (1301) can be determined based on the received syntax information. After a merge candidate list is constructed for the coding block (1301), Mv1 and Mv2 can be determined according to the two merge indexes.
[0113] In an embodiment, the coding block (1301) is divided into a number of squares having a size of 4x4 samples. Corresponding to each 4x4 square, a uni-prediction motion vector (e.g., Mv1 or Mv2) or two motion vectors (forming bi-prediction motion information) are stored according to the position of the 4x4 square in each coding block (1301). As shown in the example of FIG. 13, in the 4x4 squares that do not overlap with the diagonal edges (1303) that divide the coding blocks (1301), either the uni-prediction motion vector Mv1 or Mv2 is stored. In contrast, in each of the 4x4 squares that overlap with the diagonal edges (1303) that divide the respective coding blocks (1301), two motion vectors are stored. For the coding blocks (1302), the motion vectors can be organized and stored in a similar manner to the coding blocks (1301).
[0114] A pair of bidirectional predictive motion vectors stored in 4×4 squares overlapping each diagonal edge may be derived from Mv1 and Mv2 according to the following rules, in an embodiment: (1) If Mv1 and Mv2 are motion vectors pointing in different directions (e.g., they relate to different reference picture lists L0 or L1), Mv1 and Mv2 are combined to form a pair of bidirectional predictive motion vectors. (2) If both Mv1 and Mv2 are heading in the same direction (e.g., they are related to the same reference picture list L0 (or L1)): (2.a) If the reference picture of Mv2 is the same as a picture in reference picture list L1 (or L0), Mv2 is modified to be associated with that reference picture in reference picture list L1 (or L0). Mv1 and Mv2 with respect to the modified associated reference picture list are combined to form a pair of bidirectional predictive motion vectors. (2.b) If the reference picture of Mv1 is the same as a picture in reference picture list L1 (or L0), Mv1 is modified to be associated with that reference picture in reference picture list L1 (or L0). Mv1 and Mv2 with respect to the modified associated reference picture list are combined to form a pair of bidirectional predictive motion vectors. (2.c) Otherwise, only Mv1 is stored for each 4x4 square.
[0115] 14A-14D show an example of deriving a pair of bi-predictive motion vectors according to an exemplary set of rules. In FIG. 14A-14D, two reference picture lists are used: The first reference picture list L0 contains reference pictures with reference picture indices (refIdx) of 0 and 1 along with Picture Order Count (POC) numbers being POC0 and POC8 respectively, while the second reference picture list L1 contains reference pictures with reference picture indices of 0 and 1 along with POC numbers being POC8 and POC16 respectively.
[0116] Figure 14A corresponds to rule (1). As shown in Figure 14A, Mv1 is associated with POC0 in L0 and therefore has reference picture index refIdx = 0, while MV2 is associated with POC8 in L1 and therefore has reference picture index refIdx = 0. Since Mv1 and Mv2 are associated with different reference picture lists, Mv1 and Mv2 are used together as a pair of bidirectional motion vectors.
[0117] Figure 14B corresponds to rule (2.a). As shown, Mv1 and Mv2 are associated with the same reference picture list L0. Mv2 points to POC8, which is also a member of L1. Therefore, Mv2 is modified to be associated with POC8 of L1, and the value of each reference index is changed from 1 to 0.
[0118] 14C and 14D correspond to rules (2b) and (2c).
[0119] 7. Syntax Elements for Signaling Triangular Prediction Parameters In some embodiments, the triangular prediction unit mode is applied to a CU in skip or merge mode. The block size of the CU cannot be smaller than 8x8. For a CU coded in skip or merge mode, a CU level flag is signaled to indicate whether the triangular prediction unit mode is applied to the current CU. In an embodiment, when the triangular prediction unit mode is applied to a CU, a table index is signaled indicating the direction of splitting the CU into two triangular prediction units and the motion vectors (or respective merge indexes) of the two triangular prediction units. The table index is in the range of 0 to 39. A lookup table is used to derive the split direction and the motion vector from the table index.
[0120] III. Signaling and Derivation of Triangular Prediction Parameters 1. Signaling of Triangular Prediction Parameters As described above, when a triangular prediction mode is applied to a coding block, three parameters are generated: a split direction, a first merge index corresponding to a first triangular prediction unit, and a second merge index corresponding to a second triangular prediction unit. As described, in some examples, the three triangular prediction parameters are signaled from the encoder side to the decoder side by signaling a table index. Based on a lookup table (e.g., lookup table (1100) in the example of FIG. 11), the three triangular prediction parameters can be derived using the table index received at the decoder side. However, storing the lookup table at the decoder requires additional memory space, which may be a burden in some implementations of the decoder. For example, the additional memory may result in increased cost and power consumption of the decoder.
[0121] This disclosure provides a solution to solve the above problem. Specifically, instead of signaling a table index and relying on a lookup table to interpret the table index, three syntax elements are signaled from the encoder side to the decoder side. The three triangular prediction parameters (split direction and two merge indexes) can be derived or determined at the decoder side based on the three syntax elements without using a lookup table. In an embodiment, the three syntax elements can be signaled in any order for each coding block.
[0122] In an embodiment, the three syntax elements include a split direction syntax element, a first index syntax element, and a second index syntax element. The split direction syntax element may be used to determine a split direction parameter. The first and second index syntax elements may be used in combination to determine parameters of the first and second merge indexes. In an embodiment, an index (different from a table index, called a triangular prediction index) may first be derived based on the first and second index syntax elements and the split direction syntax element. The three triangular prediction parameters may then be determined based on the triangular prediction index.
[0123] There are various ways to set or code the three syntax elements to signal the information of the three triangular prediction parameters. Regarding the split direction syntax element, in an embodiment, the split direction syntax element takes a value of 0 or 1 to indicate whether the split direction is from the top left corner to the bottom right corner or from the top right corner to the bottom left corner.
[0124] With respect to the first and second index syntax elements, in an embodiment, the first index syntax element is configured to have a value of the parameter of the first merge index, while the second index syntax element is configured to have a value of the second merge index if the second merge index is smaller than the first merge index, and to have a value of the second merge index minus 1 if the second merge index is larger than the first merge index. Since the second and first merge indexes are assumed to take different values as described above, the second and first merge indexes will not be equal to each other).
[0125] As an example, in an embodiment, the merge candidate list has a length of 5 (5 merge candidates). Thus, the first index syntax element may take a value of 0, 1, 2, 3, or 4, while the second index syntax element may take a value of 0, 1, 2, or 3. For example, if the first merge index parameter has a value of 2 and the second merge index parameter has a value of 4, then the first and second index syntax elements would have values of 2 and 3, respectively, to signal the first and second merge indexes.
[0126] In an embodiment, the coding block is located at a position having coordinates (xCb, yCb) relative to a reference point in the current picture, where xCb and yCb represent the horizontal and vertical coordinates of the current coding block, respectively. In some implementations, xCb and yCb are aligned to the horizontal and vertical coordinates at a granularity of 4x4. Thus, the split direction syntax element is represented as split_dir[xCb][yCb]. The first index syntax element is represented as merge_triangle_idx0[xCb][yCb]. The second index syntax element is represented as merge_triangle_idx1[xCb][yCb].
[0127] The three syntax elements can be signaled in any order in the bitstream. For example, the three syntax elements can be signaled in one of the following orders: 1.split_dir, merge_triangle_idx0, merge_triangle_idx1; 2.split_dir, merge_triangle_idx1, merge_triangle_idx0; 3.merge_triangle_idx0, split_dir, merge_triangle_idx1; 4.merge_triangle_idx0, merge_triangle_idx1, split_dir; 5.merge_triangle_idx1, split_dir, merge_triangle_idx0; 6.merge_triangle_idx1, merge_triangle_idx0, split_dir.
[0128] 2. Deriving Triangular Prediction Parameters 2.1 Deriving Triangular Prediction Parameters Based on Syntax Elements In an embodiment, the three triangular prediction parameters are derived based on the three syntax elements received at the decoder side. For example, the split direction parameter can be determined according to the value of the split direction syntax element. The first merge index parameter can be determined to have a value of the first index syntax element. If the second index syntax element has a value smaller than the first index syntax element, the second merge index parameter can be determined to have a value of the second index syntax element. On the other hand, if the second index syntax element has a value equal to or greater than the first index syntax element, the second merge index parameter can be determined to have a value of the second index syntax element plus one.
[0129] Below is a pseudo-code example that implements the above derivation process: m = merge_triangle_idx0[xCb][yCb]; n = merge_triangle_idx1[xCb][yCb]; n = n + (n >= m ? 1 : 0), where m and n represent the first and second merge index parameters, respectively, and merge_triangle_idx0[xCb][yCb] and merge_triangle_idx1[xCb][yCb] represent the first and second index syntax elements, respectively.
[0130] 2.2 Deriving triangular prediction parameters based on triangular prediction indices derived from syntax elements In an embodiment, a triangular prediction index is first derived based on the three syntax elements received at the decoder side. Then, three triangular prediction parameters are determined based on the triangular prediction index. For example, the values of the three syntax elements in binary bits can be combined into a bit string that forms a triangular prediction index. Then, the bits of each syntax element can be extracted from the bit string and used to determine the three triangular prediction parameters.
[0131] In an embodiment, the triangular prediction index is: mergeTriangleIdx[xCb][yCb] = a * merge_triangle_idx0[xCb][yCb] + b * merge_triangle_idx1[xCb][yCb] + c * split_dir[xCb][yCb] where mergeTriangleIdx[xCb][yCb] represents triangular prediction, a, b, c are integer constants, and merge_triangle_idx0[xCb][yCb], merge_triangle_idx1[xCb][yCb] and split_dir[xCb][yCb] represent the three signaled syntax elements, namely the first index syntax element, the second index syntax element, and the split direction syntax element, respectively.
[0132] As an example, in the above example where the merge candidate list contains five merge candidates, the constants can take on the following values: a = 8, b = 2, c = 1. In this scenario, the above linear function is equivalent to shifting the value of the first index syntax element left by 3 bits, shifting the value of the second index syntax element left by 2 bits, and then combining the bits of the three syntax elements into a bit string via an addition operation.
[0133] In other examples, the merge candidate list may have a length different from that of 5. Thus, the first and second merge index parameters may have values in a range different from [0,4]. Also, the respective first and second index syntax elements may have values in a different range. Thus, the constants a, b, and c may take on different values to properly combine the three syntax elements into a bit string. Additionally, the order of the three syntax elements may be arranged in a different way than in the above example.
[0134] After the triangular prediction index is determined as above, it is possible to determine three triangular prediction parameters based on the determined triangular prediction index. In one embodiment, corresponding to the above embodiment where a=8, b=2, and c=1, the split direction is: triangleDir = mergeTriangleIdx[xCb][yCb] & 1 where triangleDir represents the split direction parameter, and the last digit of the triangle prediction index is extracted by a binary AND operation (&) to be the value of the split direction parameter.
[0135] In an embodiment corresponding to the above example where a=8, b=2, c=1, the first and second merge indexes may be determined according to the following pseudocode: m = mergeTriangleIdx[xCb][yCb] >> 3; / / eliminate the last 3 bits n = (mergeTriangleIdx[xCb][yCb] >> 1) & 3; / / take the second and third from the end n = n + (n >= m ? 1 : 0) As shown here, the bits in the triangular prediction index except the last 3 bits are used as the first merge index parameter. The penultimate and third last bits of the triangular prediction index are used as the second merge index parameter if the value of the penultimate and third last bits plus 1 is less than the first merge index value. Otherwise, the value of the penultimate and third last bits plus 1 is used as the second merge index parameter.
[0136] 2.3 Adaptive Construction of Syntax Elements In some embodiments, the first and second index syntax elements can be configured to express different meanings depending on the split direction of the coding unit. Or, in other words, the first and second index syntax elements can be coded differently depending on which split direction is used to split the coding unit. For example, corresponding to different split directions, the probability distributions of the values of the two merge indexes can be different due to the characteristics of the current picture or local features within the current picture. Thus, the two index syntax elements can be adaptively coded according to the respective split directions, saving bits used for coding the index syntax elements.
[0137] For example, as shown in Figure 9, two triangular prediction units PU1 and PU2 are defined corresponding to the respective partition directions of coding blocks (910) and (920). If a first partition direction from top left to bottom right is used as in coding block (910), a first index syntax element may be used to carry the merge index corresponding to PU1, while a second index syntax element may be used to carry the merge index information corresponding to PU2. On the other hand, if a second partition direction from top right to bottom left is used as in coding block (920), a first index syntax element may be used to carry the merge index corresponding to PU2, while a second index syntax element may be used to carry the merge index information corresponding to PU1.
[0138] Corresponding to the adaptive coding of the index syntax element at the encoder side, appropriate decoding operations can be performed at the decoder side. A first example of pseudocode for implementing adaptive decoding of the index syntax element is shown below: if (triangleDir == 0) { m = merge_triangle_idx0[xCb][yCb]; n = merge_triangle_idx1[xCb][yCb]; n = n + (n >= m ? 1 : 0); } else { n = merge_triangle_idx0[xCb][yCb]; m = merge_triangle_idx1[xCb][yCb]; m = m + (m >= n ? 1 : 0); }
[0139] A second example of pseudocode for implementing adaptive decoding of the index syntax element when triangular prediction indexes are used is shown below: if (triangleDir == 0) { m = mergeTriangleIdx[xCb][yCb] >> 3; n = (mergeTriangleIdx[xCb][yCb] >> 1) & 3; n = n + (n >= m ? 1 : 0); } else { n = mergeTriangleIdx[xCb][yCb] >> 3 m = (mergeTriangleIdx[xCb][yCb] >> 1) & 3; m = m + (m >= n ? 1 : 0); }
[0140] In the first and second pseudocode examples above, the split direction (represented as TriangleDir) can be determined directly from the split direction syntax element, or in different embodiments, can be determined according to the triangular prediction index.
[0141] 3. Entropy coding of the three syntax elements 3.1 Binarization of the three syntax elements The three syntax elements (split direction syntax element, first and second index syntax elements) used to signal the three triangular prediction parameters (split direction, first and second merge index) can be coded with different binarization methods in various embodiments.
[0142] In one embodiment, the first index syntax element is coded with a truncated unary coding. In another embodiment, the first index syntax element is coded with a truncated binary coding. In one example, the maximum valid value of the first index syntax element is equal to 4. In another embodiment, a combination of a prefix and fixed length binarization is used to code the first index syntax element. In one example, a prefix bin is signaled first to indicate whether the first index syntax element is zero. If the first index syntax element is not zero, an additional bin is coded with a fixed length to indicate the actual value of the first index syntax element. Examples of truncated unary coding, truncated binary coding, and prefix and fixed length coding (maximum valid value equal to 4) are shown in Table 1. Table 1 [Table 1]
[0143] In one embodiment, the second index syntax element is encoded with a truncated unary encoding. In another embodiment, the second index syntax element is encoded with a binary encoding (i.e., a 2-bit fixed length encoding). An example of a truncated unary encoding and a binary encoding with a maximum valid value equal to 3 is shown in Table 2. Table 2 [Table 2]
[0144] 3.2 Context-Based Coding In some embodiments, certain constraints are applied to the probability model used for entropy coding of the three syntax elements for signaling the three triangular prediction parameters.
[0145] In one embodiment, we are limited to using at most N total context coding bins for the three syntax elements of a coding block processed in triangular prediction mode, where N is an integer that can be 0, 1, 2, 3, etc. In one embodiment, when N is equal to 0, all bins of these three syntax elements can be coded with equal probability.
[0146] In one embodiment, when N is equal to 1, there is only one context coded bin in the group of three syntax elements. In one example, one bin in the split direction syntax element is context coded, and the remaining bins in the split direction syntax element and all bins in the first and second index syntax elements are coded with equal probability. In another example, one bin in the second index syntax element is context coded, the remaining bins in the second index syntax element are context coded, and all bins in the split direction syntax element and the first index syntax element are coded with equal probability.
[0147] In one embodiment, there are two context coded bins in this group of three syntax elements when N is equal to 2. In one example, one bin in the first index syntax element and another bin in the second index syntax element are context coded, and the remaining bins of these three syntax elements are all coded with equal probability.
[0148] In one embodiment, when a context model is applied to a syntax element, only the first bin of the syntax element is applied to the context model, and the remaining bins of the syntax element are coded with equal probability.
[0149] 4. Example of triangular prediction process FIG. 15 shows a flow chart outlining a process (1500) according to an embodiment of the present disclosure. The process (1500) can be used in the reconstruction of a block coded in a triangular prediction mode to generate a prediction block for the block being reconstructed. In various embodiments, the process (1500) is performed by a processing circuit, such as a processing circuit in the terminal devices (210), (220), (230), and (240), a processing circuit performing the function of a video decoder (310), a processing circuit performing the function of a video decoder (410), a processing circuit performing the function of an entropy decoder (771), an inter-decoder (780), or the like. In some embodiments, the process (1500) is performed by software instructions, and the processing circuit performs the process (1500) when the processing circuit executes the software instructions. The process starts at (S1501) and proceeds to (S1510).
[0150] At (S1510), three syntax elements (a split direction syntax element, a first index syntax element, and a second index syntax element) are received in a video bitstream at a video decoder. The three syntax elements carry information of three triangular prediction parameters (a split direction parameter, a first merge index parameter, and a second merge index parameter) of a coding block coded in a triangular prediction mode. For example, a coding unit is split into a first triangular prediction unit and a second triangular prediction unit according to a split direction indicated by the split direction syntax element. The first and second triangular prediction units may be associated with first and second merge indices, respectively, that are related to a merge candidate list constructed for the coding block.
[0151] At (S1520), three triangular prediction parameters (split direction, first merge index, and second merge index) may be determined according to the three syntax elements received at (S1510). For example, various techniques described in Section III.2 may be used to derive the three triangular prediction parameters.
[0152] In (S1530), the coding block may be reconstructed according to the split direction, the first merge index, and the second merge index determined in (S1520). For example, a merge candidate list may be constructed for the coding block. Based on the merge candidate list and the first merge index for the merge candidate list, a first uni-prediction motion vector and a first reference picture index associated with the first uni-prediction motion vector may be determined. Similarly, based on the merge candidate list and the second merge index for the merge candidate list, a second uni-prediction motion vector and a second reference picture index associated with the second uni-prediction motion vector may be determined. Then, a first prediction corresponding to the first triangular prediction unit may be determined according to the first uni-prediction motion vector and the first reference picture index associated with the first uni-prediction motion vector. Similarly, a second prediction corresponding to the second triangular prediction unit may be determined according to the second uni-prediction motion vector and the second reference picture index associated with the second uni-prediction motion vector. Then, based on the first and second predictions, an adaptive weighting process may be applied to samples along the diagonal edge between the first and second triangular prediction units to derive a final prediction for the coding block. The process (1500) may proceed to (S1599) and end thereat.
[0153] IV. Computer Systems The techniques described above can be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 16 illustrates a computer system (1600) suitable for implementing certain embodiments of the disclosed subject matter.
[0154] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to produce code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly, or via interpretation, microcode execution, etc.
[0155] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0156] 16 with respect to computer system (1600) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (1600).
[0157] The computer system (1600) may include certain human interface input devices that may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swiping, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereo video).
[0158] The input human interface devices may include one or more (only one of each depicted) of a keyboard (1601), a mouse (1602), a trackpad (1603), a touch screen (1610), a data glove (not shown), a joystick (1605), a microphone (1606), a scanner (1607), and a camera (1608).
[0159] The computer system (1600) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., a touch screen (1610), data gloves (not shown), or joystick (1605), although it is also possible that there are haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers (1609), headphones (not shown)), visual output devices (e.g., screens (1610), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, some of which may be capable of outputting two-dimensional visual output, or three or more dimensional visual output by means such as stereographic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown).
[0160] The computer system (1600) may also include human accessible storage devices and their associated media, such as optical media (1621), including CD / DVD ROM / RW (1620) having media such as CDs / DVDs, thumb drives (1622), removable hard drives or solid state drives (1623), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0161] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitional signals.
[0162] The computer system (1600) may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, and the like. Examples of networks include Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, and the like), TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial TV, vehicular and industrial including CANBus, and the like. Certain networks generally require an external network interface adapter (e.g., a USB port on the computer system (1600)) that connects to a particular general-purpose data port or peripheral bus (1649); other networks are generally integrated into the core of the computer system (1600) by attachment to a system bus as described below (e.g., an Ethernet interface in a PC computer system, or a cellular network interface in a smartphone computer system). Using any of these networks, the computer system (1600) may communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast television), unidirectional transmit only (e.g., a CAN bus to a particular CAN bus device), or bidirectional, for example, to other computer systems using local or wide digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.
[0163] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to a core (1640) of the computer system (1600).
[0164] The cores (1640) may include one or more central processing units (CPUs) (1641), graphics processing units (GPUs) (1642), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1643), hardware accelerators for specific tasks (1644), etc. These devices may be connected via a system bus (1648) along with read only memory (ROM) (1645), random access memory (1646), and internal mass storage (1647) such as an internal non-user accessible hard drive, SSD, etc. In some computer systems, the system bus (1648) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1648) or via a peripheral bus (1649). Peripheral bus architectures include PCI, USB, etc.
[0165] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) may execute certain instructions that, in combination, may constitute the aforementioned computer code. The computer code may be stored in a ROM (1645) or a RAM (1646). Temporary data may also be stored in the RAM (1646), and persistent data may be stored, for example, in an internal mass storage (1647). Fast storage and retrieval from any memory device may be made possible by the use of a cache memory that may be closely associated with one or more of the CPU (1641), GPU (1642), mass storage (1647), ROM (1645), RAM (1646), etc.
[0166] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be specially designed and constructed for the purposes of the present disclosure, or they can be well known and available to those having skill in the computer software arts.
[0167] As a non-limiting example, a computer system having the architecture (1600), and in particular the core (1640), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with the particular storage that is the core (1640), such as in-core mass storage (1647) or ROM (1645), of a non-transitory nature, in addition to user-accessible mass storage as described above. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1640). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (1640), and in particular the processor therein (including a CPU, GPU, FPGA, etc.), to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (1646) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1644)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, as appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0168] (Appendix 1) 1. A method for video decoding in a decoder, comprising: receiving a split direction syntax element, a first index syntax element, and a second index syntax element associated with a coding block of a picture, the coding block being coded in a triangular prediction mode and partitioned into a first triangular prediction unit and a second triangular prediction unit according to a split direction indicated by the split direction syntax element, the first and second index syntax elements indicating first and second merge indexes for merge candidate lists constructed for the first and second triangular prediction units, respectively; determining the split direction, the first merge index, and the second merge index based on the split direction syntax element, the first index syntax element, and the second index syntax element; reconstructing the coding block according to the determined split direction, the determined first merge index, and the determined second merge index. The method includes: (Appendix 2) The step of determining the split direction, the first merge index, and the second merge index includes: determining a triangular prediction index indicating a combination of the split direction, the first merge index, and the second merge index based on the split direction syntax element, the first index syntax element, and the second index syntax element; determining the split direction, the first merge index, and the second merge index based on the determined triangular prediction index; The method of claim 1, comprising: (Appendix 3) The step of determining the split direction, the first merge index, and the second merge index includes: determining the split direction according to the split direction indicated by the split direction syntax element; calculating the first merge index and the second merge index according to the split direction syntax element, the first index syntax element, and the second index syntax element; 3. The method according to claim 1 or 2, comprising: (Appendix 4) 4. The method of any one of claims 1-3, wherein the split direction syntax element has a value of 0 or 1. (Appendix 5) 5. The method of claim 1, wherein the first index syntax element has a value of 0, 1, 2, 3, or 4, and the second index syntax element has a value of 0, 1, 2, or 3. (Appendix 6) The step of determining the split direction, the first merge index, and the second merge index includes: determining, based on determining that the split direction syntax element is a first value, that the split direction is from the top left corner to the bottom right corner; determining, based on a determination that the split direction syntax element is a second value, that the split direction is from the upper right corner to the lower left corner; 6. The method according to any one of claims 1-5, comprising: (Appendix 7) The step of determining the split direction, the first merge index, and the second merge index includes: A triangular prediction index indicating a combination of the split direction, the first merge index, and the second merge index. Triangular prediction index = a * 1st index syntax element + b * 2nd index syntax element + c * division direction syntax element 7. The method of any one of claims 1-6, comprising determining according to: (Appendix 8) 8. The method according to claim 7, wherein a is equal to 8, b is equal to 2, and c is equal to 1. (Appendix 9) The step of determining the split direction, the first merge index, and the second merge index includes: 8. The method of claim 7, comprising determining the split direction, indicating whether the coding block is split from the top left corner to the bottom right corner or from the top right corner to the bottom left corner, according to a least significant bit of the triangular prediction index. (Appendix 10) The step of determining the split direction, the first merge index, and the second merge index includes: 8. The method of claim 7, comprising determining the first and second merge indexes based on the triangular prediction index. (Appendix 11) The step of determining the split direction, the first merge index, and the second merge index includes: determining a split direction indicating whether the coding block is split in one of a first direction from a top left corner to a bottom right corner and a second direction from a top right corner to a bottom left corner; When the division direction is a first direction of the first and second directions, determining the first merge index from bits of the triangular prediction index excluding the last three bits; determining the second merge index to be the value represented by the third and second last bits of the triangular prediction index when the value represented by the third and second last bits of the triangular prediction index is less than the determined first merge index; determining the second merge index to be the value represented by the third and second last bits of the triangular prediction index plus one if the value represented by the third and second last bits of the triangular prediction index is greater than or equal to the determined first merge index; When the division direction is the second of the first and second directions, determining the second merge index from bits of the triangular prediction index excluding the last three bits; determining the first merge index to be the value represented by the third and second last bits of the triangular prediction index when the value represented by the third and second last bits of the triangular prediction index is less than the determined second merge index; determining the first merge index to be the value represented by the third and second last bits of the triangular prediction index plus one if the value represented by the third and second last bits of the triangular prediction index is greater than or equal to the determined second merge index; 8. The method of claim 7, comprising: (Appendix 12) The step of determining the split direction, the first merge index, and the second merge index includes: determining the split direction according to the split direction syntax element; determining the first merge index to have a value of the first index syntax element; determining the second merge index to have a value of the second index syntax element if the second index syntax element has a smaller value than the first index syntax element; determining the second merge index to have a value of the second index syntax element plus one if the second index syntax element has a value greater than or equal to the value of the first index syntax element; 12. The method of any one of claims 1-11, comprising: (Appendix 13) The step of determining the split direction, the first merge index, and the second merge index includes: determining a split direction indicating whether the coding block is split in one of a first direction from a top left corner to a bottom right corner and a second direction from a top right corner to a bottom left corner; When the division direction is a first direction of the first and second directions, determining the first merge index to be a value of the first index syntax element; determining the second merge index to be a value of the second index syntax element if the second index syntax element is less than the first index syntax element; if the second index syntax element is greater than or equal to the first index syntax element, determining the second merge index to be the value of the second index syntax element plus one; When the division direction is the second of the first and second directions, determining the second merge index to be a value of the first index syntax element; determining the first merge index to be the value of the second index syntax element if the second index syntax element is less than the first index syntax element; if the second index syntax element is greater than or equal to the first index syntax element, determining the first merge index to be the value of the second index syntax element plus one; 13. The method of any one of claims 1-12, comprising: (Appendix 14) 14. The method of any one of claims 1-13, wherein the first and second index syntax elements are coded differently to express different meanings depending on whether the coding block is divided from the top left corner to the bottom right corner or from the top right corner to the bottom left corner. (Appendix 15) 15. The method of any one of claims 1-14, wherein at least one of the first and second index syntax elements is encoded using a truncated unary encoding. (Appendix 16) 16. The method of any one of claims 1-15, wherein a bin in the first index syntax element and a bin in the second index syntax element are context coded. (Appendix 17) 17. The method of any one of claims 1-16, wherein a first bin of one of the split direction syntax element, the first index syntax element, and the second index syntax element is context coded. (Appendix 18) The split direction syntax element, the first index syntax element, and the second index syntax element are arranged in the bitstream in the following order: the split direction syntax element, the first index syntax element, and the second index syntax element; the split direction syntax element, the second index syntax element, and the first index syntax element; the first index syntax element, the split direction syntax element, and the second index syntax element; the first index syntax element, the second index syntax element, and the split direction syntax element; the second index syntax element, the split direction syntax element, and the first index syntax element; the second index syntax element, the first index syntax element, and the split direction syntax element; 18. The method according to any one of claims 1-17, wherein the method is transmitted in one of the following ways: (Appendix 19) 20. A video decoding apparatus comprising a processing circuit configured to perform a method according to any one of claims 1-18. (Appendix 20) A computer program product that causes a computer to carry out the method according to any one of claims 1 to 18. (Appendix 21) 1. A method for video decoding in a decoder, comprising: receiving a partition direction syntax element, a first index syntax element, and a second index syntax element associated with a coding block of a picture, the coding block being coded according to a triangular prediction mode and partitioned into a first prediction unit and a second prediction unit according to a partition direction; determining the split direction by decoding the split direction syntax element; determining a first merge index by decoding the first index syntax element, the first merge index identifying a first motion information in a merge candidate list constructed for the coding block, the first merge index element being coded according to a first one or more bins, only one of the first one or more bins being context coded; determining a second merge index by decoding the second index syntax element, the second merge index identifying second motion information in the merge candidate list constructed for the coding block, the second merge index element being coded according to a second one or more bins, only one of the second one or more bins being context coded; determining a first prediction sample of the first prediction unit according to the first motion information; determining a second prediction sample of the second prediction unit according to the second motion information; reconstructing the coding block according to the first prediction sample of the first prediction unit and the second prediction sample of the second prediction unit; The method includes: (Appendix 22) 1. A method for video encoding in an encoder, comprising: determining and sending to a decoder a split direction syntax element, a first index syntax element, and a second index syntax element associated with a coding block of a picture, the coding block being coded in a triangular prediction mode and partitioned into a first triangular prediction unit and a second triangular prediction unit according to a split direction indicated by the split direction syntax element, the first and second index syntax elements indicating first and second merge indexes for merge candidate lists constructed for the first and second triangular prediction units, respectively; the split direction, the first merge index, and the second merge index are determined based on the split direction syntax element, the first index syntax element, and the second index syntax element; The coding block is reconstructed according to the determined split direction, the determined first merge index, and the determined second merge index.
[0169] Appendix A: Abbreviations JEM: Joint Exploration Model VVC: General Purpose Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI:Video Usability Information GOPs: Group of Pictures TUs: conversion units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Predicted blocks HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: cold cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD:Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Device IC: Integrated Circuit CU: Coding Unit
Claims
1. A method for video decoding in a decoder, comprising: receiving a split direction syntax element, a first merge triangular index syntax element, and a second merge triangular index syntax element related to an encoded block of a picture, wherein the encoded block is encoded in a triangular prediction mode and is divided into a first prediction unit and a second prediction unit according to a split direction; determining the split direction, a first merge index, and a second merge index based on the split direction syntax element, the first merge triangular index syntax element, and the second merge triangular index syntax element, wherein the first merge index identifies first motion information in a merge candidate of the encoded block, and the second merge index identifies second motion information in the merge candidate of the encoded block; determining first prediction samples of the first prediction unit according to the first motion information; determining second prediction samples of the second prediction unit according to the second motion information; reconstructing the encoded block according to the first prediction samples of the first prediction unit and the second prediction samples of the second prediction unit; wherein the step of determining the split direction, the first merge index, and the second merge index comprises: determining a triangular prediction index indicating the split direction, the first merge index, and the second merge index according to: triangular prediction index = a * first merge triangular index syntax element + b * second merge triangular index syntax element + c * split direction syntax element where a, b, and c are integers.
2. The method according to claim 1, wherein the split direction syntax element is 1 bit.
3. The method according to claim 1, wherein the split direction syntax element indicates a value of 0 or 1.
4. In the method according to claim 1, the step of determining the split direction comprises: determining that the split direction is from the upper left corner to the lower right corner based on the split direction syntax element indicating a first value; and Based on the division direction syntax element indicating the second value, determining that the division direction is from the upper right corner to the lower left corner; A method including the above. **Claim 5** A video processing apparatus including a processing circuit configured to execute the method according to any one of Claims 1 to 4. **Claim 6** A computer program for causing a computer to execute the method according to any one of Claims 1 to 4.
Citation Information
Patent Citations
Encoding device, decoding device, encoding method, and decoding method
WO2019138998A1
Position dependent storage of motion information
WO2020094078A1
Method for encoding / decoding image signal and device therefor
WO2020096428A1
An encoder, a decoder and corresponding methods for inter prediction
WO2020106190A1
Triangle motion information for video coding
WO2020118064A1