Coding process for the geometric partitioning mode

JP7686738B2Active Publication Date: 2025-06-02HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023222162
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-28
Filing Date
2023-12-28
Publication Date
2025-06-02
Estimated Expiration
2040-09-28

AI Technical Summary

Technical Problem

The challenge of efficiently compressing video data for transmission and storage while maintaining high quality is significant due to the substantial amount of data required, which strains limited network and memory resources.

Method used

A method and apparatus for video coding that utilizes geometric partitioning modes, including decoding processes for current coding blocks based on splitting mode index values and angle index values, improving buffer utilization and decoding efficiency.

Benefits of technology

Enhances compression ratios with minimal sacrifice in picture quality by optimizing the decoding process through geometric partitioning, thereby improving buffer utilization and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000072_0000
    Figure 00000072_0000
  • Figure 00000073_0000
    Figure 00000073_0000
  • Figure 00000074_0000
    Figure 00000074_0000
Patent Text Reader

Abstract

To provide a coding process for a geometric partition mode.SOLUTION: A method of coding implemented by a decoding device includes obtaining a splitting mode index value for the current coding block, obtaining an angle index value angleIdx for the current coding block according to a splitting mode index value and a table specifying an angle index value angleIdx on the basis of the splitting mode index value, setting an index value partIdx according to the angle index value angleIdx, and decoding the current coding block according to the index value partIdx.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims priority to International Patent Application No. PCT / EP2019 / 076805, filed October 3, 2019. The disclosures of the aforementioned patent applications are incorporated herein by reference in their entireties.

[0002] FIELD OF THE DISCLOSURE Embodiments of the present application relate generally to the field of picture processing, and more particularly to geometric partitions. [Background technology]

[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVDs and Blu-ray discs, video content collection and editing systems, and camcorders in security applications.

[0004] The amount of video data required to depict even a relatively short video can be substantial, which can become difficult when the data is streamed or otherwise communicated over a communications network that has limited bandwidth capacity. Therefore, video data is typically compressed before being communicated over modern telecommunications networks. The size of the video can also be an issue when the video is stored on a storage device, since memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data prior to transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and a constantly increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little sacrifice in picture quality are desirable. Summary of the Invention

[0005] Embodiments of the present application provide apparatus and methods for encoding and decoding according to the independent claims.

[0006] These and other objects are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, the description and the drawings.

[0007] A first aspect of the present invention provides a method of coding implemented by a decoding device, the method comprising: obtaining a split mode index value for a current coding block; obtaining an angle index value angleIdx for the current coding block in accordance with the split mode index value and a table that specifies an angle index value angleIdx based on the split mode index value; setting an index value partIdx in accordance with the angle index value angleIdx; and decoding the current coding block in accordance with the index value partIdx.

[0008] According to an embodiment of the present invention, the coding block is decoded according to the index value. The decoding process may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. Therefore, the buffer utilization and the decoding efficiency are improved.

[0009] In one implementation form, the split mode index value is used to indicate which geometric partition mode is used for the current coding block, for example, geo_partition_idx or merge_gpm_partition_idx.

[0010] In one example, merge_gpm_partition_idx[x0][y0] or geo_partition_idx specifies the partitioning shape for the geometric partitioning merge mode. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0011] When merge_gpm_partition_idx[x0][y0] or go_partition_idx[x0][y0] is not present, it is inferred to be equal to 0.

[0012] In one example, the split mode index value may be obtained by parsing an index value coded in a video bitstream, or the split mode index value may be determined according to a syntax value parsed from the video bitstream.

[0013] The bit stream may be obtained according to a wireless or wired network. The bit stream may be transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or infrared, radio, microwave, WIFI, Bluetooth, LTE, or 5G wireless technologies.

[0014] In one embodiment, the bitstream is a sequence of bits, e.g., in the form of a Network Abstraction Layer (NAL) unit stream or a byte stream, forming a representation of a sequence of access units (AUs) forming one or more Coded Video Sequences (CVSs).

[0015] In some embodiments, for a decoding process, a decoder side reads the bitstream and derives the decoded picture from the bitstream, and for an encoding process, an encoder side generates the bitstream.

[0016] Typically, a bitstream will contain syntax elements that are formed by a syntax structure. Syntax element: An element of data that is represented in a bitstream. Syntax structure: zero or more syntax elements that occur together in a bitstream in a specified order.

[0017] In a particular example, a bitstream format specifies the relationship between Network Abstraction Layer (NAL) unit streams and byte streams, both of which are referred to as bitstreams.

[0018] A bitstream can be in one of two formats, e.g., the NAL unit stream format or the byte stream format. The NAL unit stream format is conceptually the more "basic" type. The NAL unit stream format contains a sequence of syntactic structures called NAL units. This sequence is ordered in decoding order. There are constraints imposed on the decoding order (and) of the NAL units in a NAL unit stream.

[0019] A byte stream format may be constructed from the NAL unit stream format by ordering the NAL units in decoding order and prefixing each NAL unit with a start code prefix and zero or more zero-valued bytes to form a stream of bytes. The NAL unit stream format may be extracted from the byte stream format by searching for the location of a unique start code prefix pattern in this stream of bytes.

[0020] This term specifies one embodiment of the relationship between the source and the decoded pictures provided via the bitstream.

[0021] The video source represented by a bitstream is a sequence of pictures in decoding order.

[0022] Typically, the value of merge_gpm_partition_idx[x0][y0] is decoded from the bitstream. In one example, the value range of merge_gpm_partition_idx[][] is 0 to 63, inclusive. In one example, the decoding process of merge_gpm_partition_idx[][] is "bypassed."

[0023] When merge_gpm_partition_idx[x0][y0] is not present, it is inferred to be equal to 0.

[0024] In one implementation form, the angle index value angleIdx is used for the geometric partition of the current coding block.

[0025] In one example, the angle index value angleIdx specifies the angle index of a geometric partition, hereinafter the angle index value angleIdx is also referred to as the partition angle variable angleIdx.

[0026] The value of the angle parameter for the current block is obtained according to the index value of the division mode and the value of a predefined lookup table.

[0027] In one embodiment, the partition angle variable angleIdx (angle parameter) and distance variable distanceIdx of the geometric partitioning mode are set according to the value of merge_gpm_partition_idx[xCb][yCb] (indicator) specified in the following table. In an implementation form, this relationship may be implemented according to Table 1 or according to a function. [Table 1]

[0028] In one implementation, the index value partIdx satisfies the following:

number

[0029] Here, threshold1 and threshold2 are integer values, and threshold1 is less than threshold2.

[0030] Thus, the index value partIdx is set to 1 when the angle index value angleIdx is in the interval between threshold1 and threshold2 or is equal to threshold1 or threshold2, otherwise the index value partIdx is set to 0 when the angle index value angleIdx is outside this interval. According to the index value partIdx, the current coding block is decoded. For example, the index value partIdx defines which of two sub-blocks is the first sub-block (partIdx=0) and which is the second sub-block to be decoded (partIdx=1).

[0031] In one example, threshold1 is 13 and threshold2 is 27.

[0032] In one example, the index value may be a fracture index, denoted as isFlip, where the index value isFlip also satisfies:

number

[0033] The decoding process may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. Therefore, the buffer utilization and decoding efficiency are improved.

[0034] In one implementation form, decoding the coding block includes storing motion information for the current block according to the index value partIdx.

[0035] A second aspect of the present invention provides a video decoder comprising: a parsing unit configured to obtain a partition mode index value for a current coding block; an angle index value obtaining unit configured to obtain an angle index value angleIdx for the current coding block according to a partition mode index value and a table specifying an angle index value angleIdx based on the partition mode index value; a setting unit configured to set an index value partIdx according to the angle index value angleIdx; and a processing unit configured to decode the current coding block according to the index value partIdx.

[0036] In one implementation, the index value partIdx satisfies the following:

number

[0037] In one implementation, threshold1 is 13 and threshold2 is 27.

[0038] In one implementation, the processing unit is configured to store the motion information for the current block according to the index value partIdx.

[0039] In one implementation, the partition mode index value is used to indicate which geometric partition mode is to be used for the current coding block.

[0040] In one implementation, angleIdx is used for the geometric partition of the current coding block.

[0041] The method according to the first aspect of the invention may be performed by a device according to the second aspect of the invention. Further features and implementation forms of said method correspond to those of the apparatus according to the second aspect of the invention. In one embodiment, a decoder is disclosed comprising a processing circuit for performing the method according to any one of the above embodiments and implementations.

[0042] In one embodiment, a computer program product is disclosed that includes program code for performing a method according to any one of the above embodiments and implementations.

[0043] In one embodiment, one or more processors; A decoder or encoder is provided that includes: a non-transitory computer-readable storage medium coupled to a processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder or encoder to perform a method according to any one of the above embodiments and implementations.

[0044] In one embodiment, a non-transitory storage medium is provided that includes an encoded bitstream decoded by an image decoding device, the bitstream being generated by dividing a frame of a video signal or image signal into a number of blocks and including a number of syntax elements, the number of syntax elements including an indicator (syntax) according to any one of the above embodiments and implementations.

[0045] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. [Brief description of the drawings]

[0046] In the following, embodiments of the invention are explained in more detail with reference to the accompanying figures and drawings.

[0047] [Figure 1A] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Diagram 2] FIG. 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention. [Diagram 3] 2 is a block diagram illustrating an exemplary structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] FIG. 1 is a block diagram showing an example of an encoding device or a decoding device. [Diagram 5] FIG. 13 is a block diagram showing another example of an encoding device or a decoding device. [Figure 6a] An example of co-located blocks is shown. [Figure 6b] An example of spatially adjacent blocks is shown. [Figure 7] Several examples of triangular prediction modes are given. [Figure 8] Some examples of sub-block prediction modes are given. [Figure 9] Some examples of block partitions are given below. [Figure 10] Some examples of block partitions are given below. [Figure 11] Some examples of block partitions are given below. [Figure 12] Some examples of block partitions are given below. [Figure 13] An example implementation of a predefined lookup table for stepD is shown below. [Figure 14] An example implementation of a predefined lookup table for f() is shown below. [Figure 15] An example of a quantization aspect involving a predefined lookup table for stepD is shown. [Figure 16] An example of a quantization scheme where a maximum ρmax is defined for a given coding block is given. [Figure 17] An example of a quantization scheme in which alternative maximum ρmax is defined for a given coding block is given. [Figure 18] FIG. 1 is a block diagram showing the structure of an example of a content supply system. [Figure 19] FIG. 2 is a block diagram illustrating the structure of an example of a terminal device. [Figure 20] Another embodiment of the present invention has been presented. [Figure 21] Another embodiment of the present invention has been presented. [Figure 22] 2 is a flow chart illustrating a method according to the present invention. [Diagram 23] FIG. 1 is a block diagram showing an apparatus according to the present invention; [Figure 24] FIG. 2 is a block diagram showing an embodiment of a decoder according to the present invention;

[0048] Hereinafter, the same reference signs, unless expressly specified otherwise, refer to identical or at least functionally equivalent features. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0049] In the following description, reference is made to the accompanying drawings, which form a part of this disclosure, and which are intended to illustrate certain aspects of the embodiments of the present invention or in which the embodiments of the present invention may be used. It is understood that the embodiments of the present invention may be used in other ways and include structural or logical changes not shown in the drawings. Therefore, the following detailed description should not be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0050] For example, it is understood that disclosure related to a described method may also be true of a corresponding device or system configured to perform the method, and vice versa. For example, if one or more particular method steps are described, the corresponding device may include one or more units, e.g., functional units, to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units each perform one or more of the multiple steps), even if such one or more units are not explicitly described or shown in the figures. On the other hand, for example, if a particular apparatus is described based on one or more units, e.g., functional units, the corresponding method may include a step to perform the functionality of the one or more units (e.g., one step performs the functionality of the one or more units, or multiple steps, each of the multiple units performs the functionality of one or more of the multiple units), even if such one or more steps are not explicitly described or shown in the figures. Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.

[0051] Video coding typically refers to the processing of a series of pictures forming a video or a video sequence. Instead of the term "picture", the terms "frame" or "image" may be used synonymously in the field of video coding. Video coding (or encoding in general) includes two parts: video encoding and video decoding. Video coding is performed at the source side and typically involves processing (e.g., by compression) of an original video picture to reduce the amount of data required to represent the original video picture (for more efficient storage and / or transmission). Video decoding is performed at the destination side and typically involves processing in the reverse direction compared to the encoder to reconstruct the video picture. The embodiments referring to the "coding" of a video picture (or, in general, a picture) should be understood to relate to the "encoding" or "decoding" of the video picture or the respective video sequence. The combination of the encoding unit and the decoding unit is also called a CODEC (Coding and Decoding).

[0052] In the case of lossless video coding, the original video picture can be reconstructed and the reconstructed video picture has the same quality as the original video picture (assuming there is no transmission or other data loss during storage or transmission). In the case of lossy video coding, further compression, e.g. by quantization, is performed to reduce the amount of data representing the video picture, but it cannot be completely reconstructed at the decoder and the quality of the reconstructed video picture is lower or worse than the quality of the original video picture.

[0053] Some video coding standards belong to the group of "lossy hybrid video codecs" (e.g., combining spatial and temporal prediction in the sample domain with 2D transform coding to apply quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks and coding is typically performed at the block level. In other words, at the encoder, a video is typically processed, e.g., encoded, at the block (video block) level, e.g., by generating a predictive block using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracting the predictive block from a current block (the block currently being / to be processed) to obtain a residual block, transforming the residual block and quantizing the residual block in the transform domain to reduce (compress) the amount of data to be transmitted, while at the decoder, the inverse processing compared to the encoder is applied to the coded or compressed block to reconstruct the current block for presentation. Additionally, the encoder replicates the decoder processing loop so that both generate the same predictions (e.g., intra-prediction and inter-prediction) and / or processes, e.g., reconstructions for coding subsequent blocks. In the following embodiment of the video coding system 10, a video encoder 20 and a video decoder 30 are described based on Figures 1-3.

[0054] 1A is a schematic block diagram illustrating an example coding system 10, e.g., video coding system 10 (or coding system 10 for short), that may employ techniques of the present application. A video encoder 20 (or encoder 20 for short) and a video decoder 30 (or decoder 30 for short) of video coding system 10 represent examples of devices that may be configured to perform techniques in accordance with various examples described herein. As shown in FIG. 1A, coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14 for decoding, e.g., encoded picture data 13.

[0055] The source device 12 includes an encoder 20 and may additionally, for example optionally, include a picture source 16, a pre-processor (or pre-processing unit) 18, for example a picture pre-processor 18, and a communication interface or unit 22.

[0056] Picture source 16 may include any kind of picture capture device, e.g., a camera for capturing real world pictures, and / or any kind of picture generation device, e.g., a computer graphics processor for generating computer animated pictures, or any kind of other device for obtaining and / or providing real world pictures, computer generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures).

[0057] The picture source may be any type of memory or storage that stores any of the pictures mentioned above.

[0058] To distinguish between the preprocessor 18 and the processing performed by the preprocessing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17. The preprocessor 18 is configured to receive the (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture data 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the preprocessing unit 18 may be any component. The video encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21 (described in further detail below, for example, based on FIG. 2). The communications interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and to transmit the encoded picture data 21 (or any further processed version thereof) via the communications channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0059] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally, for example optionally, include a communications interface or unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0060] The communications interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or a further processed version thereof), e.g. directly from the source device 12 or from any other source, e.g. a storage device, e.g. a coded picture data storage device, and to provide the encoded picture data 21 to a decoder 30.

[0061] The communications interface 22 and the communications interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communications link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any type of combination thereof.

[0062] The communications interface 22 may be configured, for example, to package the encoded picture data 21 in a suitable format, for example into packets, and / or to process the encoded picture data using any type of transmission coding or processing for transmission over a communications link or network.

[0063] The communications interface 28, forming a counterpart to the communications interface 22, may for example be configured to receive the transmitted data and to process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpackaging to obtain the encoded picture data 21.

[0064] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrow of communication channel 13 in FIG. 1A pointing from source device 12 to destination device 14, or as bidirectional communication interfaces, and may be configured to send and receive messages, e.g., set up connections, to acknowledge and exchange communication links and / or data transmissions, e.g., encoded picture data transmissions, and any other information related to the transmission.

[0065] The decoder 30 is configured to receive the encoded picture data 21 and to provide decoded picture data 31 or decoded pictures 31 (described in further detail below, e.g. based on FIG. 3 or FIG. 5).

[0066] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31, e.g. the decoded picture 31, to obtain post-processed picture data 33, e.g. the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, e.g., color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare the decoded picture data 31, e.g., for display by a display device 34. The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g., to a user or viewer. The display device 34 may be or include any kind of display for presenting the reconstructed picture, e.g., an integrated display or an external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0067] 1A depicts source device 12 and destination device 14 as separate devices, an embodiment of the devices may include both or both functionality of source device 12 or corresponding functionality, and destination device 14 or corresponding functionality. In such an embodiment, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented by the same hardware and / or software, separate hardware and / or software, or any combination thereof.

[0068] As will be clear to one of skill in the art based on the description, the presence and (exact) division of functionality of different units or functionalities within source device 12 and / or destination device 14, as shown in FIG. 1A, may vary depending on the actual device and application.

[0069] The encoder 20 (e.g., video encoder 20), the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30 may be implemented by processing circuitry such as that shown in FIG. 1B, one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 may be implemented via processing circuitry 46 to embody various modules as described with respect to the encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via processing circuitry 46 to embody various modules as described with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations as described below. As shown in FIG. 5, if the techniques are implemented partially in software, a device can store instructions for the software on a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either video encoder 20 or video decoder 30 may be integrated as part of a combined encoder / decoder (CODED) in a single device, for example, as shown in FIG. 1B.

[0070] The source device 12 and destination device 14 may include any of a wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content delivery server), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system at all or any type of operating system.

[0071] In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices. In some cases, video coding system 10 shown in FIG. 1A is merely an example, and the techniques of the present application may be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between encoding and decoding devices. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data in memory and / or retrieve data from memory and decode it.

[0072] For ease of explanation, embodiments of the present invention are described herein with reference to, for example, High Efficiency Video Coding (HEVC) or Generic Video Coding (VVC) reference software, the next generation video encoding standard developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Coding Experts Group (MPEG) Joint Collaboration Team on Video Coding (JCT-VC). Those skilled in the art will appreciate that embodiments of the present invention are not limited to HEVC or VVC.

[0073] Encoder and encoding method

[0074] FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or an input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output 272 (or an output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder according to a hybrid video codec.

[0075] The residual computation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 are sometimes referred to as forming a forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are sometimes referred to as forming a backward signal path of the video encoder 20, which corresponds to the signal path of the decoder (see video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DRB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming a “built-in decoder” of the video encoder 20.

[0076] Picture & Picture Partitioning (Picture & Block)

[0077] The encoder 20 may be arranged to receive, for example via an input 201, a picture 17 (or picture data 17), for example a picture of a video or a sequence of pictures forming a video sequence. The received picture or picture data may be a preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to the picture 17. Picture 17 may also be referred to as a current picture or a picture to be coded (especially in video coding to distinguish the current picture from other pictures, for example previously coded and / or decoded pictures of the same video sequence, for example a video sequence which also includes the current picture).

[0078] A (digital) picture is or can be considered as a two-dimensional array or matrix of samples with intensity values. The samples in the array may also be called picture elements (an abbreviated form of picture element) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used, e.g. a picture may represent or contain three sample arrays. In an RBG format or color space, a picture contains corresponding red, green and blue sample arrays. However, in video coding, each pixel is typically represented in a luma and chrominance format or color space, e.g. YCbCr, which contains a luma component, denoted Y, and two chrominance components, denoted Cb and Cr. The luma (or luma for short) component Y represents the luminance or gray level intensity (e.g. grayscale pictures), and the two chrominance (or chroma for short) components Cb and Cr represent the chrominance or color information components. Thus, a picture in YCbCr format contains a luma sample array of luma sample values ​​(Y) and two chrominance sample arrays of chrominance values ​​(Cb and Cr). A picture in RGB format may be converted or transformed to YCbCr format and vice versa, a process also known as color conversion or transformation. If the picture is monochrome, the picture may contain only a luma sample array. Thus, a picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0079] An embodiment of the video encoder 20 may include a picture partitioning unit (not shown in FIG. 2 ) configured to partition a picture 17 into multiple (typically non-overlapping) picture blocks 203. These blocks may also be called root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size for all pictures of a video sequence and a corresponding grid defining the block sizes, or to vary the block size between pictures, subsets or groups of pictures, and to divide each picture into the corresponding blocks.

[0080] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, several or all of the blocks forming picture 17. Picture block 203 may also be referred to as a current picture block or a picture block to be coded.

[0081] Similar to the picture 17, the picture block 203 is or is considered to be a two-dimensional array or matrix of samples with intensity values ​​(sample values), although again with smaller dimensions than the picture 17. In other words, the block 203 may contain, for example, one sample array (e.g., a luma array in case of a monochrome picture 17, or a luma or chroma array in case of a color picture) or three sample arrays (e.g., a luma and two chroma arrays in case of a color picture 17) or any other number and / or type of arrays depending on the color format applied. The orientation of the block 203 and the number of samples in the vertical direction (or axis) define the size of the block 203. Thus, the block may be, for example, an MxN (M columns by N rows) array of samples, or an MxN array of transform coefficients.

[0082] The embodiment of video encoder 20 shown in FIG. 2 may be configured to encode picture 17 on a block-by-block basis, eg, encoding and prediction are performed on each block 203 .

[0083] Residual calculation

[0084] The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be described later) on a sample-by-sample (pixel-by-pixel) basis, for example by subtracting sample values ​​of the prediction block 265 from sample values ​​of the picture block 203, to obtain the residual block 205 in the sample domain.

[0085] conversion

[0086] The transform processing unit 206 may be configured to apply a transform, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values ​​of the residual block 205 to obtain transform coefficients 207 in a transform domain. The transform coefficients 207 may also be referred to as transformed residual coefficients, and may represent the residual block 205 in the transform domain.

[0087] Transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as the transform specified for H.265 / HEVC. Compared to an orthogonal DCT transform, such an integer approximation is typically scaled by a particular factor. In order to preserve the norm of the residual blocks processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transform process. The scaling factor is typically selected based on certain constraints, such as scaling factors that are powers of two for shift operations, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. A particular scaling factor may be specified for the inverse transform, e.g., by inverse transform processing unit 212 (and a corresponding inverse transform, e.g., by inverse transform processing unit 312 in video decoder 30), and a corresponding scaling factor for the forward transform, e.g., by transform processing unit 206 in encoder 20, may be specified accordingly.

[0088] An embodiment of video encoder 20 (respectively, transform processing unit 206) may be configured to output the transform parameters, e.g., directly or encoded or compressed, e.g., via entropy coding unit 270, so that, e.g., video decoder 30 may receive and use the transform parameters for decoding.

[0089] Quantization

[0090] The quantization unit 208 may be configured to quantize the transform coefficients 207, for example by applying color quantization or vector quantization, to obtain quantized coefficients 209. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.

[0091] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, in scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization and a larger quantization step size corresponds to coarser quantization. The applicable quantization step sizes may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index to a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size) and a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may include division by the quantization step size, and corresponding and / or inverse dequantization, e.g., by the inverse quantization unit 210, may include multiplication by the quantization step size. Some standards, e.g., HEVC, embodiments may be configured to use a quantization parameter to determine the quantization step size. In general, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified due to the scaling used in the fixed-point approximation of the quantization step size and quantization parameter equations. In one embodiment, the scaling of the inverse transform and dequantization may be combined. Alternatively, customized quantization tables may be used and signaled from the encoder to the decoder, e.g., in the bitstream. Quantization is a lossy operation where the loss increases with increasing quantization step size.An embodiment of video encoder 20 (respectively, quantization unit 208) may be configured to output the encoded quantization parameters, e.g., directly or via entropy encoding unit 270, such that video decoder 30, for example, may receive and apply the quantization parameters for decoding.

[0092] inverse quantization

[0093] The inverse quantization unit 210 is configured to apply the inverse quantization of the quantization unit 208 to the quantized coefficients, e.g., by applying the inverse of the quantization scheme applied by the quantization unit 208, based on or using the same quantization step size as the quantization unit 208, to obtain dequantized coefficients 211. The dequantized coefficients 211 are also referred to as dequantized residual coefficients 211 and typically correspond to the transform coefficients 207, although they are not identical to the transform coefficients due to losses due to quantization.

[0094] Reverse transformation

[0095] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT), an inverse discrete sine transform, or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 is also referred to as a transform block 213.

[0096] Reconstruction

[0097] The reconstruction unit 214 (e.g. an adder or summer 214) is configured to add the transform block 213 (the reconstructed residual block 213) to the prediction block 365, e.g. by adding sample values ​​of the reconstructed residual block 213 and sample values ​​of the prediction block 265 sample-by-sample to obtain a reconstructed block 215 in the sample domain.

[0098] filtering

[0099] The loop filter unit 220 (or “loop filter” 220 for short) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or in general, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is illustrated as an in-loop filter in FIG. 2, in other configurations the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstructed block 221. Embodiments of video encoder 20 (respectively, loop filter unit 220) may be configured to output loop filter parameters, e.g., encoded directly or via entropy encoding unit 270, such that, e.g., decoder 30 may receive and apply the same loop filter parameters or the respective loop filters for decoding.

[0100] Decoded Picture Buffer

[0101] The decoded picture buffer (DRB) 230 may be a memory that stores reference pictures or, in general, reference picture data for encoding the video data by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memories, including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks, e.g., previously reconstructed and filtered blocks 221, of the same current picture or a different picture, e.g., a previously reconstructed picture, and / or may provide a fully previously reconstructed, e.g., decoded picture (and corresponding reference blocks and samples)) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), e.g., for inter prediction. The decoded picture buffer (DRB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, or in general, unfiltered reconstructed samples, e.g., if the reconstructed blocks 215 have not been filtered by the loop filter unit 220 or are any other further processed version of a reconstructed block or sample.

[0102] Mode Selection (Partitioning & Prediction)

[0103] The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive original picture data, e.g., the original block 203 (current block 203 of current picture 17), and reconstructed picture data, e.g., filtered or unfiltered, of the same (current) picture, and / or one or more previously decoded pictures, e.g., reconstructed samples or blocks from the decoded picture buffer 230 or another buffer (e.g., line buffer, not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter prediction or intra prediction, to obtain a prediction block 265 or predictor 265.

[0104] The mode selection unit 260 may be configured to determine or select a partitioning for the current block prediction mode (including no partitioning) and a prediction mode (e.g., intra or inter prediction mode) and generate a corresponding prediction block 265 used for calculation of the residual block 205 and reconstruction of the reconstructed block 215.

[0105] An embodiment of the mode selection unit 260 may be configured to select a partitioning and prediction mode (e.g., from those supported or available by the mode selection unit 260) that provides the best match, in other words, the one that provides the smallest residual (smallest residual means better compression for transmission or storage), or the smallest signaling overhead (smallest signaling means better compression for transmission or storage), or a consideration or balance of both. The mode selection unit 260 may be configured to determine the partitioning and prediction mode based on rate-distortion optimization (RDO), e.g., to select a prediction mode that provides the smallest rate-distortion. Terms such as "best," "lowest," "optimal" in this context do not necessarily refer to an overall "best," "lowest," "optimal," etc., but may refer to a termination criterion or selection criteria, such as, for example, a value above or below a threshold, or the fulfillment of other constraints that may lead to a "suboptimal selection" but reduce complexity and processing time.

[0106] In other words, the partitioning unit 262 may be configured to partition the block 203 into smaller block partitions or sub-blocks (again forming a block), e.g., using quad tree partitioning (QT), binary partitioning (BT), triple tree partitioning (TT) or any combination thereof in an iterative manner, and to perform, e.g., prediction for each block partition or sub-block, wherein the mode selection includes selection of a tree structure of the partitioned block 203, and a prediction mode is applied to each of the block partitions or sub-blocks.

[0107] The partitioning (eg, by partitioning unit 260) and prediction processes (eg, by inter prediction unit 244 and intra prediction unit 254) performed by exemplary video encoder 20 are described in more detail below.

[0108] Partitioning

[0109] The partitioning unit 262 may partition (or split) the current block 203 into smaller partitions, e.g., smaller blocks of square or rectangular size. These smaller blocks (also called sub-blocks) may be partitioned into even smaller partitions. This is also called tree partitioning or hierarchical tree partitioning, where a root block, e.g., root tree level 0 (hierarchical level 0, depth 0), may be partitioned recursively into two or more blocks at the next lower tree level, e.g., nodes at tree level 1 (hierarchical level 1, depth 1), which may be partitioned again into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), until the partitioning is terminated, e.g., because a termination criterion is met, e.g., a maximum tree depth or a minimum block size is reached. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree using a partitioning into two partitions is called a binary tree (BT), a tree using a partitioning into three partitions is called a ternary tree (TT), and a tree using a partitioning into four partitions is called a quad tree (QT). As mentioned above, as used herein, the term "block" may refer to a portion of a picture, particularly a square or rectangular portion. For example, with reference to HEVC and VVC, a block may correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit, and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).

[0110] For example, a coding tree unit (CTU) may be or may include a CTB of luma samples, two corresponding CTBs of chroma samples of a picture with three sample arrays, or a CTB of samples of a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding tree block (CTB) may be an NxN block of samples of N of a value such that the division of a component into CTBs is a partitioning. A coding unit (CU) may be or may include a CTB of luma samples, two corresponding coding blocks of chroma samples of a picture with three sample arrays, or a coding block of samples of a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding block (CB) may be an NxN block of samples of M and N sample values ​​such that the division of a CTB into coding blocks is a partitioning.

[0111] In an embodiment, for example according to HEVC, coding tree units (CTUs) may be split into CUs by using a quad-tree structure, denoted as coding tree. The decision whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU may be further split into one, two or four PUs according to a PU split type. Within one PU, the same prediction process is applied and related information is transmitted to the decoder on a PU basis. After obtaining the residual blocks by applying a prediction process based on the PU split type, the CUs may be partitioned into transform units (TUs) according to another quad-tree structure similar to the coding tree of the CU. In an embodiment, for example according to the latest video coding standard currently under development, called Generic Video Coding (VVC), a quad-tree and binary tree (QTBT) partitioning is used to partition the coding blocks, for example. In the QTBT block structure, the CUs may have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quad-tree structure. The quad-tree leaf nodes are further partitioned by a binary tree or a ternary (or triple) tree structure. The partitioning tree leaf nodes are called coding units (CUs) and the segments are used for prediction and transform processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. It has also been proposed that multiple partitions, e.g. triple-tree partitions, in parallel, can be used with the QTBT block structure.

[0112] In one example, mode selection unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.

[0113] As mentioned above, video encoder 20 is configured to determine or select a best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, intra prediction modes and / or inter prediction modes.

[0114] Intra prediction

[0115] The set of intra prediction modes may include 35 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, e.g., as defined in HEVC, or may include 67 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, e.g., as defined for VVC.

[0116] The intra prediction unit 254 is configured to generate an intra prediction block 265 using reconstructed samples of neighboring blocks of the same current picture according to an intra prediction mode of the set of intra prediction modes.

[0117] The intra prediction unit 254 (or, generally, the mode selection unit 260) is further configured to output intra prediction parameters (or, generally, information indicating the selected intra prediction mode for the block) to the entropy coding unit 270 in the form of a syntax element 266 for inclusion in the encoded picture data 21, so that, for example, the video decoder 30 can receive and use the prediction parameters for decoding.

[0118] Inter Prediction

[0119] The set of inter predictions (or possible inter prediction modes) depends on the available reference pictures (e.g., pictures that are at least partially decoded, e.g., stored in DBP230) and other inter prediction parameters, e.g., whether the entire reference picture or only a part of the reference picture, e.g., a search window area around the area of ​​the current block, is used to search for the best matching reference block, and / or, e.g., whether pixel interpolation, e.g., half / semi-pel and / or quarter-pel interpolation, is applied. In addition to the above prediction modes, skip mode and / or direct mode may be applied. The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain the picture block 203 (current picture block 203 of current picture 17) and the decoded picture 231, or at least one or more previously reconstructed blocks, e.g. reconstructed blocks of one or more other / different previously decoded pictures 231, for motion estimation. For example, a video sequence may include the current picture and the previously decoded picture 231, in other words the current picture and the previously decoded picture 231 may be part of pictures forming a video sequence or may form a sequence of pictures.

[0120] The encoder 20 may be configured, for example, to select a reference block from multiple reference blocks of the same or different pictures from multiple other pictures and provide the reference picture (or reference picture index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-prediction parameter to the motion estimation unit, where this offset is also called a motion vector (MV).

[0121] The motion compensation unit is configured to obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit may involve fetching or generating a prediction block based on the motion / block vectors determined by motion estimation, possibly performing interpolation up to sub-pixel accuracy. Interpolation filtering can generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code the picture block. Upon receiving the motion vector for the PU of the current picture block, the motion compensation unit may locate the prediction block to which the motion vector points in one of the reference picture lists.

[0122] The motion compensation unit may also generate syntax elements associated with the blocks and video slices for use by video decoder 30 in decoding picture blocks of the video slices.

[0123] Entropy Coding

[0124] The entropy encoding unit 270 may, for example, apply an entropy encoding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context-adaptive VLC scheme (CAVLC), an arithmetic coding scheme, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding methodology or technique) or bypass (non-compression) to the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements to obtain encoded picture data 21, which may be output via output 272, for example, in the form of an encoded bitstream 21, so that, for example, a video decoder 30 can receive and use the parameters for decoding. The encoded bitstream 21 may be transmitted to the video decoder 30 or may be stored in a memory for later transmission or retrieval by the video decoder 30. Other structural variations of the video encoder 20 may be used to encode the video stream. For example, the non-transform-based encoder 20 may directly quantize the residual signal for a particular block or frame without the transform processing unit 206. In another implementation, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.

[0125] Decoder and decoding method

[0126] 3 illustrates an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive picture data 21 (e.g., encoded bitstream 21) encoded by, for example, an encoder 20 to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, for example data representing picture blocks and associated syntax elements of an encoded video slice.

[0127] 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or may include a motion compensation unit. The video decoder 30 may, in some examples, perform a decoding path that is generally the reverse of the encoding path described with respect to the video encoder 100 from FIG. 2.

[0128] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DRB) 230, inter prediction unit 344, and intra prediction unit 354 are also referred to as forming a "built-in decoder" of video encoder 20. Thus, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Thus, the descriptions provided for the respective units and functions of video decoder 30 apply to the respective units and functions of video encoder 20.

[0129] Entropy Decoding

[0130] The entropy decoding unit 304 is configured to analyze the bitstream 21 (or in general, the coded picture data 21) and, for example, perform entropy decoding on the coded picture data 21 to obtain, for example, the quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), for example any or all of inter prediction parameters (e.g. reference picture indexes and motion vectors), intra prediction parameters (e.g. intra prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme, as described with respect to the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter prediction parameters, intra prediction parameters and / or other syntax elements to the mode selection unit 360 and to provide other parameters to other units of the decoder 30. Video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.

[0131] inverse quantization

[0132] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or, in general, information related to inverse quantization) and quantized coefficients from the encoded picture data 21 (e.g., by parsing and / or decoding, e.g., by entropy decoding unit 304) and apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain dequantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter determined by the video encoder 20 for each video block in a video slice to determine the degree of quantization, and thus the degree of inverse quantization to be applied.

[0133] Reverse transformation

[0134] The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213. The transform may be an inverse transform, e.g., an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may be further configured to receive transform parameters or corresponding information from the coded picture data 21 (e.g., by analysis and / or decoding, e.g., by the entropy decoding unit 304) to determine a transform to be applied to the dequantized coefficients 311.

[0135] Reconstruction

[0136] A reconstruction unit 314 (e.g., an adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365, e.g., by adding sample values ​​of the reconstructed residual block 313 and sample values ​​of the prediction block 365, to obtain a reconstructed block 315 in the sample domain.

[0137] filtering

[0138] Loop filter unit 320 (either in the coding loop or after the coding loop) is configured to filter reconstructed block 315 to obtain filtered block 321, e.g., to smooth pixel transitions or otherwise improve video quality. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 320 is illustrated as an in-loop filter in FIG. 3, in other configurations loop filter unit 320 may be implemented as a post-loop filter.

[0139] Decoded Picture Buffer

[0140] The decoded video blocks 321 of the picture are then stored in a decoded picture buffer 330, which stores them as reference pictures for other pictures and / or subsequent motion compensation for output display, respectively. The decoder 30 is configured to output the decoded picture 311, for presentation or viewing to a user, e.g., via an output 312.

[0141] prediction

[0142] The inter prediction unit 344 may be identical to the inter prediction unit 244 (in particular the motion compensation unit), and the intra prediction unit 354 may be identical in function to the inter prediction unit 254 and performs the splitting or partitioning decision and prediction based on partitioning and / or prediction parameters or respective information received from the decoded picture data 21 (e.g. by analysis and / or decoding, e.g. by the entropy decoding unit 304). The mode selection unit 360 may be configured to perform prediction (intra prediction or inter prediction) for each block based on the reconstructed picture, block or respective samples (filtered or unfiltered) to obtain a prediction block 365.

[0143] When the video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of mode selection unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. When the video picture is coded as an inter-coded (e.g., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of mode selection unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from entropy decoding unit 304. For inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may configure the reference frame lists, List 0 and List 1, using a default configuration technique based on the reference pictures stored in DPB 330.

[0144] Mode selection unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and to generate predictive blocks for the current video block to be decoded using the prediction information. For example, mode selection unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to code the video blocks of the video slice, an inter-prediction slice type (e.g., a B slice, a P slice, or a GPB slice), configuration information for one or more of a reference picture list for the slice, a motion vector for each inter-coded video block of the slice, an inter-prediction status for each inter-coded video block of the slice, and other information for decoding the video blocks in the current video slice.

[0145] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may generate an output video stream without a loop filtering unit 320. For example, a non-transform-based decoder 30 may directly inverse quantize the residual signal for a particular block or frame without an inverse transform processing unit 312. In another implementation, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0146] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and output to the next step, for example, after the interpolation filtering, the motion vector derivation, or the loop filtering, further operations such as clipping or shifting may be performed on the processing result of the interpolation filtering, the motion vector derivation, or the loop filtering.

[0147] It should be noted that further operations may be applied to the derived motion vector of the current block (including but not limited to control point motion vector in affine mode, sub-block motion vector in affine, planar, ATMVP mode, temporal motion vector, etc.). For example, the value of a motion vector is constrained to a predefined range according to its representation bit. If the representation bit of a motion vector is bitDepth, the range is -2^(bitDepth-1)~2^(bitDepth-1)-1, where "^" means exponentiation. For example, if bitDepth is set to 16, the range is -32768~32767, and if bitDepth is set to 18, the range is -131072~131071. Here, we provide two methods to constrain the motion vector.

[0148] Method 1: Remove the overflow MSB (Most Significant Bit) using flow arithmetic.

number

number

[0149] Method 2: Remove the overflow MSB by clipping the value.

number

number

[0150] 4 is a schematic diagram of a video coding device 400 according to one embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 of FIG. 1A, or an encoder, such as the video encoder 20 of FIG. 1A.

[0151] The video coding device 400 includes an ingress port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing the data, a transmitter unit (Tx) 440 and an egress port 450 (or output port 450) for transmitting the data, and a memory 460 for storing the data. The video coding device 400 may also include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the ingress port 410, the receiver unit 420, the transmitter unit 440, and the egress port 450 for the entry and exit of optical or electrical signals.

[0152] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 is in communication with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 470 provides a substantial improvement to the functionality of the video coding device 400, resulting in the transformation of the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.

[0153] Memory 460 may include one or more disks, tape drives, and solid state drives, may be used as overflow data storage devices, store programs when such programs are selected for execution, and store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (RSAM).

[0154] FIG. 5 is a simplified block diagram of an apparatus 500 that may be used as either or both of the source device 12 and destination device 14 from FIG. 1 in accordance with an example embodiment.

[0155] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices capable of manipulating or processing information, now existing or later developed. Although the disclosed implementations may be implemented with a single processor, e.g., processor 502, as shown, multiple processors may be used to achieve advantages in speed and efficiency. The memory 504 in the device 500 may be a read-only memory (ROM) device or a random access memory (RAM) device in the implementation. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that are accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and an application program 510, which includes at least one program that allows the processor 502 to execute the methods described herein. For example, application programs 510 can include applications 1-N, which further include a video coding application that performs the methods described herein. Apparatus 500 can also further include one or more output devices, such as a display 518. Display 518, in one example, can be a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. Display 518 can be coupled to processor 502 via bus 512.

[0156] Although shown here as a single bus, bus 512 of device 500 may be comprised of multiple buses. Additionally, secondary storage 514 may be directly coupled to other components of device 500 or may be accessed over a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 may be implemented in a wide variety of configurations.

[0157] 18 is a block diagram showing a content supply system 3100 for implementing a content distribution service. The content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any type of combination thereof, etc.

[0158] The capture device 3102 may generate data and encode the data by the encoding method shown in the above embodiment. Alternatively, the capture device 3102 may distribute the data to a streaming server (not shown), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 may include, but is not limited to, a camera, a smartphone or pad, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include a source device 12 as described above. If the data includes video, the video encoder 20 included in the capture device 3102 may actually perform the video encoding process. If the data includes audio (e.g., voice), the audio encoder included in the capture device 3102 may actually perform the audio encoding process. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, e.g., in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 separately delivers the encoded audio data and the encoded video data to the terminal device 3106 .

[0159] In the content supply system 3100, the terminal device 310 receives and plays the encoded data. The terminal device 3106 can be a device having data receiving and recovery capability and capable of decoding the aforementioned encoded data, such as a Smartphone or Pad 3108, a computer or laptop 3110, a Network Video Recorder (NVR) / Digital Video Recorder (DVR) 3112, a TV 3114, a Set Top Box 3116, a video conferencing system 3118, a video surveillance system 3120, a Personal Digital Assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the capture device 3106 may include the source device 14 as described above. If the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. If the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.

[0160] In a terminal device with a display, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant 3122, or an in-vehicle device 3124, the terminal device can provide the decoded data to its display. In a device without a display, such as a STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is contacted thereto to receive and show the decoded data.

[0161] When each device in the system performs encoding or decoding, a picture encoding device or a picture decoding device can be used as shown in the above embodiment. FIG. 19 is a diagram showing an example of the configuration of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progression unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real Time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any kind of combination thereof. After the protocol progression unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, for example, in a video conference system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and audio decoder 3208 without passing through the demultiplexer 3204 .

[0162] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206 including the video decoder 30 described in the previous embodiment decodes the video ES by the decoding method shown in the above embodiment to generate video frames, and transmits the data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames, and transmits the data to a synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. Y) before being supplied to the synchronization unit 3212. Alternatively, the audio frames may be stored in a buffer (not shown in FIG. Y) before being supplied to the synchronization unit 3212.

[0163] The synchronization unit 3212 synchronizes the video and audio frames and provides the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information, which may be coded in a syntax using time stamps for the presentation of the encoded audio and visual data and for the delivery of the data stream itself.

[0164] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.

[0165] In an example for merging candidate list construction according to ITU-T H.265, the merging candidate list is constructed based on the following candidates: 1. Up to four spatial candidate MVPs derived from five spatially adjacent blocks 2. One temporal candidate derived from two temporally co-located blocks; 3. Additional candidates including combined bi-prediction candidates 4. Zero Motion Vector Candidates

[0166] spatial candidate

[0167] The motion information of spatially neighboring blocks is first added to the merge candidate list as a motion information candidate (in one example, the merge candidate list may be an empty list before the first motion vector is added to the merge candidate list). Here, the neighboring blocks considered for insertion into the merge list are shown in Figure 6b. For inter-prediction block merging, up to four candidates are inserted into the merge list by sequentially checking A1, B1, B0, A0, B2 in that order.

[0168] The motion information may include information about whether one or two reference picture lists are used, as well as all the motion data including reference indexes and motion vectors for each reference picture list.

[0169] In one example, after checking whether neighboring blocks are available and contain motion information, some additional redundancy checks are performed before taking all the motion data of the neighboring blocks as motion information candidates.

[0170] These redundancy checks can be split into two categories, serving two different purposes. Category 1 avoids having candidates with redundant motion data in the list. Category 2 prevents merging two partitions that could be expressed by other means producing redundant syntax.

[0171] temporal candidate

[0172] Fig. 6a shows the coordinates of blocks for which temporal motion information candidates are searched. Co-located blocks are blocks that have the same x,y coordinates as the current block, but are in a different picture (one of the reference pictures). Temporal motion information candidates are added to the merge list if the list is not full (in one example, the merge list is not full when the amount of candidates in the merge list is less than a threshold, e.g., the threshold may be 4, 5, 6, etc.).

[0173] Generated candidates

[0174] If the merge list is still not full after insertion of the spatial and temporal motion information candidates, the generated candidates are added to fill the list. The list size is indicated in the sequence parameter set and is fixed throughout the entire coded video sequence.

[0175] The merge list construction process in ITU-T H.265 and VVC outputs a list of motion information candidates. The merge list construction process for VVC is described in section "8.3.2.2 Derivation process for luma motion vectors for merge mode" of document JVET-L1001_v2 Versatile Video Coding (Draft 3), which is publicly available at http: / / phenix.it-sudparis.eu / jvet / . The term motion information refers to the motion data required to perform the motion compensated prediction process. Motion information typically refers to the following information: Whether the block applies uni-prediction or bi-prediction ID of the reference picture used for prediction (two IDs, if the block applies bi-prediction). · Motion vectors (two motion vectors if the block is bi-predicted) ·Additional Information

[0176] In VVC and H.265, the list of candidates that is the output of the merge list construction includes N candidate motion information. The number N is typically included in the bitstream and may be a positive integer, such as 5, 6, etc. The candidates included in the constructed merge list may include uni-prediction or bi-prediction information. This means that the candidates selected from the merge list may exhibit bi-prediction behavior.

[0177] Bi-prediction

[0178] A special mode of inter prediction is called "bi-prediction", in which two motion vectors are used to predict a block. The motion vectors can point to the same or different reference pictures, which can be indicated by a reference picture list ID and a reference picture index. For example, a first motion vector may point to the first picture in reference picture list L0, and a second motion vector may point to the first picture in reference picture list L1. Two reference picture lists (e.g., L0 and L1) may be maintained, where the picture pointed to by the first motion vector is selected from list L0, and the picture pointed to by the second motion vector is selected from list L1.

[0179] In one example, if the motion information indicates bi-prediction, the motion information includes two parts. · L0 part: Motion vectors and reference picture indices pointing to entries in the reference picture list L0. · L1 part: Motion vectors and reference picture indices pointing to entries in the reference picture list L1.

[0180] Picture Order Count (POC): A variable associated with each picture that uniquely identifies the associated picture among all pictures in the Coded Video Sequence (CVS) and indicates the position of the associated picture in the output order relative to the output order positions of other pictures in the same CVS that are output from the decoded picture buffer when the associated picture is output from the decoded picture buffer.

[0181] Each of the reference picture lists L0 and L1 may contain one or more reference pictures, each identified by a POC. The association of each reference index and the POC value may be signaled in the bitstream. As an example, the L0 and L1 reference picture lists may contain the following reference pictures: [Table 2]

[0182] In the above example, the first entry of reference picture list L1 (indicated by reference index 0) is a reference picture with POC value 13. The second entry of reference picture list L1 (indicated by reference index 1) is a reference picture with POC value 14.

[0183] Triangle Prediction Mode

[0184] The concept of triangular prediction mode is to disclose triangular partition for motion compensation prediction. As shown in Fig. 7, two triangular prediction units are used for a CU in either diagonal or anti-diagonal direction. Each triangular prediction unit in a CU is inter predicted using a uni-prediction motion vector and a reference frame index derived from the uni-prediction candidate list. After the samples related to each triangular prediction unit are predicted by motion compensation or intra-picture prediction, an adaptive weighting process is performed on the diagonal edges. Then, a transform and quantization process is applied to the CU. Note that this mode applies to skip and merge modes.

[0185] In triangular prediction mode, a block is split into two triangular parts (as in FIG. 7) and each part may be predicted using one motion vector. The motion vector used to predict one triangular part (denoted as PU1) may be different from the motion vector used to predict the other triangular part (denoted as PU2). Note that in one example, to reduce the complexity of performing the triangular prediction mode, each part may be predicted using a single motion vector (uni-prediction). In other words, PU1 and PU2 may not be predicted using bi-prediction involving two motion vectors.

[0186] Sub-block prediction mode

[0187] Triangular prediction mode is a special case of sub-block prediction. In sub-block prediction, a block is divided into two blocks. In the above example, two block division directions (45 degree and 135 degree partitions) are shown, and other partition angles and partition ratios are also possible (example of FIG. 8).

[0188] In some examples, the block is split into two parts and each part is applied by uni-prediction. Sub-block prediction represents a generalized version of triangular prediction.

[0189] Triangular prediction mode is a special case of sub-block prediction mode in which a block is divided into two blocks. In the above example, two block division directions (45 degree and 135 degree partitions) are shown. Other partition angles and partition ratios are also possible (for example, the example of FIG. 8).

[0190] In some examples, a block is split into two sub-blocks and each part (sub-block) is predicted with uni-prediction.

[0191] In one example, the following steps are applied to obtain prediction samples according to the sub-block partition mode: Step 1: A coding block is considered to be divided into two sub-blocks according to a geometric model. This model may result in a division of the block by a separation line (e.g. a straight line) as illustrated in figures 9 to 12. In this step, according to the geometric model, the samples in the coding block are considered to be located in two sub-blocks. Sub-block A or sub-block B contains some (but not all) of the samples in the current coding block. The notion of sub-block A or sub-block B may be indicated by a parameter related to the separation line. Step 2: Obtain a first prediction mode for the first sub-block and a second prediction mode for the second sub-block. In one example, the first prediction mode is not the same as the second prediction mode. In one example, the prediction mode (the first prediction mode or the second prediction mode) may be an inter prediction mode, and the information for the inter prediction mode may include a reference picture index and a motion vector. In another example, the prediction mode may be an intra prediction mode, and the information for the intra prediction mode may include an intra prediction mode index. Step 3: Generate a first forecast value and a second forecast value using the first forecast mode and the second forecast mode, respectively. Step 4: Obtain a combination value of the predicted sample according to the combination of the first predicted value and the second predicted value. As an example, in step 1, it is considered that the coding block is divided into two sub-blocks in different ways.

[0192] 9 shows an example of a partition of a coding block, where a separation line 1250 divides the block into two sub-blocks. To describe the line 1250, two parameters are signaled, one parameter is the angle alpha 1210 and the other parameter is the distance dist 1230.

[0193] In some embodiments, as shown in FIG. 9, the angle is measured between the x-axis and the separation line, and the distance is measured by the length of a vector perpendicular to the separation line and passing through the center of the current block. In another example, FIG. 10 shows another way to represent the separation line, where the example angle and distance differs from the example shown in FIG. 9. In some examples, in step 4, the division disclosed in step 1 is used in a combination of the first prediction and the second prediction to obtain a final prediction. In one example, a blending operation is applied in step 4 to remove any artifacts (edgy or jagged appearance along the separation line). The blending operation may be described as a filtering operation along the separation line.

[0194] At the encoder side, the separating line (parameters defining the line, e.g., angle and distance) are determined based on a rate-distortion based cost function. The determined line parameters are coded into a bitstream. At the decoder side, the line parameters are decoded (obtained) according to the bitstream.

[0195] In one example, for three video channels including a luma component and two chroma components, a first prediction and a second prediction are generated for each channel.

[0196] Because there are many possibilities to split a coding block into two sub-blocks, signaling (coding) the split requires too many bits. Because the angle and distance values ​​have many different values ​​that would require too much side information to be signaled in the bitstream, a quantization scheme is applied to the angle and distance side information to improve coding efficiency.

[0197] In one example, a quantized angle parameter alphaIdx and a quantized distance parameter distanceIdx are signaled for the partitioning parameters.

[0198] In one example quantization scheme, the angle and distance values ​​may be quantized by a linear uniform quantizer according to the following:

number

number

[0199] In some examples, a method is disclosed for partitioning a rectangular coding block by a line, where the line is parameterized by a pair of parameters representing a quantized angle and a quantized distance value, the quantized distance value being derived according to a quantization process depending on the angle value and the aspect ratio of the coding block.

[0200] In one example, the distance may be quantized such that a given value range of distanceIdx is filled, e.g., a value range of 0 to 3 inclusive. In another example, the distance may be quantized for a given coding block such that the separation lines for a given angleIdx and distanceIdx are evenly distributed and no separation lines lie outside the region of a given coding block.

[0201] In the first step, the maximum distance ρ max can be derived depending on the angle, and the distance 0 <dist<ρ max , so that all separation lines of are confined to the coding block (i.e., intersect with a coding block boundary). This is illustrated in Figure 15 for a coding block of size 16x4 luma samples.

[0202] In one example, the maximum distance ρ max can be derived as a function that depends on the angle alphaR and the size of the coding block according to:

number

number

[0203] In another example, the value of the angle-dependent distance quantization step size may be derived as follows:

number

[0204] In another example, the maximum distance ρmax may be derived as a function that depends on the angle and size of the coding block according to:

number

[0205] In one example, the value of Δdist depends on the value representing the angle, the parameter representing the width and the parameter representing the height may be stored in a pre-computed look-up table to avoid repeated calculation of Δdist during the encoding or decoding process.

[0206] In one example, the value of Δdist may be scaled and rounded for purposes of using integer arithmetic as follows:

number

[0207] In one example, it is to further store a pre-calculated value of stepD based on an aspect ratio denoted whRatio, which depends on the width and height of the coding block. Furthermore, the pre-calculated value of stepD is stored based on a (normalized) angle value angleN, which is an index value related to the angle of the first quadrant of the Euclidean plane (e.g., 0≦angleN*Δalpha≦90°). An example of the application of such a look-up table is shown in FIG. 13. In one example, the following steps are applied to obtain predicted values ​​for the samples of the coding block:

[0208] Step 1: For samples in the current coding block (decoded block or encoded block), the sample distance (sample_dist) is calculated.

[0209] In some examples, the sample distance may represent the horizontal or vertical distance of the sample to the separation line (the separation line is used to indicate that the coding block is divided into two sub-blocks), or a combination of the vertical and horizontal distances. The sample is represented by its coordinates (x,y) relative to the top left sample of the coding block. Sample coordinates and sample_dist are illustrated in FIG. 11 and FIG. 12. The sub-blocks are not necessarily rectangular, but may have a triangular or trapezoidal shape. In one example, the first parameter represents a quantized angle value (angleIdx) and the second parameter represents a quantized distance value (distanceIdx). The two parameters represent a line equation. In one example, the distance 1230 may be obtained according to distanceIdx (the second parameter), and the angle alpha (1210) may be obtained according to angleIdx (the first parameter). The distance 1230 may be the distance to the center of the coding block, and the angle may be the angle between the separation line and a horizontal (or equivalently vertical) line passing through the center point of the coding block.

[0210] In one example, in step 1, a coding block is considered to be partitioned into two sub-blocks in various ways. Figure 9 shows an example of a partition of a coding block, where a separating line 1250 is used to indicate that the block is divided into two sub-blocks. To account for the line 1250, one parameter angle alpha 1210 is signaled in the bitstream.

[0211] In some embodiments, the angle is measured between the x-axis and the separation line, and the distance is measured by the length of a vector that is perpendicular to the separation line and passes through the center of the current block, as shown in Figure 9. In another example, Figure 10 shows another way to represent the separation line, with example angles and distances that differ from the example shown in Figure 9.

[0212] Step 2: Calculate weighting coefficients using the calculated sample_dist, which are used for combining the first and second predicted values ​​corresponding to the sample. In one example, the weighting coefficients are denoted as sampleWeight1 and sampleWeight2, which refer to the weights corresponding to the first and second predicted values.

[0213] Step 3: The sum of the predicted sample at the sample coordinate (x, y) is calculated according to the first predicted value at the coordinate (x, y), the second predicted value at the coordinate (x, y), sampleWeight1 and sampleWeight2.

[0214] In one example, Step 1 in the above example may include the following steps:

[0215] Step 1.1: Get the index value of the angle parameter (alphaN) for the current block, the width value of the current block (W) and the height value of the current block (H). W and H are the width and height of the current block in number of samples. For example, a coding block with both width and height equal to 8 is a square block containing 64 samples. In another example, W and H are the width and height of the current block in number of luma samples.

[0216] Step 1.2: According to the value of W and the value of H, obtain the value of the ratio whRatio, which represents the ratio of the width and height of the current coding block.

[0217] Step 1.3: Obtain the stepD value according to the lookup table, the alpha value, and the whRatio value. In one example, as shown in FIG. 13, the alpha value and the whRatio value are used as index values ​​of the lookup table.

[0218] Step 1.4: The value of sample_dist is calculated according to the value of stepD.

[0219] In another example, Step 1 in the above example may include the following steps.

[0220] Step 1.1: Get the angle parameter (alphaN) for the current block, the distance index value (distanceIdx), the width value of the current block (W), and the height value of the current block (H).

[0221] Step 1.2: According to the value of W and the value of H, obtain the value of the ratio whRatio, which represents the ratio of the width and height of the current coding block.

[0222] Step 1.3: According to the lookup table, the value of alpha and the value of whRatio, a stepD value is obtained, in one example, the value of alphaN and the value of whRatio are used as index values ​​of the lookup table, as shown in Figure 13. In one example, the stepD value represents the quantization step size for the sample distance calculation process.

[0223] Step 1.4: The value of sample_dist is calculated according to the value of stepD, the value of distanceIdx, the value of angle (alphaN), the value of W, and the value of H. In one example, the value of whRatio is obtained using the formula:

number

[0224] In another example, the value of whRatio is calculated as whRatio=(W>=H)?W / H:H / W. In one example, the value of the angle alpha can be obtained from the bitstream (at the decoder). In one example, the value range of the angle is a quantized value range of 0 to 31 (inclusive), denoted as angleIdx. In one example, the quantized angle value can only take 32 different values ​​(so the values ​​0 to 31 are sufficient to represent which angle value is selected). In another example, the value range of the angle value can be between 0 and 15, which means that 16 different quantized angle values ​​can be selected. Note that in general, the angle value can be an integer value equal to or greater than zero.

[0225] In one example, the value of alphaN is an index value obtained from the bitstream, or the value of alpha is calculated based on the value of the indicator obtained from the bitstream. For example, the value of alphaN may be calculated according to the following formula:

number

[0226] In another example, the value alphaN may be calculated according to one of the following formulas:

number

[0227] In one example, the value of sample_dist is obtained according to the following formula:

number

[0228] In one example, functions f1() and f2() are implemented as lookup tables. In one example, functions f1() and f2() represent incremental changes in the value of sample_dist for changes in the values ​​of x and y. In some examples, f1(index) represents that the value of sample_dist changes in increments of 1 unit in the value of x (where a unit may be an increment equal to 1), while f2(index) represents that the value of sample_dist changes in increments of 1 unit in the value of y. The index values ​​may be obtained from the values ​​of indicators in the bitstream.

[0229] In another example, the value of sample_dist is obtained according to the formula:

number

[0230] In one example, function f() is implemented as a lookup table. Function f() represents the incremental change in the value of sample_dist for changes in the values ​​of x and y. In the example, f(index1) represents that the value of sample_dist changes with an increment of one unit in the value of x, and f(index2) represents that the value of sample_dist changes with an increment of one unit in the value of y. The values ​​of index1 and index2 are indices into the table (with integer values ​​equal to or greater than 0) that can be obtained according to the value of the indicator in the bitstream.

[0231] In one example, the implementation of function f() is shown in Figure 14. In this example, the value of idx is the input parameter (which can be index1 or index2) and the output of the function is denoted as f(idx). In one example, f() is an implementation form of the cosine function using integer arithmetic, and idx (the input index value) represents the quantized angle value.

[0232] In one embodiment, the stepD value represents a quantized distance value for the sample distance calculation.

[0233] In one example, the value of stepD is obtained according to the value of whRatio and the value of angle (alpha), as shown in Figure 13. In one example, the value of stepD can be obtained as stepD = lookupTable [alphaN] [whRatio], where the value of alphaN is an index value obtained from the bitstream, or the value of alphaN is calculated based on the value of the indicator obtained from the bitstream. For example, alpha can be calculated according to the following formula:

number

[0234] In another example,

number

number

[0235] In one example, the value of sample_dist is obtained according to distanceIdx*stepD*scaleStep, where distanceIdx is an index value obtained according to the bitstream and the value of scaleStep is obtained according to either the block width value or the block height value. The result of the multiplication represents the distance of the separation line to the center point of the coding block (having coordinates x=W / 2 and y=H / 2).

[0236] In one example, the lookup table is a predefined table. A predefined table has the following advantages: Obtaining the distance to the sample separation line is usually complex and requires solving trigonometric equations that are not acceptable when implementing video coding standards targeted at mass-produced consumer products.

[0237] In some examples, the sample distances are obtained according to a lookup table (which may be predefined), which pre-calculates intermediate results according to whRatio and alpha, which are already according to integer arithmetic (so in one example, all stepD values ​​are integers). The intermediate results obtained using the lookup table are carefully selected for the following reasons: The lookup table contains intermediate calculation results (trigonometric calculations) for complex operations, reducing the implementation complexity. · The size of the table is kept small (this requires memory).

[0238] In another example, the value of sample_dist is obtained according to distanceIdx*(stepD+T)*scaleStep, where T is an offset value having an integer value. In one example, the value of T is 32.

[0239] In another example (from both the decoder and encoder perspective): The value of sample_dist in the above example indicates the distance between the sample in the current coding sub-block and the separation line, as shown in FIG. 11 or FIG.

[0240] The two sub-blocks are indexed with 0 and 1. The first sub-block is indexed with 0 and the second sub-block is indexed with 1.

[0241] In the geometric partition mode, the processing order of two sub-blocks is determined based on the sub-block index. The sub-block with index 0 will be processed first, followed by the sub-block with index 1. The processing order may be used in the sample weight derivation process, the motion information storage process, the motion vector derivation process, etc.

[0242] In one example, the index value of a subblock is determined based on the sample_dist values ​​of the samples in the subblock.

[0243] If the sample_dist value of each sample in a subblock is less than (or equal to) zero, then the index value for the subblock is set to zero.

[0244] If the sample_dist value of each sample in a subblock is greater than (or equal to) zero, then the index value for the subblock is set to one.

[0245] In the same example, the index values ​​for the sub-blocks may be modified (eg, inverted).

[0246] In one example, the index values ​​of the sub-blocks may be modified based on predefined sample positions.

[0247] In one example, the index value for a sub-block located at a predefined position of the coding block (e.g., the bottom left position) is always set as 0.

[0248] In one example, if the predefined sample position belongs to a second sub-block having an index value of 1, the index value for this sub-block will be modified to 0 (the index values ​​for the other sub-blocks will be modified to 1).

[0249] In one example, if the predefined sample position belongs to the first sub-block with index 0, the sub-block index is not changed.

[0250] For the example shown in Figure 20 and Figure 21. The predefined sample position is the bottom left sample position of the current coding block.

[0251] In FIG. 20, the bottom left sample belongs to the first sub-block with index 0, and the sub-block indices are unchanged.

[0252] In Figure 21, the bottom left sample belongs to the second sub-block with index 1 (Figure 21a), therefore the indices of the two sub-blocks will be swapped or inverted as shown in Figure 21b, and the process order will be based on the swapped or inverted indices.

[0253] In one example, the index assignment and inversion process may be performed in the following steps. Step 1: Split the current coding block into two sub-blocks based on the split mode. Step 2: Calculate the sample_dist value for the samples in one sub-block.

[0254] If the sample_dist value is negative or equal to zero, the index value for this subblock is set to 0 (corresponding to the first subblock) and the index values ​​for the other subblocks are set to 1 (corresponding to the second subblock). or, If the sample_dist value is equal to a positive value, the index value for this subblock is set to 1 (corresponding to the second subblock) and the index values ​​for the other subblocks are set to 0 (corresponding to the first subblock).

[0255] Step 3: If the sample_dist value of the predefined sample position is a positive value, exchange the indexes of the two sub-blocks.

[0256] If the sample_dist value for a defined sample position is negative, the indices of the two sub-blocks are not changed.

[0257] In one embodiment of the present invention, a method of coding implemented by a decoding or encoding device is disclosed.

[0258] The method includes the following.

[0259] S101: Obtain a division (or partition) mode for the current coding block.

[0260] The partition mode is used to split the coding block into sub-blocks. The partition mode may be an example of a defined format having an angle and / or a distance. The angle with respect to the horizontal direction is described as a direction perpendicular to the separation line. The separation line is used to indicate that the current coding block is split into two sub-blocks. The distance describes the distance from the center position of the current coding block to the separation line. In some examples, the partition mode is represented according to a value of an indicator in the bitstream. The value of the indicator is used to obtain an angle parameter value and a distance parameter value. According to one example, the pair (first parameter (angle), second parameter (distance)) may be derived from a predefined method or function. In other words, the partition mode indicator is parsed from the decoder bitstream. The indicator is defined as an input to the method or function. The output of the method or function is the pair (first parameter and second parameter).

[0261] An example of an indicator (geo_partition_idx) and a predefined method or function is as follows:

[0262] Function 1 Indicator X is an input (0 to 63).

number

[0263] Another example of the indicator (geo_partition_idx) and the predefined method or function is as follows:

[0264] Function 2 Indicator X is an input (0 to 63).

number

[0265] S102: The current coding block is considered to be split into two sub-blocks (sub-block A, sub-block B) based on a split mode.

[0266] In this step, according to the partitioning mode, the samples in the coding block are considered to be located in two sub-blocks. Sub-block A or sub-block B includes some (but not all) of the samples in the current coding block. Sub-block A or sub-block B may be represented according to the sign of sample_dist of each sample. sample_dist may be obtained according to the above examples and embodiments.

[0267] In the same example, this step can refer to the above-mentioned embodiment or example regarding the division of the coding block into two sub-blocks. For example, the coding block may be divided into two sub-blocks according to a geometric model. This model may result in the division of the block by a separation line (e.g., a straight line), as illustrated in Figures 9 to 12.

[0268] 9 shows an example for a partition of a coding block, where the coding block is considered to be divided into two sub-blocks according to a separation line 1250. To explain the line 1250, two parameters are signaled, one parameter being the angle alpha 1210 and the other parameter being the distance dist 1230.

[0269] In some embodiments, as shown in FIG. 9, the angle is measured between the x-axis and the separation line, and the distance is measured by the length of a vector perpendicular to the separation line and passing through the center of the current block.

[0270] In another example, Fig. 10 shows another way to represent the separation line, and the example of angle and distance is different from the example shown in Fig. 9. At the encoder side, the separation line (parameters defining the line, e.g., angle and distance) are determined based on a rate-distortion based cost function. The determined line parameters are coded into a bitstream. At the decoder side, the line parameters are decoded (obtained) according to the bitstream.

[0271] One example is the case of three video channels containing a luminance component (luma) and two chrominance components (chroma). This splitting process may be performed on either the luma or the chroma.

[0272] In one embodiment, for a sample in a coding block, the distance value is used to determine in which subblock the sample is located (subblock A or subblock B). In one example, a distance value is obtained for each sample in a coding block.

[0273] In one example, when the distance value corresponding to a sample is less than or equal to 0, the sample is within subblock A (or subblock B), and when the distance value corresponding to a sample is greater than 0, the sample is within subblock B (or subblock A).

[0274] In another example, when the distance value corresponding to a sample is less than 0, the sample is within subblock A (or subblock B), and when the distance value corresponding to a sample is greater than or equal to 0, the sample is within subblock B (or subblock A).

[0275] The process of obtaining the distance value (eg, sample_dist) corresponding to the sample may refer to the above examples and embodiments, for example, sample_dist represents the distance value corresponding to the sample.

[0276] S103: Set the index value of sub-block A according to a predefined sample position of the current coding block.

[0277] In some examples, the predetermined sample position may be a bottom-left, bottom-right, top-left, or top-right position, or a center position of the current coding block.

[0278] When the predefined sample position is located within subblock A, the index value of subblock A will be set equal to a first value, and in this example, the index value of subblock B will be set equal to a second value. or, When the predefined sample position is located in subblock A (e.g., when the predefined sample position is located in subblock B), the index value of subblock A will be set equal to a second value, and in this example, the index value of subblock B will be set equal to a first value. The processing order corresponding to the first value is before the processing order corresponding to the second value, and the subblock in which the predefined sample position is located will always be processed before other subblocks. The processing may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. In some examples, in a geometric partition mode, the processing order of two subblocks is determined based on the indexes of the subblocks. The subblock with index 0 will be processed first, followed by the subblock with index 1.

[0279] The first value or the second value may be an integer value. In one example, the first value may be 0 or 1, or other value, and the second value may be 0 or 1, or other value. The first value and the second value are different.

[0280] In one example, whether a predefined sample position is located within subblock A may be determined based on a sample distance value (e.g., sample_dist) for the predefined sample position. For example, if a sample distance value corresponding to a sample in subblock A is greater than 0 and a sample distance value corresponding to the predefined sample position is less than 0, then the given sample position is not within subblock A but is within subblock B. In this example, each sample in subblock B has a sample distance value less than 0. In some other examples, each sample in subblock B may have a sample distance value less than or equal to 0, depending on whether a sample corresponding to a sample distance value equal to 0 in subblock B is considered. In another example, whether a predefined sample position is located within subblock A may be determined based on a coordinate value for the predefined sample position.

[0281] EMBODIMENT 2

[0282] In another embodiment, the sub-block or block index is set by a predefined function or look-up table, the input values ​​of the predefined function or look-up table may be variable or parameter values, and the output of the predefined or look-up table is the sub-block or block index value.

[0283] In one example, the index of a block may be a split index denoted as isFlip, where the index value isFlip satisfies:

number

[0284] In one example, input values ​​for a predefined function or look-up table may be parsed from the bitstream.

[0285] In one example, input values ​​for a predefined function or look-up table may be derived based on the value of the indication information.

[0286] In one example, the input value is the angle index value of the division boundary, which is derived from the syntax geo_partition_idx.

[0287] In one example, the index value for subblock A of the geometric partition may be derived from the following lookup table:

[0288] If the value of partIdx[AngleIdx] is equal to 0, it means that the part index of the sub-block A split off from the geometric partition is 0, and if the value of partIdx[AngleIdx] is equal to 1, it means that the part index of the sub-block A split off from the geometric partition is 1. [Table 5]

[0289] Table 4 can be converted into a function or equation form as follows:

number

[0290] In one example, subblock A and subblock B are split from one geometric partition, so that subblock B will have the opposite index of subblock A (e.g., value 1 corresponding to value 0, or value 0 corresponding to value 1).

[0291] In some examples, the processing order corresponding to the first value (0 in the above example) is before the processing order corresponding to the second value (1 in the above example). Thus, the sub-block in which a given sample position is located is always processed before the other sub-block. The processing may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. In some examples, in a geometric partition mode, the processing order of two sub-blocks is determined based on the indexes of the sub-blocks. The sub-block with index 0 will be processed first, followed by the sub-block with index 1.

[0292] In some examples, in the geometric partition mode, the predictor values ​​are equal to predictorA*Mask1+predictorB*Mask2, where the predictorA values ​​or the predictorB values ​​may be derived from other modes (e.g., merge mode), and Mask1 and Mask2 are blending masks for the subblock splits that form the geometric partition. The blending masks are pre-computed based on other indicators.

[0293] In one example, angleIdx and distanceIdx are derived from geo_partition_idx, which is parsed from the bitstream.

[0294] The distance between the sample position and the division boundary is calculated based on angleIdx and distanceIdx, and the blending mask depends on this distance using a predefined lookup table.

[0295] In one example, an index value (e.g., 0 or 1) of a subblock split off from a geometric partition may be used for indicator mask selection for the subblock.

[0296] In one example, if the index value of a subblock is a first value (e.g., 0), Mask1 is used for this subblock to generate a predictor value, and if the index value of a subblock is a second value (e.g., 1), Mask2 is used for this subblock to generate a predictor value.

[0297] In one example, if sub-block A has an index value of 0, then Mask1 is used for sub-block A and Mask2 is used for sub-block B (in this example, sub-block B has an index value of 1).

number

[0298] In one example, if sub-block A has an index value of 1, then Mask2 is used for sub-block A and Mask1 is used for sub-block B (in this example, sub-block B has an index value of 0).

number

[0299] In its general form it looks like this:

number

[0300] FIG. 22 shows a coding method implemented by a decoding device, the method including:

[0301] S2201: Obtain a division mode index value for the current coding block.

[0302] In one implementation form, the split mode index value is used to indicate which geometric partition mode is used for the current coding block, for example, geo_partition_idx or merge_gpm_partition_idx.

[0303] In one example, merge_gpm_partition_idx[x0][y0] or geo_partition_idx specifies the partitioning shape for the geometric partitioning merge mode. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0304] When merge_gpm_partition_idx[x0][y0] or go_partition_idx[x0][y0] is not present, it is inferred to be equal to 0.

[0305] In one example, the split mode index value may be obtained by parsing an index value coded in a video bitstream, or the split mode index value may be determined according to a syntax value parsed from the video bitstream.

[0306] The bit stream may be obtained according to a wireless or wired network. The bit stream may be transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or infrared, radio, microwave, WIFI, Bluetooth, LTE, or 5G wireless technologies.

[0307] In one embodiment, the bitstream is a sequence of bits, e.g., in the form of a Network Abstraction Layer (NAL) unit stream or a byte stream, forming a representation of a sequence of access units (AUs) forming one or more Coded Video Sequences (CVSs).

[0308] In some embodiments, for a decoding process, a decoder side reads the bitstream and derives the decoded picture from the bitstream, and for an encoding process, an encoder side generates the bitstream.

[0309] Typically, a bitstream will contain syntax elements that are formed by a syntax structure. Syntax element: An element of data that is represented in the bitstream. Syntax structure: zero or more syntax elements that occur together in a bitstream in a specified order.

[0310] In a particular example, a bitstream format specifies the relationship between Network Abstraction Layer (NAL) unit streams and byte streams, both of which are referred to as bitstreams.

[0311] A bitstream can be in one of two formats, e.g., the NAL unit stream format or the byte stream format. The NAL unit stream format is conceptually the more "basic" type. The NAL unit stream format contains a sequence of syntactic structures called NAL units. This sequence is ordered in decoding order. There are constraints imposed on the decoding order (and) of the NAL units in a NAL unit stream.

[0312] A byte stream format may be constructed from the NAL unit stream format by ordering the NAL units in decoding order and prefixing each NAL unit with a start code prefix and zero or more zero-valued bytes to form a stream of bytes. The NAL unit stream format may be extracted from the byte stream format by searching for the location of a unique start code prefix pattern in this stream of bytes.

[0313] This term specifies one embodiment of the relationship between the source and the decoded pictures provided via the bitstream.

[0314] The video source represented by a bitstream is a sequence of pictures in decoding order.

[0315] Typically, the value of merge_gpm_partition_idx[x0][y0] is decoded from the bitstream. In one example, the value range of merge_gpm_partition_idx[][] is 0 to 63, inclusive. In one example, the decoding process of merge_gpm_partition_idx[][] is "bypassed." When merge_gpm_partition_idx[x0][y0] is not present, it is inferred to be equal to 0.

[0316] S2202: Obtain an angle index value angleIdx for the current coding block according to the division mode index value and the table.

[0317] In one implementation, angleIdx is used for the geometric partition of the current coding block.

[0318] In one example, angleIdx specifies the angle index of a geometric partition.

[0319] The value of the angle parameter for the current block is obtained according to the index value of the division mode and the value of a predefined lookup table.

[0320] In one embodiment, the partition angle variable angleIdx (angle parameter) and distance variable distanceIdx of the geometric partitioning mode are set according to the value of merge_gpm_partition_idx[xCb][yCb] (indicator) specified in the following table. In an implementation form, this relationship may be implemented according to Table 1 or according to a function. [Table 6]

[0321] S2203: The index value partIdx is set according to the angle index value angleIdx.

[0322] In one implementation, the index value partIdx satisfies the following:

number

[0323] In one example, threshold1 is 13 and threshold2 is 27.

[0324] In one example, the index value may be a fracture index, denoted as isFlip, where the index value isFlip also satisfies:

number

[0325] S2204: Decode the current coding block according to the index value partIdx.

[0326] The decoding process may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. Therefore, the buffer utilization and decoding efficiency are improved.

[0327] In one implementation form, decoding the coding block includes storing motion information for the current block according to the index value partIdx.

[0328] According to the above embodiment, the coding block is decoded according to the index value. The decoding process may be a sample weight derivation process, a motion information storage process, a motion vector derivation process, etc. Therefore, the buffer utilization and the decoding efficiency are improved.

[0329] FIG. 23 shows a video decoder including: a parsing unit 2301 configured to obtain a partition mode index value for a current coding block; an angle index value obtaining unit 2302 configured to obtain an angle index value angleIdx for the current coding block according to the partition mode index value and a table; a setting unit 2303 configured to set an index value partIdx according to the angle index value angleIdx; and a processing unit 2304 configured to decode the current coding block according to the index value partIdx.

[0330] The method according to FIG. 22 may be performed by a decoder according to FIG.

[0331] Further features and implementation forms of the method correspond to those of the decoder.

[0332] One specific example is the motion vector storage process for the geometric partitioning mode.

[0333] This process is invoked when decoding a coding unit with MergeGpmFlag[xCb][yCb] set to 1.

[0334] The inputs to this process are: The luma position (xCb, yCb) that specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, The variable cbWidth, which specifies the width of the current coding block in luma samples; The variable cbHeight, which specifies the height of the current coding block in luma samples; A variable, angleIdx, that specifies the angle index for the geometric partition, A variable distanceIdx that specifies the distance index of the geometric partition, Luma motion vectors mvA and mvB at 1 / 16 fractional sample precision, Reference indices refIdxA and refIdxB, Predictive list flags preListFlagA and preListFlagB. The variables numSbX and numSbY, which specify the number of 4x4 blocks in the current coding block in the horizontal and vertical directions, are set equal to cbWidth>>2 and cbHeight>>2, respectively. The variables displacementX, displacementY, isFlip, and shiftHor are derived as follows:

number

[0335] The variables offsetX and offsetY are derived as follows: If ShiftHor is equal to 0, the following applies:

number

number

[0336] The variable motionIdx is calculated based on the array disLut specified in Table 37 as follows:

number

number

[0337] If sType is equal to 0, the following applies:

number

number

number

[0338] For x=0..3 and y=0..3 the following assignments are made:

number

[0339] 24 is a block diagram illustrating an embodiment of a decoder according to the present invention. The decoder 2400 includes a processor 2401 and a non-transitory computer-readable storage medium 2402 coupled to the processor 2401 and storing programming instructions for execution by the processor 2401, which, when executed by the processor, cause the decoder to perform a method according to the first aspect or any implementation form or embodiment thereof.

[0340] Example 1. A method of coding implemented by a decoding device or an encoding device, comprising: Obtaining the division mode for the current coding block; Splitting the current coding block into two sub-blocks (sub-block A, sub-block B) based on a split mode; Setting the index value of sub-block A according to a predefined function or a predefined table.

[0341] Example 2. The method of example 1, wherein the input value for the predefined function or predefined table is an angle index value for the current coding block (e.g., the angle value is used for the geometric partition of the current coding block).

[0342] Example 3. Predefined functions

number

[0343] [Table 7] where AngleIdx is the angle index value for the current coding block, and partIdx[AngleIdx] is the index value of sub-block A.

[0344] Example 5. A method of coding implemented by a decoding device or an encoding device, comprising: Obtaining the value of the angle parameter for the current block; Obtaining the value of the distance parameter for the current block; calculating a sample distance value for a given sample in the current block (e.g., a sample located at a bottom-left, bottom-right, top-left or top-right position in the current block) according to values ​​of the angle parameter and the distance parameter; and setting two index values ​​for two sub-blocks in the current block based on a sample distance value for the predefined samples.

[0345] Example 6. Setting two index values ​​for two subblocks in a current block based on the sample distance values ​​for predefined samples: setting an index value for a sub-block in which the defined sample is located equal to a first value; 6. The method of example 5, further comprising: setting index values ​​for the other subblocks equal to a second value, wherein a processing order corresponding to the first value is before a processing order corresponding to the second value.

[0346] Example 7. The method of any one of Examples 1-6, further comprising selecting a value or blend mask (eg, Mask1 or Mask2) based on an index value for the subblock.

[0347] Example 8. A decoder (30) or encoder including processing circuitry for implementing the method according to any one of examples 1 to 7.

[0348] Example 9. A computer program product comprising program code for carrying out the method according to any one of claims 1 to 7.

[0349] Example 10. A decoder or encoder, one or more processors; A decoder or encoder comprising: a non-transitory computer-readable storage medium coupled to a processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder to perform the method of any one of Examples 1-7.

[0350] Mathematical Operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real-valued division are defined. The numbering and counting conventions generally start at 0. For example, "first" is equivalent to number 0, "second" is equivalent to number 1, etc.

[0351] Arithmetic Operators The following arithmetic operators are defined as follows: + Add - subtraction (as a two-argument operator) or negation (as a unary prefix operator) * Multiplication, including matrix multiplication x y Power. Specifies the power of the exponent y of x. In other contexts, such notations are used as superscripts that are not intended to be interpreted as powers. / Integer division with the result truncated towards zero. For example, 7 / 4 and -7 / -4 round down to 1, and -7 / 4 and 7 / -4 round down to -1. ÷ Used to indicate division in mathematical expressions where truncation or rounding is not intended.

[0352]

number

[0353]

number

[0354] Logical Operators The following logical operators are defined as follows: x&&y The Boolean logic "and" of x and y x||y The Boolean logic "or" of x and y Boolean logic "negation" x?y:zIf x is TRUE or not 0, it is evaluated with the value of y, otherwise it is evaluated with the value of z.

[0355] Relational Operators The following relational operators are defined as follows: > Greater than >= Greater than or equal to < Less than <= Less than or equal == Equal != Not equal When a relational operator is applied to a syntactic element or variable to which the value "na" (not applicable) is assigned, the value "na" is treated as a non-duplicate value of that syntactic element or variable. The value "na" is considered not equal to other values.

[0356] Bitwise operators The following bitwise operators are defined as follows. & Bitwise "and". When operating on integer arguments, it operates on the two's complement representation of integer values. When operating on binary arguments that contain fewer bits than the other argument, the shorter argument is extended by adding bits equal to 0 up to the larger bit count. | Bitwise "or". When operating on integer arguments, it operates on the two's complement representation of integer values. When operating on binary arguments that contain fewer bits than the other argument, the shorter argument is extended by adding bits equal to 0 up to the larger bit count. ^ Bitwise "exclusive or". When operating on integer arguments, it operates on the two's complement representation of integer values. When operating on binary arguments that contain fewer bits than the other argument, the shorter argument is extended by adding bits equal to 0 up to the larger bit count. x>>y x is the arithmetic right shift of the two's complement integer representation of y binary digits. This function is defined only for non-negative integer values of y. The bits shifted into the most significant bit (MSB) as a result of the right shift have a value equal to the MSB of x before the shift operation. x<<y x is the arithmetic left shift of the two's complement integer representation of y binary digits. This function is defined only for non-negative integer values of y. The bits shifted into the least significant bit (LSB) as a result of the left shift have a value equal to 0.

[0357] Assignment operators The following assignment operators are defined as follows. = Assignment operator ++ increment, i.e. x++ is equivalent to x=x+1, and when used in an array index, is evaluated at the value of the variable before the increment operation. -- Decrement, i.e. x++ is equivalent to x=x+1, and when used in an array index, is evaluated at the value of the variable before the decrement operation. += Increment by the specified amount, i.e. x+=3 is equivalent to x=x+3 and x+=(-3) is equivalent to x=x+(-3). -= Decrement by the specified amount, i.e. x-=3 is equivalent to x=x-3 and x-=(-3) is equivalent to x=x-(-3).

[0358] Range Notation To specify a range of values, the following notation is used: x=y..zx takes integer values ​​from y to z, where x, y and z are integers and z is greater than y.

[0359] Mathematical Functions The following mathematical functions are defined:

[0360]

number

[0361]

number

[0362]

number

[0363]

number

[0364]

number

[0365]

number

[0366] Operation precedence order If the precedence of an expression is not stated explicitly using parentheses, the following rules apply: - Operations with higher precedence are evaluated before operations with lower precedence. - Operations of equal precedence are evaluated sequentially from left to right. The following table specifies the priority of operations from highest to lowest, with a higher position in the table indicating a higher priority. For those operators that are also used in the C programming language, the precedence used herein is the same as the precedence used in the C programming language.

[0367] [Table 8] Table: Operation precedence from highest (top of table) to lowest (bottom of table)

[0368] Logical operations in text In this text, a description of logical operations written mathematically in the form

[0369]

number

[0370]

number

[0371]

number

[0372]

number

[0373]

number

[0374]

number

[0375] 18 is a block diagram showing a content supply system 3100 for realizing a content distribution service. The content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 may include, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any kind of combination thereof. The capture device 3102 may generate data and encode the data by the encoding method shown in the above embodiment. Alternatively, the capture device 3102 may distribute data to a streaming server (not shown), which encodes the data and transmits the encoded data to the terminal device 3106.

[0376] The capture device 3102 may include, but is not limited to, a camera, a smartphone or pad, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include a source device 12 as described above. If the data includes video, a video encoder 20 included in the capture device 3102 may actually perform a video encoding process. If the data includes audio (i.e., voice), an audio encoder included in the capture device 3102 may actually perform an audio encoding process. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, e.g., in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.

[0377] In the content supply system 3100, the terminal device 310 receives and plays the encoded data. The terminal device 3106 can be a device having data receiving and recovery capability and capable of decoding the aforementioned encoded data, such as a Smartphone or Pad 3108, a computer or laptop 3110, a Network Video Recorder (NVR) / Digital Video Recorder (DVR) 3112, a TV 3114, a Set Top Box 3116, a video conferencing system 3118, a video surveillance system 3120, a Personal Digital Assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the capture device 3106 may include the source device 14 as described above. If the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. If the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.

[0378] In a terminal device with a display, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can provide the decoded data to its display. In a device without a display, such as a STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is contacted thereto to receive and show the decoded data.

[0379] If each device in the system performs encoding or decoding, it can use a picture encoding device or a picture decoding device as shown in the previous embodiment.

[0380] 19 is a diagram showing an example of the configuration of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progression unit 3202 analyzes the transmission protocol of the stream. The protocol may include, but is not limited to, Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real Time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any kind of combination thereof.

[0381] After the protocol progression unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, such as a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0382] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206 including the video decoder 30 described in the previous embodiment decodes the video ES by the decoding method shown in the previous embodiment to generate video frames, and transmits the data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames, and transmits the data to a synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. 19) before being supplied to the synchronization unit 3212. Alternatively, the audio frames may be stored in a buffer (not shown in FIG. 19) before being supplied to the synchronization unit 3212.

[0383] The synchronization unit 3212 synchronizes the video and audio frames and provides the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information, which may be coded in a syntax using time stamps for the presentation of the encoded audio and visual data and for the delivery of the data stream itself.

[0384] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.

[0385] The present invention is not limited to the above-mentioned system, and either the picture encoding device or the picture decoding device in the above-mentioned embodiments can be incorporated into other systems, for example, an automobile system.

[0386] Although embodiments of the present invention have been described primarily in terms of video coding, it should be noted that embodiments of coding system 10, encoder 20 and decoder 30 (as well as corresponding systems 10), and also other embodiments described herein, may be configured for still picture processing or coding, e.g., processing or coding of an individual picture independent of any preceding or subsequent pictures, as in video coding. In general, when picture processing coding is limited to a single picture 17, only inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also called tools or techniques) of the video encoder 20 and video decoder 30, e.g., residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, partitioning 262 / 362, intraframe prediction 254 / 354, and / or loop filtering 220, 320, and entropy encoding 270 and entropy decoding 304, may be equally used for still picture processing.

[0387] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein with reference to, for example, the encoder 20 and the decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted over a communication medium as one or more instructions or code and executed by a hardware-based processing unit.

[0388] A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium, which includes any medium that facilitates transfer of a computer program from one place to another, for example according to a communications protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable storage medium.

[0389] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, computer-readable storage media and data storage media should be understood to not include connections, carrier waves, signals, or other transitory media, but instead to non-transitory, tangible storage media. Disks, as used herein, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and disks reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer readable media.

[0390] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any of the foregoing structures, or other structures suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.

[0391] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). In this disclosure, various components, modules, or units are described to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, along with appropriate software and / or firmware.

Claims

1. 1. A method of coding implemented by a decoding device, comprising: receiving an encoded bitstream; obtaining a fragmentation mode index value for a current coding block by parsing the encoded bitstream; obtaining a value of an angle index for the current coding block according to the value of the split mode index and a predefined table; setting a value of an index according to the value of the angle index; storing motion information for the current coding block according to the value of the index.

2. 2. The method of claim 1 , wherein if the value of the angle index is greater than or equal to a first threshold and the value of the angle index is less than or equal to a second threshold, the value of the index is equal to 1, otherwise the value of the index is equal to 0, the first threshold and the second threshold are integer values, and the first threshold is less than the second threshold.

3. The method of claim 2 , wherein the first threshold is 13 and the second threshold is 27.

4. The method of claim 1 , wherein the value of the split mode index is used to indicate which geometric partition mode is used for the current coding block.

5. The method of claim 1 , wherein the value of the angular index is used for a geometric partition of the current coding block.

6. receiving an encoded bitstream; a parsing unit configured to obtain a value of a fragmentation mode index for a current coding block by parsing the encoded bitstream; an angle index value obtaining unit configured to obtain a value of an angle index for the current coding block according to the value of the fragmentation mode index and a predefined table; a setting unit configured to set a value of an index according to the value of the angle index; a processing unit configured to store motion information for the current coding block according to the value of the index.

7. 7. The video decoder of claim 6, wherein if the value of the angular index is greater than or equal to a first threshold and if the value of the angular index is less than or equal to a second threshold, the value of the index is equal to 1, otherwise the value of the index is equal to 0, the first threshold and the second threshold are integer values, and the first threshold is less than the second threshold.

8. 8. The video decoder of claim 7, wherein the first threshold is 13 and the second threshold is 27.

9. 7. The video decoder of claim 6, wherein the value of the fission mode index is used to indicate which geometric partition mode is used for the current coding block.

10. The video decoder of claim 6 , wherein the value of the angular index is used for a geometric partition of the current coding block.

11. 1. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to: receiving an encoded bitstream; obtaining a fragmentation mode index value for a current coding block by parsing the encoded bitstream; obtaining a value of an angle index for the current coding block according to the value of the split mode index and a predefined table; setting a value of an index according to the value of the angle index; storing motion information for the current coding block according to the value of the index.

12. 12. The non-transitory computer-readable storage medium of claim 11, wherein if the value of the angular index is greater than or equal to a first threshold and the value of the angular index is less than or equal to a second threshold, the value of the index is equal to 1, otherwise the value of the index is equal to 0, the first threshold and the second threshold are integer values, and the first threshold is less than the second threshold.

13. 13. The non-transitory computer-readable storage medium of claim 12, wherein the first threshold is 13 and the second threshold is 27.

14. 12. The non-transitory computer-readable storage medium of claim 11, wherein the value of the fission mode index is used to indicate which geometric partition mode is used for the current coding block.

15. The non-transitory computer-readable storage medium of claim 11 , wherein the value of the angular index is used for a geometric partition of the current coding block.

16. 1. A decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the decoder to: receiving an encoded bitstream; obtaining a fragmentation mode index value for a current coding block by parsing the encoded bitstream; obtaining a value of an angle index for the current coding block according to the value of the split mode index and a predefined table; setting a value of an index according to the value of the angle index; storing motion information for the current coding block according to the value of the index.

17. 17. The decoder of claim 16, wherein if the value of the angular index is greater than or equal to a first threshold and the value of the angular index is less than or equal to a second threshold, the value of the index is equal to 1, otherwise the value of the index is equal to 0, the first threshold and the second threshold are integer values, and the first threshold is less than the second threshold.

18. 18. The decoder of claim 17, wherein the first threshold is 13 and the second threshold is 27.

19. The decoder of claim 16 , wherein the value of the fission mode index is used to indicate which geometric partition mode is used for the current coding block.

20. The decoder of claim 16 , wherein the value of the angular index is used for a geometric partition of the current coding block.