Encoder, decoder, and corresponding method for simplifying the signaling of a picture header

By determining if a picture is an I picture and setting inter prediction syntax elements to default values in the picture header, the method addresses the challenge of reducing video data overhead in video coding, enhancing compression efficiency and simplifying signaling.

JP7697075B2Active Publication Date: 2025-06-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024017638
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2024-02-08
Publication Date
2025-06-23
Estimated Expiration
2040-10-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing and transmitting video data due to the substantial amount of data required, especially in applications with limited bandwidth. Additionally, there is a need to simplify the signaling of picture headers to reduce overhead.

Method used

The proposed solution involves a method for encoding and decoding that includes parsing a bitstream to determine if a picture is an I picture. If it is, syntax elements for inter prediction are set to default values and not signaled in the picture header. For P or B pictures, these syntax elements are signaled in the picture header.

Benefits of technology

This approach simplifies the signaling of picture headers for I pictures, reducing the signaling overhead and improving the efficiency of video data compression and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697075000022
    Figure 0007697075000022
  • Figure 0007697075000023
    Figure 0007697075000023
  • Figure 0007697075000024
    Figure 0007697075000024
Patent Text Reader

Abstract

To provide an encoder, a decoder, and a corresponding method for simplifying picture header signaling.SOLUTION: A method of coding implemented by a decoding device includes a step of parsing a bitstream to obtain a flag from a picture header of the bitstream, the flag indicates whether the current picture is an I picture. When the flag indicates that the current picture is an I picture, a syntax element designed for inter prediction is extrapolated to a default value, alternatively, when the flag indicates that the current picture is a P or B picture, a syntax element designed for inter prediction is obtained from the picture header.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This patent application claims priority to U.S. Provisional Application No. 62 / 913,730, filed on October 10, 2019, and the entire content thereof is incorporated herein by reference.

[0002] [Technical Field] Embodiments of the present application (disclosure) generally relate to the field of picture processing, and more particularly, to simplifying picture signaling.

Background Art

[0003] Video coding (video encoding and / or decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real - time conversation applications such as video chat, video conferencing, DVDs and Blu - ray discs, video content acquisition and editing systems, and camcorders for security applications.

[0004] The amount of video data necessary to depict even relatively short videos can be substantial, which can cause difficulties when the data is being streamed or communicated across a communication network having a limited bandwidth capacity. Thus, video data is generally compressed before being communicated across modern telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device, as memory resources can be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data necessary to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the ever-increasing demand for higher video quality, it is desirable to simplify the signaling of picture headers. SUMMARY OF THE INVENTION

[0005] Embodiments of the present application provide an apparatus and method for encoding and decoding according to independent claims.

[0006] The above and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the detailed description, and the drawings.

[0007] According to a first aspect, the present invention relates to a method of coding implemented by a decoding device. The method includes parsing a bitstream to obtain a flag from a picture header of the bitstream, the flag indicating whether the current picture is an I picture.

[0008] When a flag indicates that the current picture is an I picture, the syntax elements designed for inter prediction are estimated to default values, or when a flag indicates that the current picture is a P or B picture, obtain the syntax elements designed for inter prediction from the picture header.

[0009] The syntax elements designed for inter prediction include one or more of the following elements, namely, pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice, pic_cu_qp_delta_subdiv_inter_slice, pic_cu_chroma_qp_offset_subdiv_inter_slice, pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_fpel_mmvd_enabled_flag or pic_disable_bdof_dmvr_flag.

[0010] According to a second aspect, the present invention relates to a coding method implemented by an encoding device. The method includes determining whether the current picture is an I picture, sending a bitstream to a decoding device, wherein the picture header of the bitstream includes a flag indicating whether the current picture is an I picture, and when the current picture is an I picture, the syntax elements designed for inter prediction are not signaled in the picture header, or when the current picture is a P or B picture, the syntax elements designed for inter prediction are signaled in the picture header.

[0011] The method according to the first aspect of the present invention can be executed by the apparatus according to the third aspect of the present invention. Further features and implementation forms of the method according to the third aspect of the present invention correspond to the features and implementation forms of the apparatus according to the first aspect of the present invention.

[0012] The method according to the second aspect of the present invention can be executed by the apparatus according to the fourth aspect of the present invention. Further features and implementation forms of the method according to the fourth aspect of the present invention correspond to the features and implementation forms of the apparatus according to the second aspect of the present invention.

[0013] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream, including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the first aspect.

[0014] According to a sixth aspect, the present invention relates to an apparatus for encoding a video stream, including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the second aspect.

[0015] According to a seventh aspect, there is proposed a computer-readable storage medium storing instructions which, when executed, cause one or more processors to be configured to code video data. The instructions cause the one or more processors to execute the method according to the first or second aspect or any possible implementation of the first or second aspect.

[0016] According to an eighth aspect, the present invention relates to a computer program including program code for executing the method according to the first or second aspect or any possible implementation of the first or second aspect when executed on a computer.

[0017] As described above, by indicating whether the current picture is an I picture in the picture header of the bitstream, when the current picture is an I picture, the syntax elements designed for inter prediction are not signaled in the picture header. Therefore, the embodiment can simplify the signaling of the picture header for all intra pictures, i.e., I pictures. Correspondingly, the signaling overhead is reduced.

[0018] The details of one or more embodiments are set forth in the accompanying drawings and the following detailed description. Other features, objects, and advantages will be apparent from the detailed description, the drawings, and the claims.

Brief Description of the Drawings

[0019] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings and figures.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10a

Figure 10b

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

[0020] Hereinafter, the same reference numerals indicate the same or at least functionally equivalent features unless otherwise explicitly specified.

Best Mode for Carrying Out the Invention

[0021] In the following description, reference is made to the accompanying drawings, which form a part of the present disclosure and illustrate specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention can be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the drawings. Therefore, the following detailed description should not be considered in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0022] For example, it is understood that the disclosures related to the described method may also apply to corresponding devices or systems configured to perform the method, and vice versa. For example, when one or more specific method steps are described, the corresponding device may include one or more units, such as functional units, for performing the one or more method steps described, even if such one or more units are not explicitly described or shown in the drawings (e.g., one unit performs one or more steps, or multiple units each perform one or more of multiple steps). On the other hand, for example, when a specific device is described based on one or more units, such as functional units, the corresponding method may include one step for performing the functions of the one or more units, even if such one or more steps are not explicitly described or shown in the drawings (e.g., one step performs the functions of one or more units, or multiple steps each perform the functions of one or more of multiple units). Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.

[0023] Typically, video coding refers to the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture", the terms "frame" or "image" may be used as synonyms in the field of video coding. Video coding (or generally coding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves performing the reverse process compared to the encoder to reconstruct the video picture. Embodiments that refer to the "coding" of a video picture (or generally a picture) are to be understood as relating to the "encoding" or "decoding" of the video picture or respective video sequence. The combination of the encoding part and the decoding part is also referred to as a CODEC (Coding and Decoding).

[0024] In the case of reversible video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss during storage or transmission). In the case of irreversible video coding, to reduce the amount of data representing the video picture, further compression, e.g., by quantization, is performed, which cannot be fully reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.

[0025] Some video coding standards belong to the group of "irreversible hybrid video codecs" (i.e., combining spatial and temporal prediction in the sample domain and 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, for example, prediction blocks are generated using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, the prediction blocks are subtracted from the current block (the block being currently processed / to be processed) to obtain a residual block, the residual block is transformed, and the residual block is quantized in the transform domain to reduce the amount of data to be transmitted (compression), so that the video is typically processed, i.e., encoded, at the block (video block) level. On the other hand, in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed block to reconstruct the current block for presentation. Further, the encoder duplicates the decoder processing loop, so that both generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for processing subsequent blocks, i.e., for coding.

[0026] Embodiments of a video coding system 10, a video encoder 20, and a video decoder 30 will be described below with reference to FIGS. 1 to 3.

[0027] FIG. 1A is a schematic block diagram showing an exemplary coding system 10 that can utilize the technology of the present application, such as a video coding system 10 (or simply coding system 10). The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that can be configured to execute the technology according to various examples described in the present application.

[0028] As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 for decoding the encoded picture data 13.

[0029] The source device 12 includes an encoder 20 and may further, i.e., optionally, include a picture source 16, a pre-processor (or pre-processing unit) 18, for example a picture pre-processor 18, and a communication interface or communication unit 22.

[0030] The picture source 16 may be any kind of picture capture device, for example a camera for capturing real-world pictures, and / or any kind of picture generation device, for example a computer graphics processor for generating computer animation pictures, or any other device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures) and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may be any kind of memory or storage for storing any of the above pictures.

[0031] In contrast to the pre-processor 18 and the processing performed by the pre-processing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.

[0032] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.

[0033] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21 (for more details, it will be described below based on, for example, FIG. 2).

[0034] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) on the communication channel 13 to other devices, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0035] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and further, optionally, may include a communication interface or communication unit 28, a postprocessor 32 (or post-processing unit 32), and a display device 34.

[0036] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or from any other source, such as a storage device, e.g., an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.

[0037] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between source device 12 and destination device 14, e.g., via a direct wired or wireless connection, or via any kind of network, e.g., a wired or wireless network or any combination thereof, or via any kind of private and public network, or any combination of any kind thereof.

[0038] Communication interface 22 may be configured to, for example, package encoded picture data 21 into an appropriate format, e.g., a packet, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.

[0039] Communication interface 28, which forms the counterpart of communication interface 22, may be configured to, for example, receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpacking to obtain the encoded picture data 21.

[0040] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface as indicated by the arrow for communication channel 13 pointing from source device 12 to destination device 14 in FIG. 1A, or as a bidirectional communication interface, e.g., to send and receive messages, e.g., to establish a connection and approve and exchange any other information related to the communication link and / or data transmission, e.g., encoded picture data transmission.

[0041] Decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (for further details, see, e.g., the following description based on FIG. 3 or FIG. 5).

[0042] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing executed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming or resampling, or any other processing for preparing the decoded picture data 31, for example, for display by the display device 34.

[0043] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 and display the picture to, for example, a user or viewer. The display device 34 may be or include any type of display that presents a reconstructed picture, for example, an integrated or external display or monitor. The display may be, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display or include any of these.

[0044] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the devices may also include both or both functions, the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.

[0045] As will be apparent to those skilled in the art based on the description, the presence and (exact) partitioning of the different units or functions within the source device 12 and / or destination device 14 as shown in FIG. 1A may vary depending on the actual device and application.

[0046] The encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30) or both the encoder 20 and decoder 30 may be implemented via a processing circuit as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 may be implemented via the processing circuit 46 to embody various modules as described with respect to the encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via the processing circuit 46 to embody various modules as described with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as described below. As shown in FIG. 5, if the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of a combined encoder / decoder (CODEC) within a single device as shown in FIG. 1B.

[0047] The source device 12 and the destination device 14 may include any of a wide range of devices, such as any type of handheld or fixed device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server, etc.), a broadcast receiver device, a broadcast transmitter device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.

[0048] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the technology of the present application may be applied to video coding settings (for example, video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, the data is retrieved from local memory, streamed over a network, etc. The video encoding device may encode the data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data into memory and / or retrieve and decode data from memory.

[0049] For the sake of convenience of explanation, embodiments of the present invention are described herein by referring to, for example, the reference software of High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), and the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.

[0050] Encoder and Encoding Method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction processing unit 254, and a partition unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.

[0051] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming the forward signal path of the encoder 20. On the other hand, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming the reverse signal path of the video encoder 20, and the reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also sometimes referred to as forming the "built-in decoder" of the video encoder 20.

[0052] Picture and picture partition (picture and block) The encoder 20 may be configured to receive, for example, via the input 201, a picture 17 (or picture data 17), for example, a sequence of pictures forming a video or video sequence. The received picture or picture data may also be the preprocessed picture 19 (preprocessed picture data 19). For the sake of brevity, the following description refers to the picture 17. The picture 17 may also be referred to as the current picture or the picture to be coded (especially in video coding, to distinguish the current picture from other pictures, for example, pictures that have been encoded and / or decoded before in the same video sequence, i.e., the video sequence that also includes the current picture).

[0053] (Digital) A picture can be or can be considered as a two-dimensional array or matrix of samples having intensity values. Samples within the array may also be referred to as pixels (abbreviation of picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used, i.e., the picture may be represented as or may include three sample arrays. In the RGB format or color space, the picture includes corresponding sample arrays of red, green, and blue. However, in video coding, each pixel typically includes a luminance component represented by Y (L may also be used instead in some cases) and two chrominance components represented by Cb and Cr, and is represented in YCbCr. The luminance (or simply luma) component Y represents brightness or gray-level intensity (e.g., in a grayscale picture, etc.). On the other hand, the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted or transformed to the YCbCr format, and vice versa, and the process is also known as color conversion or transformation. If the picture is monochrome, the picture may include only a luminance sample array. Thus, the picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0054] An embodiment of the video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size for all pictures of the video sequence and the corresponding grid defining the block size, or to vary the block size between pictures or subsets or groups of pictures and partition each picture into corresponding blocks.

[0055] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, some, or all of the blocks forming picture 17. Picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.

[0056] Similar to picture 17, picture blocks 203 can also be or be considered as a two-dimensional array or matrix of samples having intensity values (sample values), but with dimensions smaller than those of picture 17. In other words, block 203 may include, for example, one sample array (e.g., the luma array in the case of a monochrome picture 17, or the luma or chroma array in the case of a color picture), or three sample arrays (e.g., the luma and two chroma arrays in the case of a color picture 17), or any other number and / or type of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.

[0057] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode the picture 17 block by block. For example, encoding and prediction may be performed for each block 203.

[0058] An embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode the picture by using slices (also called video slices). The picture may be partitioned into one or more slices (typically non-overlapping) or encoded using the same, and each slice may include one or more blocks (e.g., CTUs).

[0059] An embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode the picture by using slice / tile groups (also called video tile groups) and / or tiles (also called video tiles). The picture may be partitioned into one or more slice / tile groups (typically non-overlapping) or encoded using the same, and each slice / tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may have, for example, a rectangular shape and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.

[0060] Residual calculation The residual calculation unit 204 may be configured to calculate the residual block 205 by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203, for example, sample by sample (pixel by pixel) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 are provided below), to obtain the residual block 205 in the sample domain.

[0061] Transformation The transformation processing unit 206 may be configured to apply a transformation, for example, a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207 may also be referred to as transform residual coefficients and may represent the residual block 205 in the transform domain.

[0062] The transformation processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the transformation specified for H.265 / HEVC. Compared with the orthogonal DCT transform, such an integer approximation is typically scaled by a specific factor. To maintain the norm of the residual block processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transform process. The scaling factor is typically selected based on specific constraints such as the scaling factor being a power of two for shift operations, the bit depth of the transform coefficients, and the trade-off between accuracy and implementation cost. A specific scaling factor may be specified, for example, for the inverse transform by the inverse transform processing unit 212 (and the corresponding inverse transform by the inverse transform processing unit 312 in the video decoder 30, for example), and the corresponding scaling factor for the forward transform by the transformation processing unit 206 in the encoder 20 may be correspondingly specified.

[0063] Embodiments of the video encoder 20 (each, the transformation processing unit 206) may be configured to output transformation parameters, for example, the type of transformation or transformations, which are encoded or compressed, for example, directly or via the entropy encoding unit 270, whereby, for example, the video decoder 30 may receive and use the transformation parameters for decoding.

[0064] Quantization The quantization unit 208 may be configured to quantize the transform coefficient 207 to obtain a quantized coefficient 209, for example, by applying scalar quantization or vector quantization. The quantized coefficient 209 may also be referred to as the quantized transform coefficient 209 or the quantized residual coefficient 209.

[0065] The quantization process may reduce the bit depth associated with some or all of the conversion coefficient 207. For example, an n-bit conversion coefficient may be truncated to an m-bit conversion coefficient during quantization, where n is greater than m. The degree of quantization may be changed by adjusting a quantization parameter (QP). For example, in scalar quantization, different scalings may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization. On the other hand, a larger quantization step size corresponds to coarser quantization. The applicable quantization step may be indicated by a quantization parameter (QP). The quantization parameter may be, for example, an index to a predetermined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), and a large quantization parameter may correspond to coarse quantization (large quantization step size), and vice versa. Quantization may include division by a quantization step size. For example, the corresponding and / or inverse dequantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, such as embodiments according to HEVC, may be configured to use a quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an expression that includes division. Due to the scaling used in the fixed-point approximation of the expressions for the quantization step size and the quantization parameter, an additional scaling factor for quantization and dequantization may be introduced to restore the norm of the residual block that can be changed. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a customized quantization table may be used, for example, signaled from the encoder to the decoder in the bitstream. Quantization is a non-invertible operation, and the loss increases as the quantization step size increases.

[0066] Embodiments of the video encoder 20 (each, quantization unit 208) may be configured to output quantization parameters (QP) that are encoded, for example, directly or via the entropy encoding unit 270, such that, for example, the video decoder 30 may receive and apply the quantization parameters for decoding.

[0067] Inverse quantization The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantized coefficients, for example, by applying the inverse of the quantization method applied by the quantization unit 208 based on or using the same quantization step size as the quantization unit 208, to obtain the inverse quantized coefficients 211. The inverse quantized coefficients 211 are also referred to as the inverse quantized residual coefficients 211 and typically are not identical to the transform coefficients due to loss by quantization, but may correspond to the transform coefficients 207.

[0068] Inverse transform The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain the reconstructed residual block 213 (or corresponding inverse quantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.

[0069] Reconstruction The reconstruction unit 214 (e.g., adder or summator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 on a sample-by-sample basis, to obtain the reconstructed block 215 in the sample domain.

[0070] Filtering The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or to improve video quality. The loop filter unit 220 may include a deblocking filter, a sample-adaptive offset (SAO) filter or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening, a smoothing filter or a collaborative filter, or any combination thereof. The loop filter unit 220 is shown in FIG. 2 as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.

[0071] Embodiments of the video encoder 20 (each, the loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information) that are encoded, for example, directly or via the entropy encoding unit 220, such that, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.

[0072] Decoded picture buffer The decoded picture buffer (DPB) 230 may be a memory for storing reference pictures or generally reference picture data for encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store the same current picture or different pictures, for example, other previously filtered blocks of a previously reconstructed picture, for example, a previously reconstructed and filtered block 221, and may provide, for example, for inter prediction, a fully previously reconstructed, i.e., decoded picture (and corresponding reference blocks and samples), and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, or generally unfiltered reconstructed samples, or other further processed versions of either the reconstructed blocks or samples, for example, if the reconstructed block 215 is not filtered by the loop filter unit 220.

[0073] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, for example, filtered and / or unfiltered reconstructed samples or blocks from the same (current) picture and / or one or more previously decoded pictures, for example, from the decoded picture buffer 230 or other buffer (e.g., a line buffer not shown). The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or predictor 265.

[0074] The mode selection unit 260 may be configured to determine or select a partition (including not partitioning) for the current block prediction mode and a prediction mode (e.g., intra or inter prediction mode), and generate a corresponding prediction block 265 to be used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.

[0075] Embodiments of the mode selection unit 260 may be configured to select a partition and prediction mode that provides the best fit or, in other words, the minimum residual (the minimum residual means better compression for transmission or storage) or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or to consider or balance both, from among those supported or available by, for example, the mode selection unit 260. The mode selection unit 260 may be configured to determine a partition and prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. Terms such as "best," "minimum," "optimal," etc. in this context do not necessarily indicate the overall "best," "minimum," "optimal," etc., but may indicate an end or selection criterion such as a value above or below a threshold, or the fulfillment of other constraints that potentially result in a "quasi-optimal selection" but reduce complexity and processing time.

[0076] In other words, the partitioning unit 262 may be configured to repeatedly use, for example, quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, to further partition block 203 into smaller block partitions or sub-blocks (which also form blocks again), and to perform prediction for each of the block partitions or sub-blocks. The mode selection may include the selection of the tree structure of the partitioned block 203, and the prediction mode is applied to each of the block partitions or sub-blocks.

[0077] The partitioning (e.g., by partition unit 260) and prediction processing (e.g., by inter prediction unit 244 and intra prediction unit 254) performed by exemplary video encoder 20 will be described in further detail below.

[0078] Partition Partition unit 262 may partition (or split) the current block 203 into smaller partitions, e.g., smaller blocks of square or rectangular size. These smaller blocks (which may also be referred to as sub-blocks) may be further partitioned into even smaller partitions. This is also referred to as a tree partition or hierarchical tree partition. For example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned, e.g., into two or more blocks at the next lower tree level, e.g., nodes at tree level 1 (hierarchical level 1, depth 1), and these blocks may be further partitioned, again, into two or more blocks at the next lower tree level, e.g., tree level 2 (hierarchical level 2, depth 2), until the partitioning ends, e.g., because an end criterion is met, e.g., the maximum tree depth or minimum block size is reached. Blocks that are not further partitioned are also referred to as leaf blocks or leaf nodes of the tree. A tree that uses a partition into two partitions is called a binary-tree (BT), a tree that uses a partition into three partitions is called a ternary-tree (TT), and a tree that uses a partition into four partitions is called a quad-tree (QT).

[0079] As described above, the term "block" as used herein may be a part of a picture, particularly a square or rectangular portion. For example, referring to HEVC and VVC, a block may be or correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).

[0080] For example, a coding tree unit (CTU) may be or include a CTB of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a picture coded using a syntax structure used to code a monochrome picture or three separate color planes and samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N, whereby the partitioning of components into CTBs is a partition. A coding unit (CU) may be or include a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture having three sample arrays, or a coding block of samples of a picture coded using a syntax structure used to code a monochrome picture or three separate color planes and samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N, whereby the partitioning of CTBs into coding blocks is a partition.

[0081] For example, in an embodiment according to HEVC, a coding tree unit (CTU) may be divided into coding units (CUs) by using a quadtree structure shown as a coding tree. A determination of whether to code a picture area using inter-picture (temporal) prediction or to code a picture area using intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process is applied, and related information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.

[0082] For example, in an embodiment according to the currently under - development latest video coding standard called Versatile Video Coding (VVC), combined quadtree and binary tree (QTBT) partitions are used, for example, to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node is further partitioned by a binary tree or a ternary (or triple) tree structure. The partition of the tree leaf node is called a coding unit (CU), and this segmentation is used for prediction and transform processing without further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. In parallel, multiple partitions, for example, triple tree partitions, may be used with the QTBT block structure.

[0083] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.

[0084] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a set of prediction modes (e.g., pre - determined). The set of prediction modes may include, for example, an intra - prediction mode and / or an inter - prediction mode.

[0085] Intra - prediction The set of intra - prediction modes may include 35 different intra - prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes as defined, for example, in HEVC, or may include 67 different intra - prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes as defined for example for VVC.

[0086] The intra - prediction unit 254 is configured to use the reconstructed samples of adjacent blocks of the same current picture to generate an intra - prediction block 265 according to a certain intra - prediction mode from the set of intra - prediction modes.

[0087] The intra - prediction unit 254 (or generally the mode selection unit 260) is further configured to output to the entropy encoding unit 270, in the form of a syntax element 226 for inclusion in the encoded picture data 21, the intra - prediction parameter (or generally information indicating the intra - prediction mode selected for the block), whereby, for example, the video decoder 30 may receive and use the prediction parameter for decoding.

[0088] Inter - prediction The set (or possibilities) of inter prediction modes depends on the available reference pictures (i.e., for example, the previous at least partially decoded pictures stored in DBP230) and other inter prediction parameters, for example, whether the entire reference picture is used to search for the best matching reference block, or only a part of the reference picture, for example, the search window area around the area of the current block, and / or, for example, whether pixel interpolation, for example, half / semi pel and / or quarter pel interpolation, is applied or not.

[0089] In addition to the above prediction modes, a skip mode and / or a direct mode may be applied.

[0090] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (ME) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, the picture block 203 (the current block 203 of the current picture 17) and the decoded picture 231, or at least one or a plurality of previously reconstructed blocks, for example, the reconstructed blocks of one or a plurality of other / different previous decoded pictures 231. For example, the video sequence may include the current picture and the previous decoded picture 231, or in other words, the current picture and the previous decoded picture 231 may be part of or form a sequence of pictures forming the video sequence.

[0091] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures of a plurality of other pictures, and provide an offset (spatial offset) between the reference picture (or reference picture index) and / or the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter prediction parameter. This offset is also called a motion vector (MV).

[0092] The motion compensation unit is configured to obtain an inter prediction parameter, for example, receive it, and perform an inter prediction based on or using the inter prediction parameter to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit may include fetching or generating a prediction block based on the motion / block vector determined by motion estimation, and optionally performing interpolation to sub-pixel accuracy. The interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code a picture block. When receiving the motion vector of the PU of the current picture block, the motion compensation unit may find the prediction block pointed to by the motion vector within one of the reference picture lists.

[0093] The motion compensation unit may also generate syntax elements related to the block and the video slice for use by the video decoder 30 when decoding a picture block of the video slice. In addition to or instead of the slice and its respective syntax elements, a tile group and / or a tile and its respective syntax elements may be generated or used.

[0094] Entropy coding Entropy encoding unit 270 applies, for example, an entropy encoding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC (CAVLC) method, arithmetic coding method, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or other entropy encoding method or technique) to the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters and / or other syntax elements, or bypasses (non-compresses) them, and is configured to obtain encoded picture data 21 that can be output via output 272 in the form of, for example, an encoded bitstream 21, whereby, for example, video decoder 30 may receive and use the parameters for decoding. The encoded bitstream 21 may be transmitted to video decoder 39, or may be stored in a memory for later transmission or retrieval by video decoder 30.

[0095] Other structural variations of video encoder 20 can be used to encode a video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal for a particular block or frame without transform processing unit 206. In other implementations, encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 coupled in a single unit.

[0096] Decoder and decoding method Figure 3 shows an example of a video decoder 30 configured to implement the technology of the present application. The video decoder 30 is configured to receive, for example, encoded picture data 21 (e.g., an encoded bitstream 21) encoded by the encoder 20 in order to obtain a decoded picture 331. The encoded picture data or bitstream includes information for multiplexing the encoded picture data, for example, data representing picture blocks of encoded video slices (and / or tile groups or tiles) and related syntax elements.

[0097] In the example of Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or may include a motion compensation unit. In some examples, the video decoder 30 may execute a decoding path generally inverse to the encoding path described with respect to the video encoder 100 in Figure 2.

[0098] As described with respect to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 may also be referred to as forming the "built-in decoder" of the video encoder 20. Thus, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Thus, the descriptions provided for each unit and function of the video 20 encoder apply, correspondingly, to each unit and function of the video decoder 30.

[0099] Entropy decoding Entropy decoding unit 304 parses the bitstream 21 (or generally the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, the quantized coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.

[0100] Inverse quantization The inverse quantization unit 310 receives a quantization parameter (QP) (or generally information regarding inverse quantization) and quantized coefficients from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304 for example), and is configured to apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain the inverse quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter determined by the video encoder 20 for each video block within a video slice (or tile or tile group) in order to determine the degree of quantization and likewise the degree of inverse quantization to be applied.

[0101] Inverse transformation The inverse transformation processing unit 312 receives the inverse quantized coefficients 311, which may also be referred to as transform coefficients 311, and is configured to apply a transformation to the inverse quantized coefficients 311 to obtain the residual block 213 reconstructed in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 313. The transformation may be an inverse transformation, e.g., inverse DCT, inverse DST, inverse integer transformation or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive a transform parameter or corresponding information from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304 for example) to determine the transformation to be applied to the inverse quantized coefficients 311.

[0102] Reconstruction The reconstruction unit 314 (e.g., adder or summator 314) is configured to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the sample domain, e.g., by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.

[0103] Filtering (Either within or after the coding loop) A loop filter unit 320 is configured to filter the reconstructed block 315 to obtain a filtered block 321, for example, to smooth pixel transitions or to improve video quality. The loop filter unit 320 may include a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening, a smoothing filter, or a collaborative filter, or any combination thereof. The loop filter unit 320 is shown in FIG. 3 as an in-loop filter, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0104] Decoded picture buffer The decoded video block 321 of the picture is then stored in a decoded picture buffer 330 that stores the decoded picture 331 for use as a reference picture for subsequent motion compensation for other pictures and / or for output of each display.

[0105] The decoder 30 is configured to output the decoded picture 331, for example, via the output 312, for presentation or viewing by a user.

[0106] Prediction The inter prediction unit 344 may be the same as the inter prediction unit 244 (particularly, the motion compensation unit), the intra prediction unit 354 may be functionally the same as the inter prediction unit 254, and based on each piece of information received from the partition and / or prediction parameters or the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304), it performs a split or partition determination and prediction. The mode application unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed picture, block, or each (filtered or unfiltered) sample to obtain the prediction block 365.

[0107] When a video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from blocks decoded prior to the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists, list 0 and list 1, using default construction techniques based on the reference pictures stored in the DPB 330. The same or similar may apply to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices), e.g., the video may be coded using I, P, or B tile groups and / or tiles.

[0108] The mode application unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors or related information and other syntax elements, and uses the prediction information to generate a prediction block for the currently decoded video block. For example, the mode application unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to code video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the slice's reference picture lists, the motion vector of each inter-coded video block of the slice, the inter prediction state for each inter-coded video block of the slice, and other information for decoding video blocks within the current video slice. The same or similar approach may be applied to or in embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices), for example, the video may be coded using I, P, or B tile groups and / or tiles.

[0109] An embodiment of the video decoder 30 as shown in FIG. 3 may be configured to partition and / or decode a picture by using slices (also referred to as video slices), the picture may be partitioned into and / or decoded using one or more (typically non-overlapping) slices, and each slice may include one or more blocks (e.g., CTUs).

[0110] An embodiment of the video decoder 30 as shown in FIG. 3 may be configured to partition and / or decode a picture by using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), the picture may be partitioned into one or more tile groups (typically non-overlapping) or decoded using the same, each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.

[0111] Other variations of the video decoder 30 can be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 can directly inverse quantize the residual signal for a particular block or frame without the inverse transform processing unit 312. In other implementations, the video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 coupled in a single unit.

[0112] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clip or shift may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.

[0113] Further operations should be noted to be applicable to the derived motion vectors of the current block (including, but not limited to, the control point motion vectors in affine mode, affine, planar, sub-block motion vectors in ATMVP mode, temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predetermined range according to its representation bits. When the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" means exponentiation. For example, when bitDepth is set equal to 16, the range is -32768 to 32767, and when bitDepth is set equal to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of 4 4×4 sub-blocks within one 8×8 block) is restricted such that the maximum difference between the integer parts of the MVs of the 4 4×4 sub-blocks is not more than N pixels, such as not more than 1 pixel. Here, two methods for restricting the motion vector according to bitDepth are provided.

[0114] Method 1: Remove the overflow MSB (Most Significant Bit) by a flow operation. ux=(mvx+2 bitDepth )%2 bitDepth (1) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (2) uy=(mvy+2 bitDepth )%2 bitDepth (3) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (4) Here, mvx is the horizontal component of the motion vector of the image block or sub-block, mvy is the vertical component of the motion vector of the image block or sub-block, and ux and uy represent intermediate values.

[0115] For example, when the value of mvx is -32769, after applying equations (1) and (2), the resulting value is 32767. In a computer system, decimal values are stored as two's complements. The two's complement of -32769 is 1,0111,1111,1111,1111 (17 bits). In this case, the MSB is discarded, and thus the resulting two's complement is 0111,1111,1111,1111 (decimal 32767), which is the same as the output by applying equations (1) and (2). ux=(mvpx+mvdx+2 bitDepth )%2 bitDepth (5) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (6) uy=(mvpy+mvdy+2 bitDepth )%2 bitDepth (7) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (8)

[0116] As shown in equations (5) to (8), the operation may be applied between the sum of mvp and mvd.

[0117] Method 2: Remove the overflow MSB by clipping the value. vx=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1,vx) vy=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1,vy) Here, vx is the horizontal component of the motion vector of an image block or sub-block, vy is the vertical component of the motion vector of an image block or sub-block, x, y, and z respectively correspond to the three input values of the MV clipping process, and the definition of the function Clip3 is as follows.

Number

[0118] Figure 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the embodiments of the disclosure as described herein. In an embodiment, the video coding device 400 may be a decoder such as the video decoder 30 of FIG. 1A or an encoder such as the video encoder 20 of FIG. 1A.

[0119] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. The video coding device 400 may also include optoelectrical (OE) components and electro-optical (EO) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for the outlet or inlet of optical or electrical signals.

[0120] Processor 430 is implemented by hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiver unit 420, a transmitter unit 440, an output port 450, and a memory 460. Processor 430 includes a coding module 470. The coding module 470 implements the embodiments of the disclosure described above. For example, the coding module 470 realizes, processes, prepares, or provides various coding operations. Accordingly, what is included in the coding module 470 provides a substantial improvement to the functions of the video coding device 400 and brings about the conversion of the video coding device 400 to different states. Alternatively, the coding module 470 is realized as instructions stored in the memory 460 and executed by the processor 430.

[0121] Memory 460 may include one or more disks, tape drives, and solid-state drives and may be used as an overflow data storage device for storing such programs when a program is selected for execution and for storing instructions and data read during the execution of the program. Memory 460 may be, for example, volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0122] FIG. 5 is a simplified block diagram of an apparatus 500 that may be used as one or both of the source device 12 and the destination device 14 from FIG. 1A according to an exemplary embodiment.

[0123] The processor 502 within the apparatus 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of manipulating or processing information, existing currently or developed in the future. The disclosed implementation can be carried out with a single processor, e.g., processor 502 as illustrated, but the advantages in terms of speed and efficiency can be achieved using more than one processor.

[0124] The memory 504 within the apparatus 500 can be, in an implementation, a read only memory (ROM) device or a random access memory (RAM) device in the implementation. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and application programs 510, and the application programs 510 include at least one program that enables the processor 502 to execute the methods described herein. For example, the application programs 510 can include applications 1 to N that further include a video coding application that executes the methods described herein.

[0125] The apparatus 500 can also include one or more output devices such as a display 518. The display 518 can be, in one example, a touch-sensitive display that combines a touch-sensitive element operable to sense touch input with a display. The display 518 can be coupled to the processor 502 via the bus 512.

[0126] Although shown here as a single bus, the bus 512 of the apparatus 500 can be composed of a plurality of buses. Further, the secondary storage 514 can be directly coupled to other components of the apparatus 500, or can be accessed via a network, and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Thus, the apparatus 500 can be implemented in a wide range of configurations.

[0127] Parameter set The parameter sets in the state-of-the-art codecs are basically similar and share the same basic design goals, namely, bitrate efficiency, error resilience, and provision of a system layer interface. In HEVC (H.265), there is a hierarchy of parameter sets including a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS), which are similar to their corresponding ones in AVC and VVC. Each slice refers to a single active PPS, SPS, and VPS to access the information used to decode the slice. The PPS contains information applicable to all slices within a picture, and thus all slices within a picture must refer to the same PPS. Slices in different pictures are also permitted to refer to the same PPS. Similarly, the SPS contains information applicable to all pictures within the same coded video sequence.

[0128] The PPS may be different for separate pictures, but it is common for many or all pictures in a coded video sequence to refer to the same PPS. Reusing parameter sets is bit-rate efficient as it avoids the need to transmit shared information multiple times. Also, the content of the parameter sets is carried over a more reliable external communication link to ensure it is not lost, or is repeated frequently within the bitstream, making it robust to losses.

[0129] In HEVC, for a given slice, each slice header includes a PPS identifier that refers to a specific PPS to identify the active parameter sets at each level of the parameter set type hierarchy. The identifier that refers to a specific SPS is within the PPS. Next, the identifier that refers to a specific VPS is within the SPS.

[0130] A parameter set is activated when the current coded slice to be decoded refers to that parameter set. All active parameter sets must be available to the decoder when they are first referenced. Parameter sets may be transmitted in-band or out-of-band and may be transmitted repeatedly.

[0131] Parameter sets may be received in any order.

[0132] These characteristics of parameter sets provide improved error resilience by overcoming some network losses of the parameter sets. Further, the use of parameter sets enables individual slices to be decoded compared to the case where a picture header containing the same information exists for a subset of the slices of a picture even if other slices in the same picture have suffered network losses.

[0133] Picture Parameter Set (PPS) The PPS contains parameters that can vary for different pictures within the same coded video sequence. However, multiple pictures may refer to the same PPS even if they have different slice coding types (I, P, and B). Including these parameters within the picture parameter set rather than in the slice header can improve bitrate efficiency and provide error resilience when the PPS is transmitted more reliably.

[0134] The PPS contains a PPS identifier and an index to the reference SPS. The remaining parameters describe the coding tools used in the slices that refer to the PPS. The coding tools include tiles, weighted prediction, sign data hiding, temporal motion vector prediction, etc., and can be enabled or disabled. The coding tool parameters signaled in the PPS include the number of reference indexes, the initial quantization parameter (QP), and the chroma QP offset. Coding tool parameters such as deblocking filter control, tile configuration, and scaling list data may also be signaled in the PPS. Exemplary PPSs according to the document Versatile Video Coding (Draft 6) of Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 (available at http: / / phenix.it-sudparis.eu / jvet / , document number: Document:JVET-O2001-vE) are shown in FIGS. 6 and 7.

[0135] Figure 7 shows a section of the PPS. In this figure, the syntax element constant_slice_header_params_enabled_flag (constant_slice_header_params_enabled_flag) indicates whether there are certain parameters in the PPS. When the value of constant_slice_header_params_enabled_flag is equal to 1, additional syntax elements such as pps_temporal_mvp_enabled_idc (730) are included in the PPS. pps_temporal_mvp_enabled_idc is a syntax element that may exist in the PPS, controls the application of temporal motion vector prediction, and indicates the following. When the value of the syntax element is 1, temporal MV prediction is disabled in slices that refer to the PPS. When the value of the syntax element is 2, temporal MV prediction is enabled in slices that refer to the PPS. When the value of the syntax element is 0, a second syntax element is included in the slice header to control the application of temporal MV prediction for the slice. According to the prior art (JVET-O2001-vE), a pps_temporal_mvp_enabled_idc equal to 0 specifies that the syntax element slice_temporal_mvp_enabled_flag exists in the slice header of slices that do not have a slice_type equal to I in slices that refer to the PPS. A pps_temporal_mvp_enabled_idc equal to 1 or 2 specifies that slice_temporal_mvp_enabled_flag does not exist in the slice header of slices that refer to the PPS. A pps_temporal_mvp_enabled_idc equal to 3 is reserved for future use by ITU-T|ISO / IEC.

[0136] In FIG. 7, another example of dep_quant_enabled_flag(720) is shown. A pps_dep_quant_enabled_idc equal to 0 specifies that the syntax element dep_quant_enabled_flag is present in the slice header of the slice that refers to the PPS. A pps_dep_quant_enabled_idc equal to 1 or 2 specifies that the syntax element dep_quant_enabled_flag is not present in the slice header of the slice that refers to the PPS. A pps_dep_quant_enabled_idc equal to 3 is reserved for future use by ITU-T|ISO / IEC. When dep_quant_enabled_flag controls the application of dependent quantization for a slice, a value of zero corresponds to disabling dependent quantization, and a value of 1 corresponds to enabling dependent quantization. When pps_dep_quant_enabled_idc is not equal to zero, dep_quant_enabled_flag is not included in the slice header, and instead its value is assumed to be equal to pps_dep_quant_enabled_idc - 1.

[0137] When constant_slice_header_params_enabled_flag is true, the syntax elements included in the PPS are shown in FIG. 7. The syntax elements specify the default ones included in the PPS when constant_slice_header_params_enabled_flag is true and have the following common characteristics.

[0138] Each syntax element within the PPS has a corresponding syntax element within the slice header. For example, pps_dep_quant_enabled_idc is a syntax element within the PPS, while dep_quant_enabled_flag is the corresponding syntax element within the slice header.

[0139] If the syntax element in the PPS has a value of zero, the corresponding one in the slice header exists (is included) in the slice header. Otherwise, the syntax element of the slice header does not exist in the slice header.

[0140] If the syntax element in the PPS has a value different from zero, the value of the corresponding syntax element in the slice header is estimated according to the value of the syntax element in the PPS. For example, when the value of pps_dep_quant_enabled_idc is equal to 1, the value of dep_quant_enabled_idc is estimated to be pps_dep_quant_enabled_idc - 1 = 0.

[0141] In VVC and HEVC, the same video picture partitions (such as slices, tiles, sub - pictures, bricks, etc.) must refer to the same picture parameter set. Since the PPS may contain parameters applicable to the whole picture, this is a requirement of the coding standard. For example, as follows, the syntax elements included in the parentheses (610) describe how the picture is partitioned into multiple tile partitions. · tile_cols_width_minus1 specifies the width of each tile in units of CTB when the picture is evenly divided into tiles. · tile_rows_height_minus1 specifies the height of each tile in units of CTB when the picture is evenly divided into tiles. · num_tile_columns_minus1 specifies the number of tile columns in the picture. · num_tile_rows_minus1 specifies the number of tile rows in the picture.

[0142] Pictures that refer to a PPS are divided into a plurality of tile partitions specified according to the above syntax elements. If two slices of the same picture refer to two different PPSs, this can cause a conflicting situation because each PPS may indicate a different type of tile partition of the picture. Therefore, it is prohibited for two slices (or generally two partitions of a picture) to refer to two different PPSs.

[0143] In VVC, various picture partitioning mechanisms are defined, which are called slices, tiles, bricks, and sub - pictures. In this application, picture partition is a general term indicating any of a slice, a tile, a brick, or a sub - picture. A picture partition usually indicates a part of a frame that is coded independently of other parts within the same picture.

[0144] Adaptive Parameter Set In the prior - art video codec, the bitstream is composed of a sequence of data units called network abstraction layer (NAL) units. Some NAL units contain parameter sets that carry high - level information regarding the entire coded video sequence or a subset of pictures therein. Other NAL units carry samples coded in the form of slices belonging to one of various picture types. An adaptive parameter set (APS) is a parameter set used to encapsulate ALF filter control data (e.g., filter coefficients).

[0145] FIG. 8 illustrates an adaptive parameter set according to the document Versatile Video Coding (Draft 6) of Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 (available at http: / / phenix.it-sudparis.eu / jvet / , document number: Document:JVET-O2001-vE).

[0146] Slice header The slice segment header contains an index to the reference PPS. The slice segment header contains data identifying the start address of the slice. It includes the slice type (I, P, or B), picture output flag, etc., and some parameters are included only in the first slice segment of the slice. Enabling SAO separately for luma and chroma, enabling the deblocking filter operation between slices, and including the initial slice quantization parameter (QP) value, and the presence of some coding tool parameters is in the slice header if the tool is enabled in the SPS or PPS. The deblocking filter parameters may be present in either the slice segment header or the PPS.

[0147] Coded slice Each coded slice typically consists of a slice header followed by slice data. The slice header carries control information for the slice, and the slice data carries the coded samples. FIGS. 9 and 11 illustrate the slice header according to the document JVET-O2001-vE. Further, FIG. 10a illustrates the syntax structure of a coding tree unit which is part of the slice data.

[0148] Each slice is independent of other slices in the sense that the information carried by the slice is coded without depending on data from other slices within the same picture.

[0149] Picture header The picture header (PH, Picture header) is a syntax structure that contains syntax elements applicable to all slices of a coded picture. Figure 10b illustrates a picture header according to document JVET-P1006 (http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P1006-v2.zip). For example, pic_temporal_mvp_enabled_flag specifies whether the temporal motion vector predictor can be used for inter prediction for all slices related to the picture header. When pic_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slices related to the picture header shall be constrained such that the temporal motion vector predictor is not used in the decoding of the slices. Otherwise (when pic_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used in the decoding of the slices related to the picture header.

[0150] The picture header is designed for all picture types, i.e., I pictures, P pictures, and B pictures. An I picture contains only intra prediction blocks, while a P or B picture contains inter prediction blocks. The difference between a P picture and a B picture is that a B picture can contain inter prediction blocks with bidirectional prediction, while in a P picture, only unidirectional inter prediction is permitted. Note that intra prediction blocks may also exist in a P or B picture.

[0151] The picture header has a significant number of syntax elements that are applicable only to inter prediction, i.e., only to P and B pictures. These elements are not necessary for I pictures.

[0152] Embodiments of the present invention propose introducing a flag indicating whether the current picture is an I picture. When the current picture is an I picture, all syntax elements designed for inter prediction are not signaled in the picture header and are assumed to have default values.

[0153] In one example, the proposed flag is named ph_all_intra_flag. The use of this flag is shown as follows. [Table 1-1] [Table 1-2]

[0154] A partition_constraints_override_flag equal to 1 specifies that there are partition constraint parameters in the picture header. A partition_constraints_override_flag equal to 0 specifies that there are no partition constraint parameters in the picture header. When not present, the value of the partition_constraints_override_flag is assumed to be equal to 0.

[0155] pic_log2_diff_min_qt_min_cb_intra_slice_luma specifies the difference between the base-2 logarithm of the minimum size of luma samples in luma residual blocks resulting from the quadtree partitioning of a CTU, and the base-2 logarithm of the minimum coding block size of luma samples for luma CUs in a slice having a slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_min_qt_min_cb_intra_slice_luma shall be in the range from 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, the value of pic_log2_diff_min_qt_min_cb_luma is assumed to be equal to sps_log2_diff_min_qt_min_cb_intra_slice_luma.

[0156] pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the base-2 logarithm of the minimum size of luma samples in luma residual blocks resulting from the quadtree partitioning of a CTU, and the base-2 logarithm of the minimum luma coding block size of luma samples for luma CUs in a slice having a slice_type equal to 0(B) or 1(P) associated with the picture header. The value of pic_log2_diff_min_qt_min_cb_inter_slice shall be in the range from 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, the value of pic_log2_diff_min_qt_min_cb_luma is assumed to be equal to sps_log2_diff_min_qt_min_cb_inter_slice.

[0157] pic_max_mtt_hierarchy_depth_inter_slice specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves within a slice having a slice_type equal to 0 (B) or 1 (P) related to the picture header. The value of pic_max_mtt_hierarchy_depth_inter_slice shall be in the range of 0 or more and CtbLog2SizeY - MinCbLog2SizeY or less. When it does not exist, the value of pic_max_mtt_hierarchy_depth_inter_slice is assumed to be equal to sps_max_mtt_hierarchy_depth_inter_slice.

[0158] pic_max_mtt_hierarchy_depth_intra_slice_luma specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves within a slice having a slice_type equal to 2 (I) related to the picture header. The value of pic_max_mtt_hierarchy_depth_intra_slice_luma shall be in the range of 0 or more and CtbLog2SizeY - MinCbLog2SizeY or less. When it does not exist, the value of pic_max_mtt_hierarchy_depth_intra_slice_luma is assumed to be equal to sps_max_mtt_hierarchy_depth_intra_slice_luma.

[0159] pic_log2_diff_max_bt_min_qt_intra_slice_luma specifies the difference between the base-2 logarithm of the maximum size (width or height) in luma samples of a luma coding block that can be split using binary tree splitting, and the minimum size (width or height) in luma samples of a luma leaf block resulting from the quadtree splitting of a CTU within a slice having a slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_max_bt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY. When it does not exist, the value of pic_log2_diff_max_bt_min_qt_intra_slice_luma is assumed to be equal to sps_log2_diff_max_bt_min_qt_intra_slice_luma.

[0160] pic_log2_diff_max_tt_min_qt_intra_slice_luma specifies the difference between the base-2 logarithm of the maximum size (width or height) in luma samples of a luma coding block that can be split using ternary tree splitting, and the minimum size (width or height) in luma samples of a luma leaf block resulting from the quadtree splitting of a CTU within a slice having a slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_max_tt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY. When it does not exist, the value of pic_log2_diff_max_tt_min_qt_intra_slice_luma is assumed to be equal to sps_log2_diff_max_tt_min_qt_intra_slice_luma.

[0161] pic_log2_diff_max_bt_min_qt_inter_slice specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luma samples of a luma coding block that can be divided using binary tree partitioning, and the minimum size (width or height) of the luma samples of a luma residual block resulting from the quadtree partitioning of a CTU within a slice having a slice_type equal to 0 (B) or 1 (P) associated with the picture header. The value of pic_log2_diff_max_bt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeInterY. When it does not exist, the value of pic_log2_diff_max_bt_min_qt_inter_slice is assumed to be equal to sps_log2_diff_max_bt_min_qt_inter_slice.

[0162] pic_log2_diff_max_tt_min_qt_inter_slice specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luma samples of a luma coding block that can be divided using ternary tree partitioning, and the minimum size (width or height) of the luma samples of a luma residual block resulting from the quadtree partitioning of a CTU within a slice having a slice_type equal to 0 (B) or 1 (P) associated with the picture header. The value of pic_log2_diff_max_tt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeInterY. When it does not exist, the value of pic_log2_diff_max_tt_min_qt_inter_slice is assumed to be equal to sps_log2_diff_max_tt_min_qt_inter_slice.

[0163] pic_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the difference between the logarithm to the base 2 of the minimum size of luma samples in chroma leaf blocks resulting from the quadtree partitioning of chroma CTUs having treeType equal to DUAL_TREE_CHROMA, and the logarithm to the base 2 of the minimum coding block size of luma samples for chroma CUs having treeType equal to DUAL_TREE_CHROMA within a slice having slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_min_qt_min_cb_intra_slice_chroma shall be in the range of 0 or more and CtbLog2SizeY - MinCbLog2SizeY or less. When it does not exist, the value of pic_log2_diff_min_qt_min_cb_intra_slice_chroma is assumed to be equal to sps_log2_diff_min_qt_min_cb_intra_slice_chroma.

[0164] pic_max_mtt_hierarchy_depth_intra_slice_chroma specifies the maximum hierarchical depth for chroma coding units resulting from the multi-type tree partitioning of chroma quadtree leaves having treeType equal to DUAL_TREE_CHROMA within a slice having slice_type equal to 2(I) associated with the picture header. The value of pic_max_mtt_hierarchy_depth_intra_slice_chroma shall be in the range of 0 or more and CtbLog2SizeY - MinCbLog2SizeY or less. When it does not exist, the value of pic_max_mtt_hierarchy_depth_intra_slice_chroma is assumed to be equal to sps_max_mtt_hierarchy_depth_intra_slice_chroma.

[0165] pic_log2_diff_max_bt_min_qt_intra_slice_chroma specifies the difference between the base-2 logarithm of the maximum size (width or height) in luma samples of a chroma coding block that can be partitioned using binary tree partitioning, and the minimum size (width or height) in luma samples of a chroma leaf block resulting from the quadtree partitioning of a chroma CTU having treeType equal to DUAL_TREE_CHROMA within a slice having slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_max_bt_min_qt_intra_slice_chroma shall be in the range from 0 to CtbLog2SizeY - MinQtLog2SizeIntraC. When it does not exist, the value of pic_log2_diff_max_bt_min_qt_intra_slice_chroma is assumed to be equal to sps_log2_diff_max_bt_min_qt_intra_slice_chroma.

[0166] pic_log2_diff_max_tt_min_qt_intra_slice_chroma specifies the difference between the base-2 logarithm of the maximum size (width or height) in luma samples of a chroma coding block that can be partitioned using ternary tree partitioning, and the minimum size (width or height) in luma samples of a chroma leaf block resulting from the quadtree partitioning of a chroma CTU having treeType equal to DUAL_TREE_CHROMA within a slice having slice_type equal to 2(I) associated with the picture header. The value of pic_log2_diff_max_tt_min_qt_intra_slice_chroma shall be in the range from 0 to CtbLog2SizeY - MinQtLog2SizeIntraC. When it does not exist, the value of pic_log2_diff_max_tt_min_qt_intra_slice_chroma is assumed to be equal to sps_log2_diff_max_tt_min_qt_intra_slice_chroma.

[0167] pic_cu_qp_delta_subdiv_intra_slice specifies the maximum cbSubdiv value of coding units within an intra slice that transmit cu_qp_delta_abs and cu_qp_delta_sign_flag. The value of pic_cu_qp_delta_subdiv_intra_slice shall be in the range from 0 to 2*(CtbLog2SizeY - MinQtLog2SizeIntraY + pic_max_mtt_hierarchy_depth_intra_slice_luma).

[0168] When it does not exist, the value of pic_cu_qp_delta_subdiv_intra_slice is assumed to be equal to 0.

[0169] pic_cu_qp_delta_subdiv_inter_slice specifies the maximum cbSubdiv value of coding units that transmit cu_qp_delta_abs and cu_qp_delta_sign_flag within an inter slice. The value of pic_cu_qp_delta_subdiv_inter_slice shall be in the range from 0 to 2*(CtbLog2SizeY - MinQtLog2SizeInterY + pic_max_mtt_hierarchy_depth_inter_slice).

[0170] When it does not exist, the value of pic_cu_qp_delta_subdiv_inter_slice is assumed to be equal to 0.

[0171] pic_cu_chroma_qp_offset_subdiv_intra_slice specifies the maximum cbSubdiv value of the coding unit within an intra slice that conveys the cu_chroma_qp_offset_flag. The value of pic_cu_chroma_qp_offset_subdiv_intra_slice shall be in the range of 0 to 2*(CtbLog2SizeY - MinQtLog2SizeIntraY + pic_max_mtt_hierarchy_depth_intra_slice_luma).

[0172] When it does not exist, the value of pic_cu_chroma_qp_offset_subdiv_intra_slice is assumed to be equal to 0.

[0173] pic_cu_chroma_qp_offset_subdiv_inter_slice specifies the maximum cbSubdiv value of the coding unit within an inter slice that conveys the cu_chroma_qp_offset_flag. The value of pic_cu_chroma_qp_offset_subdiv_inter_slice shall be in the range of 0 to 2*(CtbLog2SizeY - MinQtLog2SizeInterY + pic_max_mtt_hierarchy_depth_inter_slice).

[0174] When it does not exist, the value of pic_cu_chroma_qp_offset_subdiv_inter_slice is assumed to be equal to 0.

[0175] The pic_temporal_mvp_enabled_flag specifies whether the temporal motion vector predictor can be used for inter prediction for the slice associated with the picture header. If pic_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slice associated with the picture header shall be constrained such that the temporal motion vector predictor is not used in the decoding of the slice. Otherwise (if pic_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used in the decoding of the slice associated with the picture header.

[0176] When pic_temporal_mvp_enabled_flag does not exist, the following applies. - If sps_temporal_mvp_enabled_flag is equal to 0, the value of pic_temporal_mvp_enabled_flag is assumed to be equal to 0. - Otherwise (if sps_temporal_mvp_enabled_flag is equal to 1), the value of pic_temporal_mvp_enabled_flag is assumed to be equal to pps_temporal_mvp_enabled_idc - 1.

[0177] mvd_l1_zero_flag equal to 1 indicates that the mvd_coding(x0,y0,1) syntax structure is not parsed and MvdL1[x0][y0][compIdx] and MvdL1[x0][y0][cpIdx][compIdx] are set to 0 for compIdx = 0..1 and cpIdx = 0..2. mvd_l1_zero_flag equal to 0 indicates that the mvd_coding(x0,y0,1) syntax structure is parsed. When it does not exist, the value of mvd_l1_zero_flag is assumed to be equal to pps_mvd_l1_zero_idc - 1.

[0178] pic_six_minus_max_num_merge_cand specifies the maximum number of merge motion vector prediction (MVP) candidates supported in the slices related to the picture header, subtracted from 6. The maximum number of merge MVP candidates, MaxNumMergeCand, is derived as follows. MaxNumMergeCand = 6 - pic_six_minus_max_num_merge_cand (7 - 111)

[0179] The value of MaxNumMergeCand shall be in the range of 1 to 6. When it does not exist, the value of pic_six_minus_max_num_merge_cand is assumed to be equal to pps_six_minus_max_num_merge_cand_plus1 - 1.

[0180] pic_five_minus_max_num_subblock_merge_cand specifies the maximum number of sub - block - based merge motion vector prediction (MVP) candidates supported in the slice, subtracted from 5.

[0181] When pic_five_minus_max_num_subblock_merge_cand does not exist, the following applies. - If - sps_affine_enabled_flag is equal to 0, the value of pic_five_minus_max_num_subblock_merge_cand is assumed to be equal to 5 - (sps_sbtmvp_enabled_flag && pic_temporal_mvp_enabled_flag). - Otherwise (if sps_affine_enabled_flag is equal to 1), the value of pic_five_minus_max_num_subblock_merge_cand is assumed to be equal to pps_five_minus_max_num_subblock_merge_cand_plus1 - 1.

[0182] The maximum number of sub-block based merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows. MaxNumSubblockMergeCand = 5 - pic_five_minus_max_num_subblock_merge_cand (7 - 112)

[0183] The value of MaxNumSubblockMergeCand shall be in the range from 0 to 5.

[0184] pic_fpel_mmvd_enabled_flag equal to 1 specifies that the merge mode with motion vector difference uses integer sample accuracy in the slice related to the picture header. pic_fpel_mmvd_enabled_flag equal to 0 specifies that the merge mode with motion vector difference can use fractional sample accuracy in the slice related to the picture header. When it does not exist, the value of pic_fpel_mmvd_enabled_flag is presumed to be 0.

[0185] pic_disable_bdof_dmvr_flag equal to 1 specifies that neither the bidirectional optical flow inter prediction nor the decoder motion vector refinement based inter bidirectional prediction is enabled in the slice related to the picture header. pic_disable_bdof_dmvr_flag equal to 0 specifies that either the bidirectional optical flow inter prediction or the decoder motion vector refinement based inter bidirectional prediction may or may not be enabled in the slice related to the picture header. When it does not exist, the value of pic_disable_bdof_dmvr_flag is presumed to be 0.

[0186] pic_max_num_merge_cand_minus_max_num_triangle_cand specifies the maximum number of triangle merge mode candidates supported in the slice related to the picture header, subtracted from MaxNumMergeCand.

[0187] When pic_max_num_merge_cand_minus_max_num_triangle_cand does not exist, sps_triangle_enabled_flag is equal to 1, and MaxNumMergeCand is 2 or more, pic_max_num_merge_cand_minus_max_num_triangle_cand is presumed to be equal to pps_max_num_merge_cand_minus_max_num_triangle_cand_plus1 - 1.

[0188] The maximum number of triangle merge mode candidates, MaxNumTriangleMergeCand, is derived as follows. MaxNumTriangleMergeCand = MaxNumMergeCand - pic_max_num_merge_cand_minus_max_num_triangle_cand (7 - 113)

[0189] When pic_max_num_merge_cand_minus_max_num_triangle_cand exists, the value of MaxNumTriangleMergeCand shall be in the range of 2 or more and MaxNumMergeCand or less.

[0190] When pic_max_num_merge_cand_minus_max_num_triangle_cand does not exist, (sps_triangle_enabled_flag is equal to 0 or MaxNumMergeCand is less than 2), MaxNumTriangleMergeCand is set to 0.

[0191] When MaxNumTriangleMergeCand is equal to 0, the triangular merge mode is not permitted for the slices associated with the picture header.

[0192] pic_six_minus_max_num_ibc_merge_cand specifies the maximum number of Inter Block Copy (IBC) merge block vector prediction (BVP) candidates supported in the slices associated with the picture header, subtracted from 6. The maximum number of IBC merge BVP candidates, MaxNumIbcMergeCand, is derived as follows. MaxNumIbcMergeCand = 6 - pic_six_minus_max_num_ibc_merge_cand (7-114)

[0193] The value of MaxNumIbcMergeCand shall be in the range of 1 to 6 inclusive.

[0194] Note that the ph_all_intra_flag can be signaled anywhere before the first place that controls the syntax elements related to inter prediction, i.e., pic_log2_diff_min_qt_min_cb_inter_slice. Note that one or more places of ph_all_intra_flag can be removed (i.e., only control a subset of the syntax elements of the embodiments of the present invention). As another example, the use of this flag is shown as follows.

Table 2

[0195] In particular, the following methods and embodiments implemented by an encoding device are provided. The encoding device may be the video encoder 20 of FIG. 1A or the encoder 20 of FIG. 2.

[0196] According to Embodiment 1200 (refer to FIG. 12), in step 1201, the device determines whether the current picture is an I picture.

[0197] Since an I picture contains only intra prediction blocks, there is no need to signal syntax elements designed for inter prediction in the picture header of the bitstream. Therefore, when the current picture is an I picture, the syntax elements designed for inter prediction are not signaled in the picture header. In step 1203, the device sends the bitstream to the decoding device, and the picture header of the bitstream includes a flag used to indicate whether the current picture is an I picture. In this situation, the flag indicates that the current picture is an I picture. The syntax elements designed for inter prediction are not signaled in the picture header. When the current picture is an I picture, the syntax elements designed for inter prediction are estimated to be default values.

[0198] When the current picture is not an I picture, i.e., when the current picture is a P or B picture, in step 1205, the device obtains syntax elements designed for inter prediction. As described above, the syntax elements designed for inter prediction include one or more of the following elements, namely, pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice, pic_cu_qp_delta_subdiv_inter_slice, pic_cu_chroma_qp_offset_subdiv_inter_slice, pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_fpel_mmvd_enabled_flag or pic_disable_bdof_dmvr_flag.

[0199] In step 1207, the device sends the bitstream to the decoding device. The picture header of the bitstream includes not only a flag used to indicate whether the current picture is an I picture, but also syntax elements designed for inter prediction. In this situation, the flag indicates that the current picture is not an I picture.

[0200] For example, the flag is named ph_all_intra_flag. The ph_all_intra_flag is properly signaled anywhere before the first place that controls the syntax elements designed for inter prediction. The use of this flag is shown above.

[0201] The following methods and embodiments implemented by a decoding device are provided. The decoding device may be the video decoder 30 of FIG. 1A or the decoder 30 of FIG. 3. According to embodiment 1300 (see FIG. 13), in step 1301, the device receives a bitstream, parses the bitstream to obtain a flag from the picture header of the bitstream, and the flag indicates whether the current picture is an I picture.

[0202] For example, the flag is named ph_all_intra_flag. The ph_all_intra_flag flag is properly signaled anywhere before the first place that controls the syntax element designed for inter prediction. The use of this flag is shown above.

[0203] In step 1303, the device determines whether the current picture is an I picture based on the flag.

[0204] When the current picture is an I picture, the syntax element designed for inter prediction is not signaled in the picture header, and in step 1305, the syntax element designed for inter prediction is estimated to a default value.

[0205] When the current picture is not an I picture, i.e., when the current picture is a P or B picture, at step 1307, the device obtains syntax elements designed for inter prediction from the picture header of the bitstream. As described above, the syntax elements designed for inter prediction include one or more of the following elements: pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice, pic_cu_qp_delta_subdiv_inter_slice, pic_cu_chroma_qp_offset_subdiv_inter_slice, pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_fpel_mmvd_enabled_flag or pic_disable_bdof_dmvr_flag.

[0206] FIG. 14 shows an embodiment of device 1400. Device 1400 may be the video encoder 20 of FIG. 1A or the encoder 20 of FIG. 2. Device 1400 can be used to implement embodiment 1200 and the other embodiments described above.

[0207] Device 1400 according to the present disclosure includes a determination unit 1401, an acquisition unit 1402, and a signaling unit 1403. The determination unit 1401 is configured to determine whether the current picture is an I picture.

[0208] When the current picture is not an I picture, i.e., when it is a P or B picture, the acquisition unit 1042 is configured to acquire syntax elements designed for inter prediction. As described above, the syntax elements designed for inter prediction include one or more of the following elements, namely, pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice, pic_cu_qp_delta_subdiv_inter_slice, pic_cu_chroma_qp_offset_subdiv_inter_slice, pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_fpel_mmvd_enabled_flag or pic_disable_bdof_dmvr_flag.

[0209] When the current picture is an I picture, the syntax elements designed for inter prediction are estimated to default values.

[0210] The signaling unit 1403 is configured to send the bitstream to the decoding device, and the picture header of the bitstream includes a flag used to indicate whether the current picture is an I picture.

[0211] The picture header of the bitstream further includes syntax elements designed for inter prediction when the current picture is not an I picture. Since an I picture only includes intra prediction blocks, there is no need to signal the syntax elements designed for inter prediction in the picture header of the bitstream.

[0212] For example, the flag is named ph_all_intra_flag. The ph_all_intra_flag is signaled appropriately anywhere before the first place that controls the syntax element designed for inter prediction. The use of this flag is shown above.

[0213] FIG. 15 shows an embodiment of device 1500. Device 1500 may be the video decoder 30 of FIG. 1A or the decoder 30 of FIG. 3. Device 1500 can be used to implement embodiment 1300 and other embodiments above.

[0214] Device 1500 includes an acquisition unit 1501 and a determination unit 1502. The acquisition unit 1501 is configured to parse the bitstream to obtain a flag from the picture header of the bitstream, and the flag indicates whether the current picture is an I picture. For example, the flag is named ph_all_intra_flag.

[0215] The determination unit 1502 is configured to determine whether the current picture is an I picture based on the flag.

[0216] When the flag indicates that the current picture is not an I picture, i.e., a P or B picture, the acquisition unit 1501 is further configured to obtain the syntax element designed for inter prediction from the picture header. When the current picture is an I picture, the syntax element designed for inter prediction is not signaled in the picture header, and the syntax element designed for inter prediction is estimated to be a default value.

[0217] Furthermore, the following embodiments are provided here.

[0218] Embodiment 1. A coding method implemented by a decoding device, comprising: Parsing a bitstream; A step of obtaining a flag from a picture header of a bitstream, the flag indicating whether the current picture is an I picture, and a step A method including

[0219] Embodiment 2. The method according to Embodiment 1, wherein when the current picture is an I picture, the syntax elements designed for inter prediction are estimated to default values.

[0220] Embodiment 3. The method according to Embodiment 1 or 2, wherein the flag is carried in the PBSP syntax of the picture header.

[0221] Embodiment 4. The method according to any one of Embodiments 1 to 3, wherein the flag is named ph_all_intra_flag.

[0222] Embodiment 5. The method according to Embodiment 4, wherein the use of the flag is as follows.

Table 3

[0223] Embodiment 6. The method according to Embodiment 4 or 5, wherein the use of the flag is as follows.

Table 4

[0224] Embodiment 7. The method according to any one of Embodiments 4 to 6, wherein the use of the flag is as follows.

Table 5

[0225] Embodiment 8. The method according to any one of Embodiments 4 to 7, wherein the use of the flag is as follows.

Table 6

[0226] The method according to any one of Embodiments 1 to 8, wherein the I picture includes only intra prediction blocks, while the P or B picture includes inter prediction blocks.

[0227] A coding method implemented by an encoding device, comprising: signaling a flag in the picture header of the bitstream, the flag indicating whether the current picture is an I picture; transmitting the bitstream; and a method comprising.

[0228] The method according to Embodiment 10, wherein when the current picture is an I picture, all syntax elements designed for inter prediction are not signaled in the picture header.

[0229] The method according to Embodiment 10 or 11, wherein when the current picture is an I picture, the syntax elements designed for inter prediction are estimated to default values.

[0230] The method according to any one of Embodiments 10 to 12, wherein the flag is signaled in the PBSP syntax of the picture header.

[0231] The method according to any one of Embodiments 10 to 13, wherein the flag is named ph_all_intra_flag.

[0232] The method according to Embodiment 14, wherein the ph_all_intra_flag is properly signaled anywhere before the first location for controlling syntax elements related to inter prediction.

[0233] The method according to Embodiment 14 or 15, wherein the use of the flag is as follows.

Table 7

[0234] Embodiment 17. The use of the flag is as follows, the method described in any one of Embodiments 14 to 16.

Table 8

[0235] Embodiment 18. The use of the flag is as follows, the method described in any one of Embodiments 14 to 17.

Table 9

[0236] Embodiment 19. The use of the flag is as follows, the method described in any one of Embodiments 14 to 18.

Table 10

[0237] As described above, by indicating whether the current picture is an I picture in the picture header of the bitstream, when the current picture is an I picture, the syntax elements designed for inter prediction are not signaled in the picture header. Therefore, the embodiment can simplify the signaling of the picture header for all intra pictures, that is, I pictures. Correspondingly, the signaling overhead is reduced.

[0238] The term "obtain" may indicate receiving (explicitly as respective parameters from other entities / devices or modules within the same device, e.g., by parsing a bitstream for a decoder) and / or deriving (which may also be referred to as implicitly receiving, e.g., by parsing other information / parameters from a bitstream for a decoder and deriving respective information or parameters from such other information / parameters).

[0239] The following is an explanation of the encoding method, decoding method, and their application in the above embodiments, as well as the system using them.

[0240] FIG. 16 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 over a communication link 3104. The communication link may include the above communication channel 13. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types.

[0241] The capture device 3102 may generate data and encode the data by an encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may distribute the data to a streaming server (not shown in the drawings), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. In some actual scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.

[0242] In the content supply system 3100, the terminal device 310 receives and plays the encoded data. The terminal device 3106 may be a device having data reception and restoration capabilities, such as a smartphone or a pad 3108 capable of decrypting the above encoded data, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set top box (STB) 3116, a video conferencing system 3118, a video monitoring system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.

[0243] In a terminal device having its own display, such as a smartphone or a pad 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video decoder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its own display. In a terminal device without a display, such as an STB 3116, a video conferencing system 3118, or a video monitoring system 3120, an external display 3126 is contacted to itself to receive and display the decoded data.

[0244] When each device in this system performs encoding or decoding, as shown in the above embodiments, a picture encoding device or a picture decoding device can be used.

[0245] FIG. 17 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, the Real Time Streaming Protocol (RTSP), the Hyper Text Transfer Protocol (HTTP), the HTTP Live Streaming protocol (HLS), MPEG-DASH, the Real-time Transport protocol (RTP), the Real Time Messaging Protocol (RTMP), or any combination of these types.

[0246] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0247] Through inverse multiplexing processing, a video elementary stream (ES), an audio ES, and optional subtitles are generated. A video decoder 3206 including a video decoder 30 as described in the above embodiments decodes the video ES by a decoding method as shown in the above embodiments to generate video frames, and supplies this data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. 21) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in FIG. 21) before being supplied to the synchronization unit 3212.

[0248] The synchronization unit 3212 synchronizes the video frames and the audio frames, and supplies the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in the syntax using time stamps related to the presentation of the coded audio and visual data and time stamps related to the delivery of the data stream itself.

[0249] When subtitles are included in the stream, a subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frames and the audio frames, and supplies the video / audio / subtitle to a video / audio / subtitle display 3216.

[0250] The present invention is not limited to the above system, and either the picture encoding device or the picture decoding device in the above embodiments can be incorporated into other systems, such as vehicle systems.

[0251] Mathematical operator The mathematical operators used in this application are the same as those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operators such as exponentiation and division of real values are defined. The numbering and counting rules generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 1st, and so on.

[0252] Logical operators The following logical operators are defined as follows.

Table 11

[0253] Logical operators The following logical operators are defined as follows. x&&y The Boolean logical "product" of x and y x||y The Boolean logical "sum" of x and y ! Boolean logical "negation" x?y:z If x is true or not equal to 0, it is evaluated to the value of y; otherwise, it is evaluated to the value of z

[0254] Relational operators The following relational operators are defined as follows. > Greater than >= Greater than or equal to < Less than <= Less than or equal to == Equal to != Not equal to When a relational operator is applied to a syntax element or variable to which the value "na" (not applicable) is assigned, the value "na" is treated as the individual value of the syntax element or variable. The value "na" is considered not equal to any other value.

[0255] Bitwise operators The following bitwise operators are defined as follows. & Bitwise "product". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding higher-order bits equal to 0. | Bitwise "sum". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding higher-order bits equal to 0. ^ Bitwise "exclusive sum". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding higher-order bits equal to 0. x>>y Arithmetic right shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has a value equal to the MSB of x before the shift operation. x<<y Arithmetic left shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.

[0256] Assignment operators The following assignment operators are defined as follows. = Assignment operator ++ Increment. That is, x++ is equal to x=x+1. When used in an array index, it is evaluated to the value of the variable before the increment operation. -- Decrement. That is, x-- is equal to x=x-1. When used in an array index, it is evaluated to the value of the variable before the decrement operation. += Increment by the specified amount. That is, x += 3 is equivalent to x = x + 3, and x += (-3) is equivalent to x = x + (-3). -= Decrement by the specified amount. That is, x -= 3 is equivalent to x = x - 3, and x -= (-3) is equivalent to x = x - (-3).

[0257] Range notation The following notations are used to specify a range of values. x = y..z x takes integer values greater than or equal to y and less than or equal to z, where x, y, and z are integers and z is greater than y.

[0258] Mathematical functions The following mathematical functions are defined.

Number

Number

Number

Number

Number

Number

Number

[0259] Operator precedence When the precedence of an expression is not explicitly indicated by the use of parentheses, the following rules apply. - Operations with higher precedence are evaluated before any operations with lower precedence. - Operations with the same precedence are evaluated sequentially from left to right.

[0260] The following table specifies the precedence of operations from highest to lowest, where a higher position in the table indicates a higher precedence.

[0261] For operators also used in the C programming language, the precedence used in this specification is the same as the precedence used in the C programming language.

Table 12

[0262] Text description of logical operations In the text, the following format: if (condition 0) Statement 0 else (condition 1) Statement 1 ... else / * Reference note regarding the remaining conditions * / Statement n A statement of logical operation described mathematically as follows may be described in the following manner. ... As follows / ... The following applies: - If condition 0, then statement 0 - Otherwise, if condition 1, then statement 1 -... - Otherwise (reference note regarding the remaining conditions), statement n Each "if... then... else if... then... else..." statement in the text is introduced by "if... then... as follows" or "if... then... the following applies" immediately following "if...". The last condition of "if... then... else if... then... else..." is always "otherwise,...". The alternating "if... then... else if... then... else..." statements can be identified by matching them with "otherwise,..." at the end of "if... then... as follows" or "if... then... the following applies".

[0263] In the text, the following format: if (condition 0a && condition 0b) Statement 0 else if (condition 1a || condition 1b) Statement 1 ... else Statement n A statement of logical operation described mathematically as follows may be described in the following manner. ... As follows / ... The following applies: - If all of the following conditions are true, then statement 0: - Condition 0a - Condition 0b - Otherwise, if one or more of the following conditions are true, Statement 1: - Condition 1a - Condition 1b - … - Otherwise, Statement n

[0264] In the text, for the following format: if (Condition 0) Statement 0 if (Condition 1) Statement 1 Logical operation statements mathematically described in the form of, for example, may be described in the following manner. When Condition 0, Statement 0 When Condition 1, Statement 1

[0265] Although embodiments of the present invention have been mainly described based on video coding, embodiments of the coding system 10, the encoder 20, and the decoder 30 (and correspondingly the system 10), as well as other embodiments described herein, may also be configured for still picture processing or coding, i.e., processing or coding of individual pictures independent of any previous or subsequent pictures as in video coding. It should be noted that generally, when picture processing coding is limited to a single picture 17, not all of the inter prediction units 244 (encoder) and 344 (decoder) may be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 may be equally used for still picture processing, for example, residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, as well as entropy coding 270 and entropy decoding 304.

[0266] For example, embodiments of the encoder 20 and decoder 30, and the functions described herein with respect to the encoder 20 and decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted over a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0267] By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transitory, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data with a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0268] The commands may be executed by one or more processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Further, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding or incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuits or logic elements.

[0269] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize the functional aspects of devices configured to execute the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as noted above, the various units may be combined with appropriate software and / or firmware and coupled to a codec hardware unit, or provided by a collection of interoperable hardware units including one or more of the processors as described above.

Claims

1. 1. A method of coding implemented by an encoding device, comprising: encoding a first flag and a second flag into a picture header of a bitstream, the first flag indicating that all coded slices of a current picture have a slice type equal to 2 (I) or that one or more coded slices in the current picture have a slice type equal to 0 (B) or 1 (P), and the second flag indicating whether a partition constraint parameter is present in the picture header; encoding a syntax element designed for inter prediction into the picture header when the first flag indicates that the one or more coded slices in the current picture have the slice type equal to 0 (B) or 1 (P) and the second flag indicates that the partition constraint parameters are present in the picture header, the second flag being a partition_constraints_override_flag; The method includes:

2. 2. The method of claim 1, wherein when all coded slices of the current picture have the slice type equal to 2(I), the syntax elements designed for inter prediction are inferred to be equal to a value of a syntax element designed for inter prediction in a sequence parameter set (SPS) of the bitstream.

3. The syntax elements designed for inter prediction include pic_log2_diff_min_qt_min_cb_inter_slice; pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the logarithm of the base 2 of the minimum size in luma samples of a luma reef block resulting from a quadtree partitioning of a coding tree unit (CTU) and the logarithm of the base 2 of the minimum luma coding block size in luma samples for a luma coding unit (CU) in a slice having a P or B picture associated with the picture header, or 3. The method of claim 1 or 2, wherein when pic_log2_diff_min_qt_min_cb_inter_slice is not present, the value of pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to sps_log2_diff_min_qt_min_cb_inter_slice.

4. The syntax elements designed for inter prediction include pic_max_mtt_hierarchy_depth_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice specifies the maximum hierarchical depth for coding units resulting from a multitype tree split of quadtree leaves within a slice, or 4. A method according to claim 1, wherein when pic_max_mtt_hierarchy_depth_inter_slice is not present, the value of pic_max_mtt_hierarchy_depth_inter_slice is inferred to be equal to sps_max_mtt_hierarchy_depth_inter_slice.

5. When pic_max_mtt_hierarchy_depth_inter_slice is not equal to 0, the syntax elements designed for inter prediction further include pic_log2_diff_max_bt_min_qt_inter_slice; pic_log2_diff_max_bt_min_qt_inter_slice specifies the difference between the logarithm in base 2 of the maximum size in luma samples (width or height) of a luma coding block that can be divided using binary tree partitioning and the logarithm in base 2 of the minimum size in luma samples (width or height) of a luma reef block resulting from quadtree partitioning of CTUs in the slice, or 5. The method of claim 4, wherein when pic_log2_diff_max_bt_min_qt_inter_slice is not present, the value of pic_log2_diff_max_bt_min_qt_inter_slice is inferred to be equal to sps_log2_diff_max_bt_min_qt_inter_slice.

6. When pic_max_mtt_hierarchy_depth_inter_slice is not equal to 0, the syntax elements designed for inter prediction further include pic_log2_diff_max_tt_min_qt_inter_slice; pic_log2_diff_max_tt_min_qt_inter_slice specifies the difference between the base 2 logarithm of the maximum size in luma samples (width or height) of a luma coding block that can be divided using ternary tree decomposition and the base 2 logarithm of the minimum size in luma samples (width or height) of a luma reef block resulting from quadtree decomposition of CTUs in a slice, or The method of claim 4 or 5, wherein when pic_log2_diff_max_tt_min_qt_inter_slice is not present, the value of pic_log2_diff_max_tt_min_qt_inter_slice is inferred to be equal to sps_log2_diff_max_tt_min_qt_inter_slice.

7. The syntax elements designed for inter prediction include pic_cu_qp_delta_subdiv_inter_slice, pic_cu_qp_delta_subdiv_inter_slice specifies the maximum cbSubdiv value of the coding units that carry cu_qp_delta_abs and cu_qp_delta_sign_flag in the inter slice, or 7. The method of claim 1, wherein when pic_cu_qp_delta_subdiv_inter_slice is not present, the value of pic_cu_qp_delta_subdiv_inter_slice is inferred to be equal to 0.

8. The syntax elements designed for inter prediction include pic_cu_chroma_qp_offset_subdiv_inter_slice; pic_cu_chroma_qp_offset_subdiv_inter_slice specifies the maximum cbSubdiv value of the coding unit in the inter slice that carries cu_chroma_qp_offset_flag, or The method of claim 1 , wherein when pic_cu_chroma_qp_offset_subdiv_inter_slice is not present, the value of pic_cu_chroma_qp_offset_subdiv_inter_slice is inferred to be equal to 0.

9. The syntax elements designed for inter prediction include pic_temporal_mvp_enabled_flag; pic_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction for the slice associated with said picture header, or 9. The method of claim 1, wherein when pic_temporal_mvp_enabled_flag is not present, the value of pic_temporal_mvp_enabled_flag is inferred to be equal to 0.

10. An encoder comprising processing circuitry for carrying out the method of any one of claims 1 to 9.

11. 1. An encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors to cause the one or more processors to perform the method of any one of claims 1 to 9; Including the encoder.

12. 1. A coding device, comprising: a coding unit configured to code a first flag and a second flag into a picture header of a bitstream, the first flag indicating that all coded slices of a current picture have a slice type equal to 2 (I) or that one or more coded slices in the current picture have a slice type equal to 0 (B) or 1 (P), and the second flag indicating whether a partition constraint parameter is present in the picture header; The encoding device, wherein the encoding unit is further configured to encode syntax elements designed for inter prediction into the picture header when the first flag indicates that the one or more coded slices in the current picture have the slice type equal to 0 (B) or 1 (P) and the second flag indicates that the partition constraint parameters are present in the picture header, the second flag being partition_constraints_override_flag.

13. The device of claim 12, wherein when all coded slices of the current picture are slices having a slice_type equal to I, the syntax elements designed for inter prediction are inferred to be equal to a value of a syntax element designed for inter prediction in a sequence parameter set (SPS) of the bitstream.

14. The syntax elements designed for inter prediction include pic_log2_diff_min_qt_min_cb_inter_slice; pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the logarithm of the base 2 of the minimum size in luma samples of a luma reef block resulting from a quadtree partitioning of a coding tree unit (CTU) and the logarithm of the base 2 of the minimum luma coding block size in luma samples for a luma coding unit (CU) in the slice having slice_type equal to 0 (B) or 1 (P) in the current picture; or A device according to claim 12 or 13, wherein when pic_log2_diff_min_qt_min_cb_inter_slice is not present, the value of pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to sps_log2_diff_min_qt_min_cb_inter_slice.

15. The syntax elements designed for inter prediction include pic_max_mtt_hierarchy_depth_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice specifies the maximum hierarchical depth for coding units resulting from a multi-type tree split of quadtree leaves in slices with slice_type equal to 0 (B) or 1 (P) in the current picture, or A device according to claim 12 , wherein when pic_max_mtt_hierarchy_depth_inter_slice is not present, the value of pic_max_mtt_hierarchy_depth_inter_slice is inferred to be equal to sps_max_mtt_hierarchy_depth_inter_slice.

16. When pic_max_mtt_hierarchy_depth_inter_slice is not equal to 0, the syntax elements designed for inter prediction further include pic_log2_diff_max_bt_min_qt_inter_slice; pic_log2_diff_max_bt_min_qt_inter_slice specifies the difference between the logarithm base 2 of the maximum size in luma samples (width or height) of a luma coding block that can be divided using binary tree partitioning and the logarithm base 2 of the minimum size in luma samples (width or height) of a luma reef block resulting from quadtree partitioning of CTUs in the slice with slice_type equal to 0 (B) or 1 (P) in the current picture; or The device of claim 15, wherein when pic_log2_diff_max_bt_min_qt_inter_slice is not present, the value of pic_log2_diff_max_bt_min_qt_inter_slice is inferred to be equal to sps_log2_diff_max_bt_min_qt_inter_slice.

17. When pic_max_mtt_hierarchy_depth_inter_slice is not equal to 0, the syntax elements designed for inter prediction further include pic_log2_diff_max_tt_min_qt_inter_slice; pic_log2_diff_max_tt_min_qt_inter_slice specifies the difference between the logarithm base 2 of the maximum size in luma samples (width or height) of a luma coding block that can be divided using ternary tree division and the logarithm base 2 of the minimum size in luma samples (width or height) of a luma reef block resulting from quadtree division of CTUs in a slice with slice_type equal to 0 (B) or 1 (P) in the current picture, or 17. A device according to claim 15 or 16, wherein when pic_log2_diff_max_tt_min_qt_inter_slice is not present, the value of pic_log2_diff_max_tt_min_qt_inter_slice is inferred to be equal to sps_log2_diff_max_tt_min_qt_inter_slice.

18. The syntax elements designed for inter prediction include pic_cu_qp_delta_subdiv_inter_slice, pic_cu_qp_delta_subdiv_inter_slice specifies the maximum cbSubdiv value of the coding units that carry cu_qp_delta_abs and cu_qp_delta_sign_flag in the inter slice, or 18. A device according to claim 12, wherein when pic_cu_qp_delta_subdiv_inter_slice is not present, the value of pic_cu_qp_delta_subdiv_inter_slice is inferred to be equal to 0.

19. The syntax elements designed for inter prediction include pic_cu_chroma_qp_offset_subdiv_inter_slice; pic_cu_chroma_qp_offset_subdiv_inter_slice specifies the maximum cbSubdiv value of the coding unit in the inter slice that carries cu_chroma_qp_offset_flag, or A device according to claim 12 to 18, wherein when pic_cu_chroma_qp_offset_subdiv_inter_slice is not present, the value of pic_cu_chroma_qp_offset_subdiv_inter_slice is inferred to be equal to 0.

20. A computer program comprising a program code for carrying out the method according to any one of claims 1 to 9.

21. A non-transitory computer readable storage medium comprising computer executable instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 9.

22. An apparatus for storing a bitstream, comprising one or more storage media and a receiver, comprising: The receiver is adapted to receive one or more bitstreams generated by a method according to any one of claims 1 to 9, The one or more storage media are configured to store the one or more bitstreams.

23. 1. A device for transmitting a bitstream, comprising: At least one storage medium adapted to store at least one bitstream generated by the method according to any one of claims 1 to 9; at least one processor configured to retrieve one or more bitstreams from one of the at least one storage medium and transmit the one or more bitstreams to another device; The device that contains

24. A device for receiving and storing a bitstream, comprising a receiver, a processor, and a storage medium, comprising: The receiver is configured to receive a bitstream; the storage medium is configured to store the bitstream; the bitstream includes a picture header and a coded slice of a current picture; the picture header includes a first flag and a second flag, the first flag indicating that all coded slices of a current picture have a slice type equal to 2 (I) or one or more coded slices in the current picture have a slice type equal to 0 (B) or 1 (P), the second flag indicating whether a partition constraint parameter is present in the picture header, and when one or more coded slices in the current picture have the slice type equal to 0 (B) or 1 (P), the picture header further includes a syntax element designed for inter prediction; The processor decodes the picture header to obtain the first flag; The device, wherein the processor further decodes the picture header to obtain the syntax elements designed for inter prediction when the first flag indicates that one or more coded slices in the current picture have the slice type equal to 0 (B) or 1 (P) and the second flag indicates that the partition constraint parameter is present in the picture header.