Encoder, decoder, and corresponding method

JP7927680B2Active Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023222186
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2023-12-28
Publication Date
2026-10-01
Estimated Expiration
2040-09-30

Smart Images

  • Figure 0007927680000037
    Figure 0007927680000037
  • Figure 0007927680000038
    Figure 0007927680000038
  • Figure 0007927680000039
    Figure 0007927680000039
Patent Text Reader

Abstract

To provide a method for decoding a coded video bit stream.SOLUTION: A decoding method comprises: obtaining, from a coded video bitstream, a first syntax element specifying whether a first layer uses inter-layer prediction and one or more second syntax elements related to one or more second layers, where each second syntax element specifies whether a second layer is a direct reference layer for the first layer, where at least one second syntax element of the one or more second syntax elements has a value specifying that the second layer is the direct reference layer for the first layer, in the case that the value of the first syntax element specifies that the first layer is allowed to use the inter-layer prediction; and performing the inter-layer prediction for a picture of the first layer by using a picture of the second layer related to the at least one second syntax element as a reference picture.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present application (disclosure) generally relate to the field of image processing, and more particularly to inter-layer prediction. CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims priority to U.S. Provisional Application No. 62 / 912,046, filed on October 7, 2019, the entire subject matter of which is incorporated herein by reference. [Background Art]

[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications, for example, broadcast digital television, video transmission over the Internet and mobile networks, real-time conversational applications such as video chat, video conferencing, DVD and Blu-ray® discs, video content collection and editing systems, and camcorders for security applications.

[0003] The amount of video data required to depict a relatively short video is substantial, which can result in difficulties if the data is streamed or otherwise transmitted over a communication network with limited bandwidth. Therefore, video data is generally compressed before being transmitted over modern telecommunications networks. If the video is stored on a storage device, the video size can also be a problem, as memory resources may be limited. Video compressors often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompressor that decodes the video data. Given limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice in image quality are desirable. [Overview of the Initiative]

[0004] Embodiments of this application provide apparatus and methods for encoding and decoding according to independent claims.

[0005] The aforementioned and other objectives are achieved by the subject matter of the independent claims. Further embodiments are evident from the dependent claims, specification, and drawings.

[0006] Specific embodiments are outlined in the appended independent claims, and other embodiments are outlined in the dependent claims.

[0007] In accordance with a first aspect, the present invention relates to a method for decoding an encoded video bitstream. The method is performed by a decoding device. The method includes the steps of: obtaining a first syntax element from the encoded video bitstream that specifies whether a first layer uses interlayer prediction; obtaining one or more second syntax elements from the encoded video bitstream that relate to one or more second layers, each second syntax element specifying whether the second layer is a direct reference layer of the first layer, wherein, if the value of the first syntax element specifies that the first layer is permitted to use interlayer prediction, at least one of the one or more second syntax elements has a value that specifies that the second layer is a direct reference layer of the first layer; and performing interlayer prediction on a picture of the first layer by using a picture of the second layer related to the at least one second syntax element as a reference picture.

[0008] Alternatively, the first syntax element specifies whether or not one or more second syntax elements related to one or more second layers are present in the encoded video bitstream. Furthermore, a first syntax element equal to 1 specifies that one or more second syntax elements related to one or more second layers are not present in the encoded video bitstream, or a first syntax element equal to 0 specifies that one or more second syntax elements related to one or more second layers are present in the encoded video bitstream. Here, a layer contains a sequence of coded pictures having the same layer index. Here, the layer index of one or more second layers is smaller than the layer index of the first layer. Here, second layers associated with different second syntax elements have different layer indices. Here, one or more second syntax elements have a one-to-one correspondence with one or more second layers. Here, a bitstream is a sequence of bits that make up one or more coded video sequences (CVS). Here, a coded video sequence (CVS) is a sequence of AU. Here, the coded layer video sequence (CLVS) is a sequence of PUs that have the same nuh_layer_id value. Here, an access unit (AU) is a set of PUs that belong to different layers and contain coded pictures associated with the output from the DPB at the same time. Here, a picture unit (PU) is a set of NAL units that are related to one another according to a specified classification rule, are consecutive in the decoding order, and contain exactly one coded picture. Here, the Interlayer Reference Picture (ILRP) is a picture in the same AU as the current picture, and its nuh_layer_id is smaller than the current picture's nuh_layer_id. Here, SPS is a syntax structure that contains syntax elements that apply to zero or more of the entire CLVS. Here, if layer A uses layer B as a reference layer, then layer B is a direct reference layer of layer A. If layer A uses layer B as a reference layer, and layer B uses layer C as a reference layer, but layer A does not use layer C as a reference layer, then layer C is not a direct reference layer of layer A.

[0009] In a possible implementation of the method according to such first embodiment, the first syntax element equal to 1 specifies that the first layer does not use interlayer prediction, or the first syntax element equal to 0 specifies that the first layer is permitted to use interlayer prediction.

[0010] In any prior implementation of the first embodiment or in any possible implementation of the method according to such first embodiment, a second syntax element equal to 0 specifies that the second layer associated with the second syntax element is not a direct reference layer of the first layer, or a second syntax element equal to 1 specifies that the second layer associated with the second syntax element is a direct reference layer of the first layer.

[0011] In any prior implementation of the first embodiment or a possible implementation of the method according to such first embodiment, the step of obtaining one or more second syntax elements is performed if the value of the first syntax element specifies that the first layer is permitted to use interlayer prediction.

[0012] In any prior implementation of the first embodiment or a possible implementation of the method according to such first embodiment, the method further includes the step of performing a prediction on the picture of the first layer without using the picture of the layer associated with the at least one second syntax element as a reference picture, where the value of the first syntax element specifies that the first layer does not use interlayer prediction.

[0013] According to a second aspect, the present invention provides a method for encoding an encoded video bitstream. The method is performed by an encoder. The method includes the steps of: determining whether at least one second layer is a direct reference layer of a first layer; and encoding a syntax element into the encoded video bitstream, wherein the syntax element specifies whether the first layer uses interlayer prediction, where if none of the at least one second layer is a direct reference layer of the first layer, the value of the syntax element specifies that the first layer does not use interlayer prediction.

[0014] The step of determining whether at least one second layer is a direct reference layer of the first layer includes the steps of determining that the second layer is a direct reference layer of the first layer, based on the determination that the first rate distortion cost is less than or equal to the second rate distortion cost, and determining that the second layer is not a direct reference layer of the first layer, based on the determination that the first rate distortion cost is greater than or equal to the second rate distortion cost, where the first rate distortion cost is the cost of using the second layer as a direct reference layer of the first layer, and the second rate distortion cost is the cost of not using the second layer as a direct reference layer of the first layer.

[0015] In a possible implementation of the method according to such second aspect, the value of the syntax element specifies that the first layer is permitted to use interlayer prediction if the at least one second layer is a direct reference layer of the first layer.

[0016] According to a third aspect, the present invention relates to an apparatus for decoding an encoded video bitstream. The apparatus includes an acquisition unit and a prediction unit. The acquisition unit is configured to acquire a first syntax element from the encoded video bitstream that specifies whether a first layer uses interlayer prediction. The acquisition unit is further configured to acquire one or more second syntax elements associated with one or more second layers, each second syntax element specifying whether the second layer is a direct reference layer of the first layer. Here, if the value of the first syntax element specifies that the first layer is allowed to use interlayer prediction, at least one of the one or more second syntax elements has a value that specifies the second layer is a direct reference layer of the first layer. The prediction unit is configured to perform interlayer prediction on a picture of the first layer by using a picture of the second layer associated with at least one second syntax element as a reference picture.

[0017] In a possible implementation of the method according to such a third aspect, the first syntax element equal to 1 specifies that the first layer does not use interlayer prediction, or the first syntax element equal to 0 specifies that the first layer is permitted to use interlayer prediction.

[0018] In any prior implementation of the third aspect or in any possible implementation of the method according to such third aspect, a second syntax element equal to 0 specifies that the second layer associated with the second syntax element is not a direct reference layer of the first layer, or a second syntax element equal to 1 specifies that the second layer associated with the second syntax element is a direct reference layer of the first layer.

[0019] In a possible implementation of the method according to any preceding implementation of the third aspect or such third aspect, the prediction unit configured to obtain one or more second syntax elements is performed when the value of said first syntax element specifies that said first layer is permitted to use inter-layer prediction.

[0020] According to a fourth aspect, the present invention relates to an apparatus for encoding a coded video bitstream. The apparatus comprises a determination unit and an encoding unit. The determination unit is configured to determine whether at least one second layer is a direct reference layer of a first layer. The encoding unit is configured to encode a syntax element into the coded video bitstream, wherein the syntax element specifies whether the first layer uses inter-layer prediction, and wherein if none of the at least one second layer is a direct reference layer of the first layer, the value of the syntax element specifies that the first layer does not use inter-layer prediction.

[0021] In a possible implementation of the method according to said fourth aspect, when at least one second layer is a direct reference layer of the first layer, the value of said syntax element specifies that said first layer is permitted to use inter-layer prediction.

[0022] The method of the first aspect of the present invention may be implemented by the apparatus of the third aspect of the present invention. Further features and implementations of the method according to the third aspect of the present invention correspond to features and implementations of the apparatus according to the first aspect of the present invention.

[0023] The method of the second aspect of the present invention may be implemented by the apparatus of the fourth aspect of the present invention. Further features and implementations of the method according to the fourth aspect of the present invention correspond to features and implementations of the apparatus according to the second aspect of the present invention.

[0024] The method according to the second aspect can be extended to an implementation form corresponding to the implementation form of the first apparatus according to the first aspect. Therefore, the implementation form of the method includes the features of the corresponding implementation form of the first apparatus.

[0025] The advantages of the method according to the second embodiment are the same as the advantages of the corresponding implementation of the first apparatus according to the first embodiment.

[0026] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and memory. The memory stores instructions causing the processor to perform the method according to the first aspect.

[0027] According to a sixth aspect, the present invention relates to an apparatus for encoding a video stream, comprising a processor and memory. The memory stores instructions causing the processor to perform the method according to a second aspect.

[0028] According to the seventh aspect, a computer-readable storage medium is proposed that stores instructions configured to encode video data in one or more processors when executed. The instructions cause one or more processors to perform the method according to the first or second aspect, or any possible embodiment of the first or second aspect.

[0029] In accordance with the eighth aspect, the present invention relates to a computer program that, when executed on a computer, includes program code for performing a method according to the first or second aspect, or any possible embodiment of the first or second aspect.

[0030] In accordance with the ninth aspect, the present invention relates to a non-temporary storage medium comprising an encoded bitstream decoded by a device. The bitstream is generated by dividing a frame of a video signal or image signal into a plurality of blocks and comprises a plurality of syntax elements, wherein the plurality of syntax elements comprises a first syntax element specifying whether a first layer uses interlayer prediction, and one or more second syntax elements relating to one or more second layers, each second syntax element specifying whether the second layer is a direct reference layer of the first layer. Where the value of the first syntax element specifies that the first layer is permitted to use interlayer prediction, at least one of the one or more second syntax elements has a value specifying that the second layer is a direct reference layer of the first layer.

[0031] Details of one or more embodiments are revealed in the accompanying drawings and the following description. Other features, purposes, and advantages will be apparent from the specification, drawings, and claims.

[0032] Furthermore, the following embodiments are provided.

[0033] In one embodiment, a method for decoding an encoded video bitstream is provided. This method includes the following: A step of parsing a first syntax element that specifies whether the layer having index i uses interlayer prediction, where i is an integer and i is greater than 0. If the first condition is met, the step is to parse a second syntax element that specifies whether the layer having index j is a direct reference layer of the layer having index i, wherein the first condition includes specifying that the layer having index i can use interlayer prediction, j is equal to i-1, and any one of the layers having an index less than j is not a direct reference layer of the layer having index i. A step of predicting the picture of the layer having index i based on the value of the second syntax element.

[0034] In one embodiment, if the first syntax element specifies that the layer having index i may use interlayer prediction, then the sum of vps_direct_direct_direct_dinercy_flag[i][k] is greater than 0, where k is any integer between 0 and i-1. Here, a vps_direct_direct_direct_depency_flag equal to 1 specifies that the layer having index k is a direct reference layer of the layer having index I, and a vps_direct_direct_direct_depency_flag equal to 0 specifies that the layer having index k is not a direct reference layer of the layer having index i.

[0035] In one embodiment, if the first syntax element specifies that a layer having index i can use interlayer prediction, then at least one value of vps_direct_direct_direct_dinercy_flag[i][k] is equal to 1, where k is an integer and k is in the range from 0 to i-1. Here, vps_direct_direct_direct_depency_flag equal to 1 specifies that the layer having index k is a direct reference layer of the layer having index I, and vps_direct_direct_direct_depency_flag equal to 0 specifies that the layer having index k is not a direct reference layer of the layer having index i.

[0036] In one embodiment, the picture of a layer having index i includes a picture within the layer having index i, or a picture associated with the layer having index i.

[0037] In one embodiment, a method for decoding an encoded video bitstream is provided. This method includes the following: A step of parsing a first syntax element that specifies whether the layer having index i uses interlayer prediction, where i is an integer and i is greater than 0. The step of predicting the picture of the layer having index i by using the image of the layer having index j as the direct reference layer of the layer having index i, provided that the condition is met, where j is an integer and j is equal to i-1, and the condition includes a syntax element specifying that the layer having index i can use interlayer prediction.

[0038] In one embodiment, the picture of a layer having index i includes a picture within the layer having index i, or a picture associated with the layer having index i.

[0039] In one embodiment, a method for decoding an encoded video bitstream is provided. This method includes the following: A step of parsing syntax elements that specify whether at least one long-term reference picture (LTRP) is used for interpretation of coded pictures in a coded video sequence (CVS). Here, each picture of at least one LTRP is marked as “used for long-term reference” but is not an interlayer reference picture (ILRP). A step to predict one or more coded pictures in CVS based on the values ​​of syntax elements.

[0040] In one embodiment, a method for decoding an encoded video bitstream is provided. This method includes the following: This step determines whether a condition is met, where the entire condition includes the layer index of the current layer being greater than a preset value. If the conditions are met, the step is to parse a first syntax element that specifies whether at least one interlayer reference picture (ILRP) is used for interpretation of any coded picture in a coded video sequence (CVS). A step of predicting one or more coded pictures in CVS based on the value of the first syntax element.

[0041] In one embodiment, the preset value is 0.

[0042] In one embodiment, the condition further includes that the second syntax element (e.g., sps_video_parameter_set_id) is greater than 0.

[0043] In one embodiment, a method for decoding an encoded video bitstream is provided. This method includes the following: A step to determine whether the conditions are met, where all conditions include the layer index of the current layer being greater than a preset value and the current entry in the reference picture list structure being an ILRP entry. If the conditions are met, the step of parsing the syntax element that specifies an index to the list of direct dependent layers of the current layer. A step of predicting one or more coded pictures in CVS based on the reference picture list structure of the current entry, from which the current ILRP is obtained using an index against a list of direct dependent layers.

[0044] In one embodiment, the preset value is 1.

[0045] In one embodiment, an encoder (20) is provided, comprising a processing circuit for performing the method described in any one of claims 1 to 12.

[0046] In one embodiment, a decoder (30) is provided, comprising a processing circuit for performing the method described in any one of claims 1 to 12.

[0047] In one embodiment, a computer program product is provided which, when executed on a computer or processor, includes program code for performing the method described in any one of the preceding claims.

[0048] In one embodiment, a decoder is provided which includes the following: One or more processors, and A non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein the program, when executed by the processor, configures the decoder to perform the method according to any one of the preceding claims.

[0049] In one embodiment, an encoder is provided which includes the following: One or more processors, and A non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein the program, when executed by the processor, configures the encoder to perform the method according to any one of the preceding claims.

[0050] In one embodiment, a non-temporary computer-readable medium is provided that carries program code. When the program code is executed by a computer device, it causes the computer device to perform the method described in any one of the preceding claims. [Brief explanation of the drawing]

[0051] The following embodiments of the present invention are described in more detail with reference to the accompanying figures and drawings. [Figure 1A] Figure 1A is a block diagram showing one example of a video coding system configured to carry out an embodiment of the present invention. [Figure 1B] Figure 1B is a block diagram showing another example of a video coding system configured to carry out embodiments of the present invention. [Figure 2] Figure 2 is a block diagram showing one example of a video encoder configured to carry out an embodiment of the present invention. [Figure 3]Figure 3 is a block diagram showing one exemplary structure relating to a video decoder configured to carry out an embodiment of the present invention. [Figure 4] Figure 4 is a block diagram showing one example of an encoding or decoding device. [Figure 5] Figure 5 is a block diagram showing another example relating to an encoding or decoding device. [Figure 6] Figure 6 is a block diagram illustrating scalable coding with two layers. [Figure 7] Figure 7 is a block diagram showing an exemplary configuration of the content supply system 3100 that realizes a content distribution service. [Figure 8] Figure 8 is a block diagram showing the configuration of one example of a terminal device. [Figure 9] Figure 9 shows a flowchart relating to a decoding method according to one embodiment. [Figure 10] Figure 10 shows a flowchart relating to an encoding method according to one embodiment. [Figure 11] Figure 11 is a schematic diagram of an encoder according to one embodiment. [Figure 12] Figure 12 is a schematic diagram of a decoder according to one embodiment.

[0052] The following identical reference numerals refer to the same, or at least functionally equivalent, features unless explicitly specified otherwise. [Modes for carrying out the invention]

[0053] The following description refers to the accompanying drawings, which form part of the present disclosure and, as illustrative examples, show specific embodiments of the present invention or specific ways in which embodiments of the present invention may be used. Embodiments of the present invention may be used in other ways and are understood to include structural or logical modifications not shown in the drawings. The following detailed description should therefore not be construed as restrictive, and the scope of the present invention is defined by the accompanying claims.

[0054] For example, disclosures relating to a described method are also true for a corresponding device or system configured to perform the method, and vice versa, so that the steps to be taken are understood. For example, if one or more specific method steps are described, the corresponding device to perform the described one or more method steps may include one or more units (e.g., functional units) (e.g., one unit that performs one or more steps, or multiple units that perform each of the multiple steps). On the other hand, if a particular device is described based on one or more units, e.g., functional units, the corresponding method may include one step to perform the functionality of one or more units (e.g., one step that performs the functionality of one or more units, or multiple steps that perform the functionality of one or more of the multiple units), even if such one or more units are not explicitly described or shown in the drawings. Furthermore, unless otherwise noted, the various exemplary embodiments and / or features of the aspects described herein are understood to be achieved by combining them with each other.

[0055] Video coding typically refers to processing a series of images that form a video or video sequence. Instead of the term “picture”, the terms “frame” or “image” may be used synonymously in the field of video coding. Video coding (or coding in general) comprises two parts: video coding and video decoding. Video coding is performed at the source and typically involves processing the original video image to reduce the amount of data required to represent the video image (for more efficient storage and / or transmission). Video decoding is performed at the destination and typically involves reverse processing compared to the encoder to reconstruct the video image. Embodiments referring to “coding” of video image (or picture in general) are understood to relate to “encoding” or “decoding” of the video image or each video sequence. The combination of encoding and decoding is also referred to as CODEC (Coding and Decoding).

[0056] In lossless video coding, the original video footage can be reconstructed. That is, the reconstructed video footage has the same quality as the original video footage (assuming there is no transmission loss or other data loss during storage or transmission). In lossy video coding, further compression is performed, for example by quantization, to reduce the amount of data representing the video footage, and it cannot be fully reconstructed by the decoder. That is, the quality of the reconstructed video footage is lower or worse than the quality of the original video footage.

[0057] Several video coding standards belong to the group of “lossy hybrid video codecs” (i.e., 2D transform coding to apply spatial and temporal prediction in the sample domain and quantization in the transform domain). Each picture in a video sequence is typically divided into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in an encoder, video is typically processed, i.e., encoded, at the block (video block) level. This is done, for example, by using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to generate prediction blocks, by subtracting prediction blocks from current blocks (blocks currently being processed / to be processed) to obtain residual blocks, by transforming the residual blocks to reduce the amount of data to be transmitted, and by quantizing the residual blocks in the transform domain. In a decoder, on the other hand, the reverse processing is applied to the encoded or compressed blocks compared to the encoder, so that the current blocks are reconstructed for representation. Furthermore, the encoder duplicates the decoder processing loop so that both generate identical predictions (e.g., intra-predictions and inter-predictions) and / or reconstruct the subsequent blocks for processing, i.e., coding.

[0058] In the following embodiment relating to the video coding system 10, the video encoder 20 and video decoder 30 will be described based on Figures 1 to 3.

[0059] Figure 1A is a schematic block diagram showing an exemplary coding system 10, for example, a video coding system 10 (or short coding system 10) on which the technology of this application can be utilized. The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video coding system 10 represent an example of a device that may be configured to perform the technology according to the various embodiments described herein.

[0060] As shown in Figure 1A, the coding system 10 includes a source device 12 configured to provide encoded video data 12, which it then provides to a destination device 14 for decoding, for example, the encoded video data 13.

[0061] The source device 12 includes an encoder 20 and, optionally, a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a video preprocessor 18, and a communication interface or communication unit 22.

[0062] The picture source 16 may include, or may be any kind of video capture device, e.g., a camera for capturing real-world images, and / or any kind of video generation device, e.g., a computer graphics processor for generating computer-animated images, or any other kind of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The picture source may be any kind of memory or storage device for storing any of the above images.

[0063] In distinction from the processing performed by the preprocessor 18 and the preprocessing unit 18, the video or video data 17 is also referred to as raw video or raw video data 17.

[0064] The preprocessor 18 is configured to receive (raw) video data 17, perform preprocessing on the video data 17, and obtain preprocessed video data 19 or preprocessed video data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 18 may be an optional component.

[0065] The video encoder 20 is configured to receive pre-processed video data 19 and provide encoded video data 21 (for example, based on Figure 2, further details are described below). The communication interface 22 of the source device 12 may be configured to receive the encoded video data 21 and transmit the encoded video data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14, or any other device, for storage or direct reconstruction.

[0066] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and may also include a communication interface or communication unit 28, a post-processor 32 (or a post-processing unit 32), and a display device 34.

[0067] The communication interface 28 of the destination device 14 is configured to receive encoded video data 21 (or any further processed version thereof) for example directly from the source device 12 or from any other source, such as a storage device, such as an encoded video data storage device, and to provide the encoded video data 21 to the decoder 30.

[0068] Communication interfaces 22 and 28 provide a direct communication link between the source device 12 and the destination device 14, for example, a direct wired or wireless connection. The system may be configured to transmit or receive encoded video data 21 or encoded data 13 via, or via any type of network, such as a wired or wireless network, or any combination thereof, or any type of private and public network, or any combination thereof.

[0069] The communication interface 22 may be configured to, for example, package the encoded video data 21 into an appropriate format, such as a packet, and / or process the encoded video data using any type of transmission encoding or processing for transmission over a communication link or communication network.

[0070] A communication interface 28, which forms the counterpart to communication interface 22, can be configured, for example, to receive transmission data and process the transmission data using any kind of corresponding transmission decoding, processing, and / or depackaging to obtain encoded video data 21.

[0071] Communication interfaces 22 and 28, both of which may be configured as one-way and / or two-way communication interfaces, as indicated by the arrows on the communication channel 13 in Figure 1A pointing from the source device 12 to the destination device 14. They may also be configured to send and receive messages, for example, to acknowledge and exchange any other information relating to the communication link and / or data transmission, such as encoded video data transmission.

[0072] The decoder 30 is configured to receive encoded video data 21 and provide decoded video data 31 or decoded picture 31 (further details are described below, for example, based on Figure 3 or Figure 5).

[0073] The post-processor 32 of the destination device 14 is configured to post-process the decoded video data 31 (also called reconstructed video data), for example, the decoded picture 31, and obtains post-processed video data 33, for example, the post-processed video 33. Post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., conversion from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare, for example, the decoded video data 31 for display by the display device 34.

[0074] The display device 34 of the destination device 14 is configured to receive post-processed video data 33 for displaying an image to, for example, a user or viewer. The display device 34 may be, or include, any type of display for representing the reconstructed image, such as an integrated or external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light-emitting diode display (OLED), a plasma display, a projector, a micro-LED display, liquid crystal on silicon (LCoS), a digital optical processor (DLP), or any other type of display.

[0075] Figure 1A depicts the source device 12 and the destination device 14 as separate devices, but the embodiment of the device may also include both or both functionalities, the source device 12 or its corresponding functionality, and the destination device 14 or its corresponding functionality. In such embodiments, the source device 12 or its corresponding functionality and the destination device 14 or its corresponding functionality can be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0076] As will be apparent to those skilled in the art based on this description, the functionality or existence and (exact) division of different units within the source device 12 and / or destination device 14, as shown in Figure 1A, may vary depending on the actual device and application.

[0077] The encoder 20 (e.g., video encoder 20), the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30 may be implemented via a processing circuit as shown in Figure 1B. This could be one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 may be implemented via a processing circuit 46 to embody various modules, as described with respect to the encoder 20 in Figure 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via a processing circuit 46 to embody various modules, as described with respect to the decoder 30 in Figure 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations, as will be described later. As shown in Figure 5, when the technology is partially implemented in software, the device can store instructions for the software on a suitable non-temporary computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Either the video encoder 20 or the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, for example, as shown in Figure 1B.

[0078] The source device 12 and destination device 14 can include any wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (such as a content service server or content distribution server), broadcast receiver device, broadcast transmitter device, and may or may not use an operating system. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication. Thus, the source device 12 and destination device 14 can be wireless communication devices.

[0079] In some cases, the video coding system 10 shown in Figure 1A is merely one example, and the technology of this application may be applied to video coding configurations (e.g., video encoding or video decoding) that do not necessarily involve data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory and streamed over a network, etc. A video coding device can encode data and store it in memory, and / or a video decoding device can retrieve data from memory and decode it. In some examples, encoding and decoding do not communicate with each other but are simply performed by devices that encode data into memory and / or retrieve data from memory and decode it.

[0080] For the sake of explanation, embodiments of the present invention are described herein by reference, for example, to High-Efficiency Video Coding (HEVC), or to reference software for Versatile Video Coding (VVC), the Joint Collaboration on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG), and next-generation video coding standards developed by the ISO / IEC Motion Picture Experts Group (MPEG). A person skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.

[0081] Encoder and encoding method Figure 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of Figure 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transformation unit 206, a quantization unit 208, an inverse quantization unit 210, and an inverse transformation unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded video buffer 230, a mode selection unit 260, an entropy encoding unit 270, and an output unit 272 (or output interface 272). The mode selection unit 260 may include an inter-prediction unit 244, an intra-prediction unit 254, and a partitioning unit 262. The inter-prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in Figure 2 is also referred to as a hybrid video encoder or video encoder, depending on the hybrid video codec.

[0082] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 are referred to as forming the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded video buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 may be referred to as forming the backward signal path of the video encoder 20. Here, the backward signal path of the video encoder 20 corresponds to the signal path of the decoder (see video decoder 30 in Figure 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded video buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 are also referred to as forming the “built-in decoder” of the video encoder 20.

[0083] Picture and picture splitting (picture and block) The encoder 20 may be configured to receive, for example, a picture 7 (or video data 17), which is a video relating to a series of pictures forming a video or video sequence, via input 201. The received picture or video data may also be a pre-processed picture 19 (or pre-processed video data 19). For brevity, the following description will refer to Figure 17. Picture 17 may also be referred to as the current picture or the picture being coded (particularly in video coding to distinguish the current picture from other pictures, for example, previously encoded and / or decoded pictures relating to the same video sequence, i.e., the video sequence which also includes the current picture).

[0084] A (digital) image is, or can be considered as, a two-dimensional array or matrix of samples with intensity values. A sample in an array may also be referred to as a pixel (a shorter form of a picture element) or pel. The number of samples in the horizontal and vertical directions (or axes) of an array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used; that is, a picture may represent or contain three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format, or color space, for example, YCbCr, which includes a luminance component represented by Y (sometimes L is also used instead) and two chrominance components represented by Cb and Cr. The luminance (or, short, luma) component Y represents brightness or gray level intensity (e.g., in a grayscale image), while the two chrominance (or, short, chroma) components Cb and Cr represent chromaticity or color information components. Therefore, a picture in the YCbCr format contains a luminance sample array of luminance sample values ​​(Y) and two chroma sample arrays of chrominance values ​​(Cb and Cr). A picture in the RGB format can be converted to or transformed into the YCbC format, and vice versa. This method is also known as color conversion or transformation. If the picture is monochrome, it may contain only a luminance sample array. Thus, an image can be, for example, an array of luminance samples in a monochrome format, or an array of luminance samples, as well as two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0085] Embodiments of the video encoder 20 may include a video partitioning unit (not shown in Figure 2) configured to divide a picture 17 into multiple (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), coding tree blocks (CTB), or coding tree units (CTU) (H.265 / HEVC and VVC). The video partitioning unit may be configured to use the same block size for all pictures in a video sequence and for corresponding grids that define the block size, or to change the block size between pictures, subsets, or groups of pictures, and to divide each picture into a corresponding block.

[0086] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, for example, one, several, or all of the blocks that make up picture 17. Picture blocks 203 may also be referred to as currently picture blocks, or picture blocks to be coded.

[0087] Similar to picture 17, picture block 203, though smaller in dimensions than picture 17, is again a two-dimensional array or matrix of samples with luminance values ​​(sample values), or can be considered as such. In other words, block 203 may contain, for example, one sample array (e.g., a lumar array in the case of monochrome image 17, or a lumar or chromar array in the case of color image), or three sample arrays (e.g., a lumar and two chromar arrays in the case of color image 17), or any other number and / or type of arrays depending on the applied color format. The number of samples in the horizontal and vertical (or axis) directions of block 203 defines the size of block 203. Thus, a block may be, for example, an M×N (M columns N rows) array of samples, or an M×N array of conversion coefficients.

[0088] As shown in Figure 2, an embodiment of the video encoder 20 may be configured to encode blocks of picture 17 block by block. For example, encoding and prediction may be performed for each block 203.

[0089] As shown in Figure 2, an embodiment of the video encoder 20 may further be configured to partition and / or encode a picture by using slices (also referred to as video slices). Here, a picture can be partitioned or encoded into one or more slices (typically non-overlapping), and each slice may contain one or more blocks (e.g., CTUs) or groups of one or more blocks (e.g., tiles (H.265 / HEVC and VC) or bricks (VVC)).

[0090] As shown in Figure 2, embodiments of the video encoder 20 may further be configured to partition and / or encode a picture by using slice / tile groups (also referred to as video tile groups) and / or tile groups (also referred to as video tiles). Here, a picture can be partitioned or encoded into one or more slice / tile groups (typically non-overlapping), and each slice / tile group may contain, for example, one or more blocks (e.g., CTUs) or one or more tiles. Here, each tile may be, for example, rectangular in shape and may contain one or more blocks (e.g., CTUs), for example, complete or divided blocks.

[0091] Residual Calculation The residual calculation unit 204 can be configured to calculate residual blocks 205 (also referred to as residual block 205) based on picture blocks 203 and prediction blocks 265 (further details about prediction blocks 265 are provided below). For example, residual blocks 205 in the sample region are obtained sample by sample (pixel by pixel) by subtracting the sample values ​​of prediction blocks 265 from the sample values ​​of picture blocks 203.

[0092] Transform The transformation processing unit 206 may be configured to apply a transformation, such as a discrete cosine transform (DCT) or discrete sine transform (DST), to the sample values ​​of the residual block 205 in order to obtain transformation coefficients 207 in the transformation domain. The transformation coefficients 207 are also referred to as transformation residual coefficients and can represent the residual block 205 in the transformation domain.

[0093] The conversion processing unit 206 may be configured to apply an integer approximation of the DCT / DST, such as the conversion specified for H.265 / HEVC. Compared to the orthogonal DCT conversion, such an integer approximation is typically scaled by a predetermined factor. An additional scaling factor is applied as part of the conversion process to preserve the norm of the residual blocks processed by the forward and inverse conversions. The scaling factor is typically selected based on predetermined constraints such as a scaling factor that is a power of two for the shift operation, the bit depth of the conversion factor, and a trade-off between precision and implementation cost. A specific scaling factor is specified, for example, for the inverse conversion by the inverse conversion processing unit 212 (and for the corresponding inverse conversion by the inverse conversion processing unit 312 in the video decoder 30). Then, a corresponding scaling factor for the forward conversion may be specified accordingly, for example, by the conversion processing unit 206 in the encoder 20.

[0094] Embodiments of the video encoder 20 (each with a conversion processing unit 206) may be configured to output, for example, a conversion parameter, such as the type of conversion, either directly or encoded or compressed via, for example, the entropy encoding unit 270. As a result, for example, the video decoder 30 can receive and use the conversion parameter for decoding.

[0095] quantization The quantization unit 208 may be configured to quantize the transformation coefficients 207 to obtain quantization coefficients 209, for example, by applying scalar quantization or vector quantization. The quantization coefficients 209 are also referred to as quantization transformation coefficients 209 or quantization residual coefficients 209.

[0096] The quantization process can reduce the bit depth associated with some or all of the 207 conversion coefficients. For example, n-bit conversion coefficients may be rounded down to m-bit conversion coefficients during quantization, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scaling can be applied to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. Applicable quantization step sizes can be indicated by the quantization parameter (QP), which can be, for example, an index to a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to finer quantization (smaller quantization step size), and a large quantization parameter may correspond to coarser quantization (larger quantization step size), or vice versa. Quantization may involve division by the quantization step size, and the corresponding and / or inverse dequantization may involve multiplication by the quantization step size, for example by the inverse quantization unit 210. Embodiments following some standards, e.g., HEVC, may be configured to use a quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of the equations, including division. Additional scaling factors may be introduced for quantization and dequantization to recover the norm of the residual block, which may be modified for the scaling used in the fixed-point approximation of the equations for the quantization step size and quantization parameter. In one embodiment, the scaling of the inverse transform and dequantization can be combined. Alternatively, a customized quantization table may be used and signaled from encoder to decoder, for example, in a bitstream.Quantization is an irreversible operation, and here, the loss increases with increasing quantization step size.

[0097] Embodiments of the video encoder 20 (each with a quantization unit 208) may be configured to output quantization parameters (QP) directly or encoded, for example, via an entropy encoding unit 270. As a result, a video decoder 30, for example, can receive and apply the quantization parameters for decoding.

[0098] Inverse Quantization The inverse quantization unit 210 is configured to apply the inverse quantization of the quantization unit 208 to the quantization coefficients by applying the reciprocal of the quantization scheme applied by the quantization unit 208, for example, based on or using the same quantization step size as the quantization unit 208, in order to obtain the dequantized coefficients 211. The dequantized coefficients 211 are also referred to as the dequantized residual coefficients 211 and correspond to the transformation coefficients 207, although they are typically not identical to the transformation coefficients due to losses due to quantization.

[0099] Inverse Transform The inverse transform processing unit 212 is configured to apply the inverse transform of the transform applied by the transform processing unit 206, such as the inverse discrete cosine transform (DCT), the inverse discrete sine transform (DST), or other inverse transforms, to obtain a reconstructed residual block 213 (or the corresponding dequantization coefficient 213) in the sample region. The reconstructed residual block 213 is also referred to as the transform block 213.

[0100] Reconstruction The reconstruction unit 214 (e.g., an adder or summator 214) is configured to add the transformed block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample value of the reconstructed residual block 213 and the sample value of the prediction block 265—sample by sample—to obtain the reconstructed block 215 in the sample region.

[0101] filtering The loop filter unit 220 (or, for short, “loop filter” 220) is configured to filter the reconstructed block 215 to obtain filtered block 221, or, more generally, to filter the reconstructed sample to obtain filtered sample values. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one embodiment, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be deblocking, SAO, and then ALF. In another example, a process called chroma mapping with chroma scaling (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. In another example, the deblocking process may also be applied to internal subblock edges, such as affine subblock edges, ATMVP subblock edges, subblock transformation (SBT) edges, and intra-subpartition (ISP) edges. The loop filter unit 220 is shown as an in-loop filter in Figure 2, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconfigured block 221.

[0102] Embodiments of the video encoder 20 (each with a loop filter unit 220) may be configured to output loop filter parameters (such as SAO filter parameters, ALF filter parameters, or LMCS parameters) either directly or encoded via the entropy encoding unit 270. As a result, the decoder 30 may receive and apply the same loop filter parameters, or the respective loop filters.

[0103] Decoded picture buffer The decoded picture buffer (DPB) 230 may be a memory for storing reference pictures for encoding video data by the video encoder 20, or it may be within general reference video data. The DPB 230 may be formed by any of various memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (ADRAM), magnetoresistive RAM (MRAM), resistive random access RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks (e.g., previously reconfigured and filtered blocks 221) relating to the same current picture or a different picture, e.g., a previously reconfigured picture. Then, for example, for interpretation, a fully reconfigured, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially reconfigured current picture (and corresponding reference blocks and samples) may be provided. The decoded picture buffer (DPB) 230 may also be configured to store one or more reconstructed unfiltered blocks 215, or, for example, if the reconstructed blocks 215 are not filtered by the loop filter unit 220, generally unfiltered reconstructed samples, or any other further processed versions of the reconstructed blocks or samples.

[0104] Mode selection (partitioning and prediction) The mode selection unit 260 includes a partitioning unit 262, an inter-prediction unit 244, and an intra-prediction unit 254, and is configured to receive or acquire original video data, e.g., the original block 203 (the current block 203 of the current picture 17), and reconstructed video data, e.g., filtered and / or unfiltered, reconstructed samples or blocks of the same (current) picture, and / or from one or more previously decoded pictures, e.g., from a decoded picture buffer 230 or other buffer (e.g., a line buffer, not shown). The reconstructed video data is used as reference picture data for predictions, e.g., inter-prediction or intra-prediction, to acquire prediction block 265 or predictor 265.

[0105] The mode selection unit 260 may be configured to determine or select partitioning for the current block prediction mode (including no partitioning) and prediction mode (e.g., intra or inter prediction mode), and to generate corresponding prediction blocks 265 used for calculating residual blocks 205 and for reconstructing reconstructed blocks 215.

[0106] Embodiments of the mode selection unit 260 may be configured to select partitioning and prediction modes (for example, from those supported or available by the mode selection unit 260) that provide the best fit, i.e., in other words, the smallest residuals (where smallest residuals mean better compression for transmission or storage), or the smallest signaling overhead (where smallest signaling overhead means better compression for transmission or storage), or consider or balance both. The mode selection unit 260 may be configured to determine the partitioning and prediction modes based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the smallest rate distortion. In this context, terms such as “best”, “minimum”, and “optimum”, etc., do not necessarily refer to the overall “best”, “minimum”, and “optimum”, etc., but may also refer to achieving termination or selection criteria that reduce complexity and processing time, which may lead to values ​​above or below a threshold, or a “sub-optimum”,

[0107] In other words, the partitioning unit 262 may be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs). The CTU 203 is then further partitioned into smaller block partitions or subblocks (which form blocks again) using, for example, quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, iteratively, and predictions can be performed, for example, on each of the block partitions or subblocks. Here, mode selection involves selecting the tree structure of the partitioned block 203, and prediction modes are applied to each of the block partitions or subblocks.

[0108] The following describes in more detail the partitioning (e.g., by the partitioning unit 260) and prediction processing (by the inter-prediction unit 244 and the intra-prediction unit 254) performed by an exemplary video encoder 20.

[0109] Partitioning The partitioning unit 262 may be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs). The partitioning unit 262 can then partition (or split) the coding tree units (CTUs) 203 into smaller partitions, for example, smaller blocks of a square or rectangular size. For a picture with three sample arrays, the CTU consists of an N×N block of luma samples and two corresponding blocks of chroma samples. The maximum allowable size of a luma block in a CTU is specified as 128×128 in the Versatile Video Coding (VVC) under development, but in the future it may be specified as a value other than 128×128, for example, 256×256. The CTUs of a picture may be clustered / grouped as slice / tile groups, tiles, or bricks. A tile covers a rectangular area of ​​a picture, and a tile can be divided into one or more bricks. A brick consists of multiple columns of CTUs within a tile. A tile that has not been partitioned into multiple bricks may be referred to as a brick. However, a brick is a true subset of a tile and is not referred to as a tile. There are two modes for tile groups supported by VVC: raster-scan slice / tile group mode and rectangular slice mode. In raster-scan tile group mode, a slice / tile group contains a sequence of tiles in the tile raster scan of a picture. In rectangular slice mode, a slice contains many bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are in the order of the slice's brick raster scan. These smaller blocks (also referred to as subblocks) can be further partitioned into even smaller partitions.This is also referred to as tree partitioning or hierarchical tree partitioning, where the root block, e.g., root tree level 0 (hierarchical level 0, depth 0), can be recursively partitioned. For example, it can be partitioned into two or more blocks of nodes at the next lower tree level, e.g., tree level 1 (hierarchical level 1, depth 1), where these blocks can be further partitioned into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), etc., until the partitioning is complete. Termination occurs when termination criteria are met, e.g., reaching the maximum tree depth or minimum block size. Blocks that are not further partitioned are also referred to as leaf blocks or leaf nodes of the tree. A tree using partitioning into two partitions is referred to as a binary tree (BT), a tree using partitioning into three partitions is referred to as a ternary tree (TT), and a tree using partitioning into four partitions is referred to as a quad tree (QT).

[0110] For example, a coding tree unit (CTU) may be, or include, a CTB for a luminous sample, two corresponding CTBs for a chroma sample of a picture having three sample arrays, or a CTB for a sample of a monochrome image or picture coded using three separate color planes and syntax structures used to code the sample. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N, such that the division of components into the CTB is a partition division. A coding unit (CU) may be, or include, a coding block for a luminous sample, two corresponding coding blocks for a chroma sample of a picture having three sample arrays, or a coding block for a sample of a monochrome image or picture coded using three separate color planes and syntax structures used to code the sample. Correspondingly, a coding block may be an M×N block of samples for some value of M and N, such that the division of the CTB into the coding block is a partition division.

[0111] In an embodiment, for example, according to HEVC, a coding tree unit (CTU) may be split into CUs by using a quad-tree structure, which is shown as the coding tree. The decision of whether or not to encode the picture region using interpicture (temporal) or intrapicture (spatial) prediction is made at the leaf CU level. Each leaf CU may be further split into one, two, or four PUs, depending on the PU splitting type. Within a single PU, the same prediction process is applied, and the relevant information is sent to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU splitting type, the leaf CU may be partitioned into transform units (TUs) according to another quad-tree structure similar to the coding tree for the CU.

[0112] In embodiments, for example, according to the latest video coding standard currently under development, referred to as Versatile Video Coding (VVC), a combined quad tree is nested in a multi-type tree using binary and terminally split segmentation structures used to partition coding tree units. In the coding tree structure within a coding tree unit, the CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quad tree. Then, the quaternary tree leaf nodes can be further partitioned by a multi-type tree structure. There are four split types in the multi-type tree structure: vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical terminally split (SPLIT_TT_VER), and horizontal terminally split (SPLIT_TT_HOR). A multi-type tree leaf node is called a coding unit (CU), and this segmentation is used for prediction and transformation processing without any further partitioning, as long as the CU is not too large relative to the maximum transformation length. This means that in most cases, CUs, PUs, and TUs have the same block size in a quad tree with a nested multi-type tree coding block structure. The exception occurs when the maximum supported transformation length is smaller than the width or height of the color components of the CU. VVC has developed a unique signaling mechanism for partition split information in a quad tree with a nested multi-type tree coding tree structure. In the signaling mechanism, a coding tree unit (CTU) is treated as the root of a quaternary tree and is first partitioned by the quaternary tree structure.Each quota tree leaf node (if large enough to allow it) is then further partitioned by a multitype tree structure. In the multitype tree structure, a first flag (mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned, and if so, a second flag (mtt_split_cu_vertical_flag) is signaled to indicate the split direction, and then a third flag (mtt_split_cu_binary_flag) is signaled to indicate whether the split is a binary split or a terminal split. Based on the values ​​of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the CU's multitype tree split mode (MttSplitMode) can be derived by the decoder based on predefined rules or tables. For a given design, such as a 64x64 lumar block and 32x32 chroma pipeline design in a VVC hardware decoder, TT splitting is prohibited if either the width or height of the lumar coding block is greater than 64, as shown in Figure 6. TT splitting is also prohibited if either the width or height of the chroma coding block is greater than 32. The pipeline design divides the picture into virtual pipeline data units (VPDUs), defined as non-overlapping units within the picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. Since the VPDU size is roughly proportional to the buffer size in most pipeline stages, it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum translation block (TB) size. However, in VVC, terminal tree (TT) and binary tree (BT) partitions can increase the VPDU size.In addition, it should be noted that if any part of a tree node block extends beyond the bottom or right side of the picture boundary, the tree node block will be forced to split until all samples of the coded CU are located inside the picture boundary. For example, the Intra Sub-Partitions (ISP) tool can divide a luminance prediction block into two or four subpartitions, either vertically or horizontally, depending on the block size.

[0113] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.

[0114] As described above, the video encoder 20 is configured to determine or select the best or most optimal prediction mode from a set of prediction modes (for example, a predetermined set). The set of prediction modes may include, for example, an intra-prediction mode and / or an inter-prediction mode.

[0115] Intra-Prediction The set of intra-prediction modes may include 35 different intra-prediction modes, for example, non-directional modes such as DC (or average) mode and planar mode, or directional modes, as defined in HEVC. Alternatively, it may include 67 different intra-prediction modes, for example, non-directional modes such as DC (or average) mode and planar mode, or directional modes, as defined in VVC. As one example, some conventional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes for non-square blocks, for example, as defined in VVC. As another example, to avoid partitioning operations for DC prediction, only longer sides are used for non-square blocks to calculate the average. The results of planar mode intra-prediction can then be further modified by the Position-Dependent Intra-Prediction Coupling (PDPC) method.

[0116] The intra-prediction unit 254 is configured to use reconfigured samples of adjacent blocks of the same current picture to generate an intra-prediction block 265 according to the intra-prediction mode of a set of intra-prediction modes.

[0117] The intra-prediction unit 254 (or, generally, the mode selection unit 260) is further configured to output intra-prediction parameters (or, generally, information indicating the selected intra-prediction mode for a block) to the entropy encoding unit 270 in the form of syntax elements 266 for inclusion in the encoded video data 21. As a result, for example, the video decoder 30 can receive and use the prediction parameters for decoding.

[0118] Interlayer prediction (including inter-layer prediction) The set of interpretation modes (or possible modes) depends on the available reference picture (i.e., a previously decoded picture stored in DBP 230, for example) and other interpretation parameters, such as whether the entire reference picture or a portion of the reference picture, for example, the search window area around the current block's region, is used to search for the best matching reference block, and / or whether pixel interpolation is applied, for example, half / 4semi-pel, quarter-pel, and / or 1 / 16-pel, or not.

[0119] In addition to the prediction modes described above, skip mode, direct mode, and / or other interpretation modes may be applied.

[0120] For example, in Extended merge prediction, the merge candidate list for such modes is constructed by including, in order, the following five types of candidates: spatial MVP from spatially adjacent CUs, temporal MVP from collocated CUs, history-based MVP from a FIFO table, pairwise average MVP, and zero MV. Bilateral matching-based decoder-side motion vector refinement (DMVR) may be applied to improve the accuracy of the MV in the merge mode. Merge mode with MVD (MMDV) is derived from the merge mode using motion vector differences. The MMVD flag is signaled immediately after sending the skip flag and merge flag to specify whether the MMVD mode is used for the CU. Adaptive motion vector resolution (AMVR) scans at the CU level may be applied. AMVR allows the MVD of a CU to be coded with different accuracies. The MVD of the current CU can be adaptively selected depending on the prediction mode of the current CU. If the CU is coded in merge mode, a combined inter / intra prediction (CIIP) mode can now be applied to the CU. To obtain the CIIP prediction, a weighted average of the inter and intra prediction signals is performed. In affine motion compensation prediction, the affine motion field of a block is described by motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). Subblock-based temporal motion vector prediction (SbTMVP) is similar to temporal motion vector prediction (TMVP) in HEVC, but now predicts the motion vector of subCUs within the CU. Bi-directional optical flow (BDOF), formerly called BIO, is a more concise version that requires much less computation, particularly in terms of the number of multiplications and the size of the multipliers.In triangle partition mode, the CU is evenly divided into two triangular partitions using either diagonal or opposite-angle partitioning. In addition, bi-prediction mode extends beyond simple averaging to allow for a weighted average of two prediction signals.

[0121] The interpretation unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither of which are shown in Figure 2). The motion estimation unit may be configured to receive or acquire, for motion estimation, a picture block 203 (the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one or more previously reconstructed blocks, e.g., one or more other / different reconstructed blocks of the previously decoded picture 231. For example, a video sequence may include the current picture and the previously decoded picture 231. Or, in other words, the current picture and the previously decoded picture 231 may be, or may be, part of a sequence of pictures that make up the video sequence.

[0122] The encoder 20 may be configured to, for example, select a reference block from multiple reference blocks relating to the same or different pictures of multiple other pictures, and provide the reference picture (or reference picture index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an interprediction parameter to the motion estimation unit. This offset is also called the motion vector (MV).

[0123] The motion compensation unit is configured to perform interprediction based on or using the interprediction parameters to obtain, for example, received interprediction parameters and interprediction block 265. Motion compensation performed by the motion compensation unit may include fetching or generating predictive blocks based on the motion / block vector determined by motion estimation and performing interpolation to sub-pixel precision. Interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate predictive blocks that can be used to encode the picture block. Now receiving a motion vector for the picture block PU, the motion compensation unit may place the predictive block pointed to by the motion vector in one of the reference picture lists.

[0124] The motion compensation unit may also generate syntax elements related to blocks and video slices for use by the video decoder 30 when decoding picture blocks of video slices. In addition to, or as an alternative to, slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be generated or used.

[0125] Entropy Coding The entropy coding unit 270 is configured to apply or bypass (uncompress) an entropy encoding algorithm or scheme (e.g., variable-length coding (VLC) scheme, context-adaptive VLC scheme (CAVLC), arithmetic coding scheme, binariconversion, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or another entropy encoding method or technique) to the quantization coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements, and obtains encoded video data 21, which can be output via output 272 in the form of an encoded bitstream 21. As a result, the video decoder 30 can receive and use the decoded parameters. The encoded bitstream 21 may be sent to the video decoder 30 or stored in memory for later transmission or retrieval by the video decoder 30.

[0126] Other structural variations of the video encoder 20 can be used to encode video streams. For example, a non-conversion-based encoder 20 can directly quantize the residual signal for a given block or frame without a conversion processing unit 206. In another embodiment, the encoder 20 can combine a quantization unit 208 and an inverse quantization unit 210 into a single unit.

[0127] Decoder and decoding method Figure 3 shows an example of a video decoder 30 configured to implement the technology of this application. The video decoder 30 is configured to receive encoded video data 21 (e.g., encoded bitstream 21), e.g., encoded by an encoder 20, in order to obtain a decoded picture 331. The encoded video data or bitstream includes information for decoding the encoded video data, e.g., data representing picture blocks (and / or tile groups or tiles) and associated syntax elements of an encoded video slice.

[0128] In the example in Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transformation processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an interpretation unit 344, and an intraprediction unit 354. The interpretation unit 344 may be or may include a motion compensation unit. In some examples, the video decoder 30 can perform a decoding path that is largely reciprocal to the encoding path described with respect to the video encoder 100 from Figure 2.

[0129] As described with respect to encoder 20, the inverse quantization unit 210, inverse processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter-prediction unit 344, and intra-prediction unit 354 are also referred to as forming the “built-in decoder” of video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse processing unit 312 may be functionally identical to the inverse processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Accordingly, the descriptions provided for each unit and function relating to video decoder 20 apply to each unit and function relating to video 30.

[0130] Entropy Decoding The entropy decoding unit 304 is configured to analyze the bitstream 21 (or, generally, encoded video data 21) and, for example, perform entropy decoding on the encoded video data 21. For example, it obtains some or all of the quantization coefficients 309 and / or decoded coded parameters (not shown in Figure 3), such as inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transformation parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme, as described with respect to the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may further be configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit 360, and the other parameters to other units of the decoder 30. The video decoder 30 can receive syntax elements at the video slice level and / or video block level. In addition to slices and their respective syntax elements, or as an alternative, tile groups and / or tiles, and their respective syntax elements may be received and / or used.

[0131] Inverse Quantization The inverse quantization unit 310 can be configured to receive quantization parameters (QP) (or information generally related to inverse quantization) and quantization coefficients from encoded video data 21 (for example, by analysis and / or decoding by the entropy decoding unit 304), and to apply inverse quantization to the decoded quantized video data 309 based on the quantization parameters, and to obtain dequantization coefficients 311, which may also be referred to as conversion coefficients 311. The inverse quantization process may include using quantization parameters determined by the video encoder 20 for each video block in a video slice (or tile or tile group) to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.

[0132] Inverse Transform The inverse transformation processing unit 312 may be configured to receive the dequantization coefficients 311, also referred to as the transformation coefficients 311, and to apply a transformation to the dequantization coefficients 311 in order to obtain the reconstructed residual block 213 in the sample region. The reconstructed residual block 213 may also be referred to as the transformation block 313. The transformation may be an inverse transformation, such as an inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may further be configured to receive transformation parameters or corresponding information from the encoded video data 21 (for example, by analysis and / or decoding by the entropy decoding unit 304) and to determine the transformation to be applied to the dequantization coefficients 311.

[0133] Reconstruction The reconstruction unit 314 (e.g., an adder or summator 314) may be configured to add a reconstruction residual block 313 to a prediction block 365 in order to obtain a reconstruction block 315 within the sample region. This can be done, for example, by adding the sample values ​​of the reconstruction residual block 313 and the sample values ​​of the prediction block 365.

[0134] filtering The loop filter unit 320 is configured to filter the reconstructed block 315 (either within or after the coding loop) to obtain the filtered block 321, which, for example, smooths pixel transitions or otherwise improves video quality. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be deblocking filter, SAO, and then ALF. In another example, a process called chroma mapping with chroma scaling (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. In another example, the deblocking filter process may also be applied to internal subblock edges, such as affine subblock edges, ATMVP subblock edges, subblock transform (SBT) edges, and intra-subpartition (ISP) edges. Although the loop filter unit 320 is shown as an in-loop filter in Figure 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0135] Decoded picture buffer The decoded video block 321 of the picture is then stored in the decoded picture buffer 330. The buffer stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for the display of each output.

[0136] The decoder 30 is configured to output, for example, via output 312, the decoded picture 311 for presentation or display to the user.

[0137] Prediction The inter-prediction unit 344 may be identical to the inter-prediction unit 244 (particularly the motion compensation unit), and the intra-prediction unit 354 may be functionally identical to the inter-prediction unit 254, and performs partitioning or partitioning decisions and predictions based on partitioning and / or prediction parameters, or respective information received from the encoded video data 21 (e.g., by analysis and / or decoding by the entropy decoding unit 304). The mode application unit 360 may be configured to perform block-by-block predictions (intra-predictions or inter-predictions) based on the reconstructed picture, block, or each sample (filtered or unfiltered) in order to obtain prediction blocks 365.

[0138] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit 354 of the mode-applying unit 360 is configured to generate a prediction block 365 for the picture block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., motion compensation unit) of the mode-applying unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 304. For inter-prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 can construct the reference frame list, list 0 and list 1 using default construction techniques based on the reference pictures stored in the DPB 330. Embodiments using tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to, or as an alternative to, slices (e.g., video slices) may be applied identically or similarly. For example, video may be coded using I, P, or B tile groups and / or tiles.

[0139] The mode application unit 360 is configured to determine predictive information about video blocks in the current video slice by analyzing motion vectors or related information and other syntax elements, and to use the predictive information to generate predictive blocks of the current video block to be decoded. For example, the mode application unit 360 uses some of the received syntax elements to determine the predictive mode used to encode the video blocks in the video slice, the slice type of interprediction (e.g., B slice, P slice, or GPB slice), configuration information for one or more reference picture lists for the slice, motion vectors for each intercoded video block in the slice, the interprediction status for each intercoded video block in the slice, and other information for decoding the video blocks in the current video slice. The same or similar may apply to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or alternatively to slices (e.g., video slices). For example, video may be encoded using I, P, or B tile groups and / or tiles.

[0140] As shown in Figure 3, an embodiment of the video decoder 30 may be configured to partition and / or decode a picture by using slices (also referred to as video slices). Here, a picture may be partitioned or decoded using one or more slices (typically non-overlapping). Each slice may contain one or more blocks (e.g., CTUs) or one or more block groups (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).

[0141] As shown in Figure 3, an embodiment of the video decoder 30 may be configured to partition and / or decode a picture by using slice / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). Here, the picture may be partitioned and / or decoded using one or more slice / tile groups (typically non-overlapping). Each slice, each slice / tile group, may contain, for example, one or more blocks (e.g., CTUs) or one or more tiles. Here, each tile may be, for example, rectangular in shape and may contain one or more blocks (e.g., CTUs), for example, complete or divided blocks.

[0142] Other variations of the video decoder 30 can be used to decode the encoded video data 21. For example, the decoder 30 can generate an output video stream without a loop filtering unit 320. For example, a non-conversion-based decoder 30 can directly dequantize the residual signal for a given block or frame without an inverse conversion processing unit 312. In another embodiment, the video decoder 30 can combine the inverse quantization unit 310 and the inverse conversion processing unit 312 into a single unit.

[0143] It should be understood that in the encoder 20 and decoder 30, the processing result of the current step may be further processed and then output for the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.

[0144] It should be noted that further operations can be applied to the currently derived motion vectors of a block (including, but not limited to, control point motion vectors in affine mode, subblock motion vectors in affine, planar, and ATMVP modes, temporal motion vectors, etc.). For example, the value of a motion vector is constrained to a predefined range according to its representation bits. When the representation bit of a motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" signifies exponentiation. For example, if bitDepth is set to equal 16, the range is -32768 to 32767, and if bitDepth is set to equal 18, the range is -131072 to 131071. For example, the derived motion vector values ​​(e.g., the MV of a 4x4 subblock within an 8x8 block) are constrained such that the maximum difference between the integer parts of the MVs of four 4x4 subblocks is less than or equal to N pixels, such that the maximum difference is less than or equal to 1 pixel. Here, we provide two methods for constraining motion vectors according to bit depth.

[0145] Figure 4 is a schematic diagram of a video coding apparatus 400 according to one embodiment of the present disclosure. The video coding apparatus 400 is suitable for carrying out the disclosed embodiment as described herein. In one embodiment, the video coding apparatus 400 may be a decoder, such as the video decoder 30 in Figure 1A, or an encoder, such as the video encoder 20 in Figure 1A.

[0146] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an exit port 450 (or output port 450) for transmitting data, and memory 460 for storing data. The video coding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the exit port 450 for the input or output of optical or electrical signals.

[0147] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiver unit 420, the transmitter unit 440, the output port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the embodiments disclosed above. For example, the coding module 470 performs, processes, prepares, or provides various coding operations. Including the coding module 470 therefore provides a substantial improvement to the functionality of the video coding apparatus 400 and results in the transformation of the video coding apparatus 400 to different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.

[0148] Memory 460 includes one or more disks, tape drives, and solid-state drives and may be used as an overflow data storage device to store the program when such program is selected for execution and to store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).

[0149] Figure 5 is a schematic block diagram of a device 500 that can be used as either or both of the source device 12 and destination device 14 in Figure 1, according to one exemplary embodiment. The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or multiple devices currently existing or to be developed that are capable of manipulating or processing information. The disclosed embodiments can be implemented using a single processor, e.g., processor 502, as illustrated, but advantages in speed and efficiency may be achieved by using more than one processor.

[0150] The memory 504 in the device 500 may be a read-only memory (ROM) device or a random access memory (RAM) device in its implementation. Any other suitable type of storage device may be used as memory 504. Memory 504 may include code and data 506 accessed by the processor 502 using the bus 512. Memory 504 may further include an operating system 508 and an application program 510. The application program 510 includes at least one program that enables the processor 502 to perform the method described herein. For example, the application program 510 may include applications 1 through N and further include a video coding application that performs the method described herein. The device 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that can operate to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0151] Although shown here as a single bus, the bus 512 of device 500 may consist of multiple buses. Furthermore, the secondary storage device 514 may be directly coupled to other components of device 500 or accessed via a network. It may also include a single integrated unit such as a memory card, or multiple units such as multiple memory cards. Device 500 can therefore be implemented in a wide variety of configurations.

[0152] Scalable coding Scalable coding includes quality-scalable (PSNR-scalable), spatial-scalable, and others. For example, as shown in Figure 6, a sequence may be downsampled to a low spatial resolution version. Both the low spatial resolution version and the original spatial resolution (high spatial resolution) version are encoded. Generally, the low spatial resolution is encoded first and then used as a reference for the high spatial resolution which is encoded later.

[0153] To describe the layer information (number, dependencies, output), a VPS (Video Parameter Set) is defined as follows: [Table 1] `vps_max_layers_minus1 plus 1` specifies the maximum number of layers allowed for each CVS that references the VPS. A value of 1 for vps_all_independent_layers_flag specifies that all layers in CVS are coded independently without using interlayer prediction. A vps_all_independent_layers_flag equal to 0 indicates that one or more layers in CVS may use interlayer prediction. If it does not exist, the value of vps_all_independent_layers_flag is assumed to be equal to 1. If vps_all_independent_layers_flag is equal to 1, then the value of vps_independent_layer_flag[i] is presumed to be equal to 1. If vps_all_independent_layers_flag is equal to 0, the value of vps_independent_layer_flag[0] is assumed to be 1. vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For two non-negative integer values ​​m and n, if m is less than n, then the value of vps_layer_id[m] is less than vps_layer_id[n]. A vps_independent_layer_flag[i] equal to 1 specifies that the layer with index i does not use interlayer prediction. A vps_independent_layer_flag[i] equal to 0 indicates that the layer with index i can use interlayer prediction and that vps_layer_dependency_flag[i] exists within the VPS. A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. If vps_direct_dinercy_dependency_flag[i][j] does not exist for i and j within the range of 0 to vps_max_layers_minus1, it is presumed to be equal to 0. The variable DirectDependentLayerIdx[i][j] specifies the j-th direct dependency layer of the i-th layer, and is derived as follows:

number

number

[0154] DPB management and reference picture marking To manage these reference pictures in the decoding process, decoded pictures must be kept in a Decoded Picture Buffer (DPB) for reference use in subsequent picture decoding. To indicate these pictures, their Picture Order Count (POC) information must be signaled directly or indirectly in the slice header. Generally, there are two reference picture lists: list0 and list1. And to signal the pictures in the lists, the reference picture index must also be included. For uni predictions, reference pictures are fetched from one reference picture list, and for bi predictions, reference pictures are fetched from two reference picture lists. All reference pictures are stored in the DPB. All pictures in the DPB are marked as "used for long-term reference," "used for short-term reference," or "unused for reference," and there is only one of the three statuses. Once a picture is marked as "not used for reference," it will never be used for reference. It can also be removed from the DPB if it is not needed for output. The status of a referenced picture can be signaled within the slice header or derived from the slice header information. A new reference picture management method called the RPL (reference picture list) method has been proposed. RPL proposes an entire set or multiple sets of reference pictures for the currently coded picture, and the reference pictures within the reference picture set are used for decoding the current picture or future (later, or next) picture. Therefore, RPL reflects the picture information in the DPB, and even if a reference picture is not currently used for referencing a picture, it is necessary to store it in the RPL if it will be used for referencing a next picture. The picture, after being reconstructed, is saved within the DPB and marked as "for short-term reference" by default. DPB management operations begin after parsing the RPL information in the slice header.

[0155] Reference picture list configuration Reference picture information may be signaled via the slice header. Additionally, several RPL candidates may exist within the Sequence Parameter Set (SPS). In this case, the slice header may include an RPL index to obtain the required RPL information without signaling the entire RPL syntax structure. Alternatively, the entire RPL syntax structure may be signaled within the slice header.

[0156] Introduction of the RPL method To save on the cost bits of RPL signaling, several RPL candidates may exist within the SPS. The picture can use the RPL index (ref_pic_list_idx[i]) to obtain RPL information from the SPS. RPL candidates are signaled as follows: [Table 2]

[0157] The semantics are as follows: A flag equal to 1 for rpl1_sim_as_rpl0_flag indicates that the syntax structures num_ref_pic_lists_in_sps[1] and ref_pic_list_struct(1,rplsidx) do not exist, and then the following applies: - The value of num_ref_pic_lists_in_sps[1] is presumed to be equal to the value of num_ref_pic_lists_in_sps[0]. The value of each syntax element in -ref_pic_list_struct(1,rplsIdx) is presumed to be equal to the value of the corresponding syntax element in ref_pic_list_struct(0,rplsIdx) for rplsIdx in the range of 0 to num_ref_pic_lists_in_sps[0]-1. num_ref_pic_lists_in_sps[i] specifies the index of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure in the SPS that has a listIdx equal to 1. The value of num_ref_pic_lists_in_sps[i] is in the range of 0 to 64.

[0158] In addition to obtaining RPL information based on the RPL index from the SPS, RPL information can also be signaled in the slice header. [Table 3] A ref_pic_list_sps_flag[i] equal to 1 specifies that the reference picture list i of the current slice is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures in the SPS, which has a listIdx equal to i. A ref_pic_list_sps_flag[i] equal to 0 specifies that the reference picture list i of the current slice is derived based on the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, which has a listIdx equal to i directly contained within the slice header of the current picture. If ref_pic_list_sps_flag[i] does not exist, the following applies: - If num_ref_pic_lists_in_sps[i] is equal to 0, then the value of ref_pic_list_sps_flag[i] is presumed to be equal to 0. -Otherwise (num_ref_pic_lists_in_sps[i] is greater than 0), if rpl1_idx_present_flag is equal to 0, the value of ref_pic_list_sps_flag[1] is presumed to be equal to ref_pic_list_sps_flag[0]. - Otherwise, the value of ref_pic_list_sps_flag[i] is presumed to be equal to pps_ref_pic_list_sps_idc[i]-1. ref_pic_list_idx[i] specifies the index of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure that has a listIdx equal to i, which is currently used to derive the reference picture list i of the picture, into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures that have a listIdx equal to i contained within the SPS. The syntax element ref_pic_list_idx[i] is represented by the Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. If it does not exist, the value of ref_pic_list_idx[i] is assumed to be equal to 0. The value of ref_pic_list_idx[i] is in the range of 0 to num_ref_pic_lists_in_sps[i]-1. If ref_pic_list_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, then the value of ref_pic_list_idx[i] is presumed to be equal to 0. If ref_pic_list_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, then the value of ref_pic_list_idx[1] is presumed to be equal to ref_pic_list_idx[0]. The variable RplsIdx[i] is derived as follows:

number

number

number

[0159] The syntax structure of RPL is as follows: [Table 4] num_ref_entries[listIdx][rplsIdx] specifies the number of entries in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. The value of num_ref_entries[listIdx][rplsIdx] must be in the range of 0 to sps_max_dec_pic_buffering_minus1+14. A value of 0 for ltrp_in_slice_header_flag[listIdx][rplsIdx] indicates that the POC LSB for the LTRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure exists within the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. An ltrp_in_slice_header_flag[listIdx][rplsIdx] equal to 1 indicates that the POC LSB for an LTRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure does not exist in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. An inter_layer_ref_pic_flag[listIdx][rplsIdx][i] equal to 1 specifies that the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is an ILRP entry. A value of 0 for inter_layer_ref_pic_flag[listIdx][rplsIdx][i] indicates that the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is not an ILRP entry. If it does not exist, the value of inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is assumed to be equal to 0. A value equal to 1, st_ref_pic_flag[listIdx][rplsIdx][i], specifies that the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is a STRP entry. A value of 0 for st_ref_pic_flag[listIdx][rplsIdx][i] indicates that the i-th entry of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is an LTRP entry. If inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 0 and st_ref_pic_flag[listIdx][rplsIdx][i] does not exist, then the value of st_ref_pic_flag[listIdx][rplsIdx][i] is presumed to be equal to 1. The variable NumLtrpEntries[listIdx][rplsIdx] is derived as follows:

number

number

number

[0160] Some general explanations regarding the RPL structure For each list, an RPL structure exists. First, num_ref_entries[listIdx][rplsIdx] is signaled to indicate the number of reference pictures in the list. ltrp_in_slice_header_flag[listIdx][rplsIdx] is used to indicate whether LSB (least significant bit) information is signaled in the slice header. If the current reference picture is not an interlayer reference picture, st_ref_pic_flag[listIdx][rplsIdx][i] indicates whether it is a long-term reference picture. If it is a short-term reference picture, POC information (abs_delta_poc_st and strp_entry_sign_flag) is signaled. If ltrp_in_in_slice_header_flag[listIdx][rplsIdx] is zero, rpls_poc_lsb_lt[listIdx][rplsIdx][j+++] is used to derive the LSB information of the currently referenced picture. The MSB (most significant bit) can be derived directly or based on information in the slice header (delta_poc_msb_present_flag[i][j] and delta_poc_msb_cycle_lt[i][j]).

[0161] Decoding process for constructing a reference picture list This process is invoked for each slice of a non-IDR picture at the start of the decoding process. Reference pictures are handled through reference indices. A reference index is the index of the reference picture list. When decoding an I slice, the reference picture list is not used in decoding the slice data. When decoding a P slice, only reference picture list 0 (i.e., RefPicList[0]) is used in decoding the slice data. When decoding a B slice, both reference picture list 0 and reference picture list 1 (i.e., RefPicList[1]) are used in decoding the slice data. At the start of the decoding process for each slice of a non-IDR picture, the reference picture lists RefPicList[0] and RefPicList[1] are derived. The reference picture lists are used in the marking of reference pictures as defined in Clause 8.3.3, or in the decoding of slice data. Note 1 - For I slices of non-IDR pictures that are not the first slice of the picture, RefPicList[0] and RefPicList[1] may be derived for bitstream compatibility checks, but their derivation is not necessary to decode the current picture or any picture that follows the current picture in the decoding order. For P slices that are not the first slice of the picture, RefPicList[1] may be derived for bitstream compatibility checks, but its derivation is not necessary to decode the current picture or any picture that follows the current picture in the decoding order. The reference picture lists RefPicList[0] and RefPicList[1] are structured as follows:

number

number

[0162] Decoding process for reference picture marking This process is called once per picture, after the decoding of the slice header and the decoding process for constructing the referenced picture list for the slice, as specified in Clause 8.3.2, but before the decoding of the slice data. This process may result in one or more referenced pictures in the DPB that are marked as “not used for reference” or “for long-term reference.” A decoded picture in a DPB can be marked as “Not for Reference,” “Short-Term Reference,” or “Long-Term Reference,” but at any given point in time during the decoding process, it is only one of these three. Assigning one of these markings to a picture implicitly removes another of these markings, where applicable. When a picture is referred to as being marked as “For Reference,” this collectively refers to pictures that are marked as either “Short-Term Reference” or “Long-Term Reference” (but not both). STRPs and ILRPs are identified by their nuh_layer_id and PicOrderCntVal values. LTRPs are identified by their nuh_layer_id value and the Log2(MaxLtPicOrderCntLsb) LSB of their PicOrderCntVal value. If the current picture is a CLVSS picture, all current referenced pictures in the DPB that have the same nuh_layer_id as the current picture (if any) will be marked as "reference not used". Otherwise, the following applies: - For each LTRP entry in RefPicList[0] or RefPicList[1], if the referenced picture is a STRP with the same nuh_layer_id as the current picture, that picture is marked as "long-term reference". - Each referenced picture that has the same nuh_layer_id as the current picture in the DPB and is not referenced by any entry in RefPicList[0] or RefPicList[1] is marked as "reference not used". - For each ILRP entry in RefPicList[0] or RefPicList[1], the referenced picture is marked as "for long-term reference".

[0163] Note that the ILRP (Interlayer Reference Picture) is marked as "for long-term reference" here.

[0164] Within SPS, there are two syntaxes related to interlayer reference information. [Table 5] If sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id for the VPS referenced by the SPS. If sps_video_parameter_set_id is equal to 0, the SPS does not reference a VPS, and the VPS is not referenced when decoding each CVS by referencing the SPS. A long_term_ref_pics_flag equal to 0 specifies that LTRP will not be used for interpretation of any coded pictures in CVS. A long_term_ref_pics_flag equal to 1 specifies that LTRP may be used for interpretation of one or more coded pictures in CVS. A value of 0 for inter_layer_ref_pics_present_flag indicates that ILRP will not be used for inter-prediction of any coded picture in CVS. A value of 1 for inter_layer_ref_pics_flag indicates that ILRP may be used for inter-prediction of one or more coded pictures in CVS. If sps_video_parameter_set_id is equal to 0, the value of inter_layer_ref_pics_present_flag is assumed to be 0.

[0165] A brief explanation is as follows: The `long_term_ref_pics_flag` flag is used to indicate whether LTRP (Long-Term Reverse Processing) may be used in the decoding process. The `inter_layer_ref_pics_present_flag` flag is used to indicate whether ILRP may be used in the decoding process.

[0166] Therefore, if inter_layer_ref_pics_present_flag is equal to 1, an ILRP used in the decoding process may exist and is marked as "for long-term reference". In this case, even if long_term_ref_pics_flag is equal to 0, an LTRP used in the decoding process exists. Thus, there is a contradiction with the semantics of long_term_ref_pics_flag. In existing methods, some syntax elements for interlayer reference information are always signaled without considering the current layer index. This invention proposes adding several conditions to the syntax elements to improve signaling efficiency.

[0167] Since long_term_ref_pics_flag is used solely to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, its semantic is modified to control the parsing of flag parsing in RPL. Syntax elements for interlayer reference information are signaled, taking into account the index of the current layer. If the information can be derived from the index of the current layer, the information does not need to be signaled.

[0168] Since long_term_ref_pics_flag is used solely to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, its semantics are modified to control the parsing of flag parsing in RPL. Syntax elements for interlayer reference information are signaled, taking into account the index of the current layer. If the information can be derived from the index of the current layer, the information does not need to be signaled.

[0169] First Embodiment of the Invention [Semantics] (Modify the semantics of long_term_ref_pics_flag to eliminate inconsistencies between LTRP and ILRP)

[0170] Since long_term_ref_pics_flag is used solely to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, the semantics are modified as follows: A long_term_ref_pics_flag equal to 1 specifies that ltrp_in_slice_header_flag and st_ref_pic_flag exist within the syntax structure ref_pic_list_struct(listIdx,rplsIdx). A long_term_ref_pics_flag equal to 0 specifies that these syntax elements do not exist within the syntax structure ref_pic_list_struct(listIdx,rplsIdx). Being equal to 0 specifies that LTRP is not used for interpretation of any coded pictures in CVS. A long_term_ref_pics_flag equal to 1 specifies that LTRP may be used for interpretation of one or more coded pictures in CVS. Furthermore, the semantics can also be modified as follows to exclude ILRP: A long_term_ref_pics_flag equal to 0 indicates that LTRP is not used for interpretation of coded pictures in CVS. A long_term_ref_pics_flag equal to 1 indicates that LTRP may be used for interpretation of one or more coded pictures in CVS, where LTRP does not include ILRP (Interlayer Reference Picture).

[0171] Second Embodiment of the Present Invention [VPS]

[0172] Proposal 1: Conditional signaling of vps_direct_direct_dependency_flag[i][j] (Interlayer reference information is signaled considering the index of the current layer, eliminating redundant signaling and improving coding efficiency.)

[0173] Option 1.A: Note that if i is equal to 1, it means that layer 1 must reference other layers. On the other hand, since only layer 0 can be a reference layer, vps_direct_direct_dependency_flag[i][j] does not need to be signaled. Only if i is greater than 1 does vps_direct_direct_dependency_flag[i][j] need to be signaled. [Table 6] A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 indicates that the layer with index j is a direct reference layer for the layer with index i. If vps_direct_dinercy_dinercy_flag[i][j] does not exist within the range of i and j from 0 to vps_max_layers_minus1, then if i is equal to 1 and vps_independent_layer_flag[i] is equal to 0, then vps_direct_direct_dependent_flag[i][j] is presumed to be equal to 1; otherwise, it is presumed to be equal to 0.

[0174] Option 1.B: In addition to the above implementation method (Option 1.A), there is also Option 1.B. This means that for i and j in the range including 0 to i-1, and for j in the range including 0 to i-2, if all values ​​of vps_direct_direct_dependency_flag[i][j] are equal to 0, then the value of vps_direct_direct_dependency_flag[i][i-1] does not need to be signaled and is presumed to be equal to 1. [Table 7] A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 indicates that the layer with index j is a direct reference layer for the layer with index i. If vps_direct_dinercy_dinercy_flag[i][j] does not exist within the range of i and j from 0 to vps_max_layers_minus1, then vps_direct_direct_dependenticy_flag[i][j] is presumed to be equal to 1 if vps_independent_layer_flag[i] is equal to 0, j is equal to 0, and the value of SumDependencyFlag is equal to 0; otherwise, it is presumed to be equal to 0.

[0175] Proposal 2: Semantic constraints of vps_direct_direct_dependency_flag[i][j] Furthermore, we can impose constraints on the semantics of vps_direct_direct_depency_flag[i][j] without changing the syntax signaling method or syntax table. Basically, for i, if the layer with index i is a dependent layer (vps_independent_layer_flag[i] is equal to 0), then at least one value of vps_direct_direct_dependency_flag[i][j] is equal to 1 for j in the range of 0 to i-1. Alternatively, the sum of vps_direct_direct_dependency_flag[i][j] should not be equal to 0 for j in the range of 0 to i-1, or it should be greater than or equal to 1 (e.g., >=1), or it should be greater than 0 (e.g., >0).

[0176] Option 2.A: A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 indicates that the layer with index j is a direct reference layer for the layer with index i. vps_direct_dinercy_dependency_flag[i][j] is assumed to be equal to 0 if, for i and j, they do not exist within the range from 0 to vps_max_layers_minus1. Here, for i and j within the range from 0 to i-1, and if vps_independent_layer_flag[i] is equal to 0, then the sum of vps_direct_direct_depency_flag[i][j] is greater than 0.

[0177] Option 2.B: A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 indicates that the layer with index j is a direct reference layer for the layer with index i. A vps_direct_dinercy_dependency_flag[i][j] is assumed to be equal to 0 if, for i and j, they do not exist within the range from 0 to vps_max_layers_minus1. Here, for i and j within the range from 0 to i-1, and if vps_independent_layer_flag[i] is equal to 0, then at least one value of vps_direct_direct_dependency_flag[i][j] is equal to 1.

[0178] Proposal 3: Proposal 1+Proposal 2

[0179] Option 3: In practice, options 1 and 2 can be combined to create other implementation methods. The same applies to Operation 1.B + Operation 2.B. [Table 8] A vps_direct_direct_dependency_flag[i][j] equal to 0 indicates that the layer with index j is not a direct reference layer for the layer with index i. A vps_direct_direct_dependency_flag[i][j] equal to 1 indicates that the layer with index j is a direct reference layer for the layer with index i. If vps_direct_direct_dependency_flag[i][j] does not exist for i and j within the range from 0 to vps_max_layers_minus1, then vps_independence_layer_flag[i] is presumed to be equal to 1 if vps_independence_layer_flag[i] is equal to 0, j is equal to i-1, and the value of SumDependencyFlag is equal to 0; otherwise, it is presumed to be equal to 0. Here, for i and j within the range from 0 to i-1, and if vps_independent_layer_flag[i] is equal to 0, then at least one value of vps_direct_direct_dependency_flag[i][j] is equal to 1.

[0180] The method of joining is not limited here and can also be done as follows: The same applies to Operation 1.A + Operation 2.B. Operation 1.A + Operation 2.A are similar. The same applies to Operation 1.B + Operation 2.A.

[0181] Third embodiment of the present invention [sps] [sps] (Interlayer reference information is signaled considering the index of the current layer, eliminating redundant signaling and improving coding efficiency.) Note here that if sps_video_parameter_set_id is equal to 0, it means that there are no multiple layers, therefore there is no need to signal inter_layer_ref_pics_flag, and the flag is 0 by default. [Table 9] A value of 0 for inter_layer_ref_pics_present_flag indicates that ILRP will not be used for inter-prediction of any coded picture in CVS. A value of 1 for inter_layer_ref_pics_flag indicates that ILRP may be used for inter-prediction of one or more coded pictures in CVS. If sps_video_parameter_set_id is 0 and inter_layer_ref_pics_flag does not exist, the value of inter_layer_ref_pics_present_flag is assumed to be equal to 0. Here, note that if GeneralLayerIdx[nuh_layer_id] is equal to 0, the current layer is layer 0 and cannot reference any other layers. Therefore, there is no need to signal inter_layer_ref_pics_present_flag, and its value is 0 by default. [Table 10] A value of 0 for inter_layer_ref_pics_present_flag indicates that ILRP will not be used for inter-prediction of coded pictures in CVS. A value of 1 for inter_layer_ref_pics_flag indicates that ILRP may be used for inter-prediction of one or more coded pictures in CVS. If sps_video_parameter_set_id is 0 and inter_layer_ref_pics_flag does not exist, the value of inter_layer_ref_pics_present_flag is assumed to be equal to 0. The following is an example of another application that codes for both of the above cases. [Table 11] A value of 0 for inter_layer_ref_pics_present_flag indicates that ILRP will not be used for inter-prediction of coded pictures in CVS. A value of 1 for inter_layer_ref_pics_flag indicates that ILRP may be used for inter-prediction of one or more coded pictures in CVS. If sps_video_parameter_set_id is 0 and inter_layer_ref_pics_flag does not exist, the value of inter_layer_ref_pics_present_flag is assumed to be equal to 0.

[0182] Fourth Embodiment of the Present Invention [RPL] Here, we note that if GeneralLayerIdx[nuh_layer_id] is equal to 1, the current layer is layer 1 and can only refer to layer 0, while layer 0's ilrp_idc must be 0. Therefore, in this case, there is no need to signal irp_idc. [Table 12] irp_idc[listIdx][rplsIdx][i] directly specifies the ILRP index of the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure for the list of dependency layers. The value of lrp_idc[listIdx][rplsIdx][i] is assumed to be within the range of 0 to GeneralLayerIdx[nuh_layer_id]-1. If GeneralLayerIdx[nuh_layer_id] is equal to 1, the value of lrp_idc[listIdx][rplsIdx][i] is assumed to be equal to 0.

[0183] Fifth Embodiment of the Present Invention [combination] It should be noted that some or all of the embodiments from Embodiments 1 to 4 can be combined to form new embodiments. For example, Embodiment 1 + Embodiment 2 + Embodiment 3 + Embodiment 4, or Embodiment 2 + Embodiment 3 + Embodiment 4, or other combinations.

[0184] The following describes the encoding method, the decoding method, and the application of a system using them, as shown in the embodiments described above.

[0185] Figure 7 is a block diagram showing a content supply system 3100 for realizing a content distribution service. This content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 may include, but is not limited to, Wi-Fi®, Ethernet®, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0186] The capture device 3102 can generate data and encode the data by an encoding method as shown in the embodiments described above. Alternatively, the capture device 3102 can distribute the data to a streaming server (not shown in the figure), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or tablet, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. If the data includes video, the video encoder 20 included in the capture device 3102 may actually perform the video encoding process. If the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform the audio encoding process. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and encoded video data separately to the terminal device 3106.

[0187] In the content supply system 3100, the terminal device 310 receives and reproduces the encoded data. The terminal device 3106 may be any device having a data receiving and recovering capability that can decode such encoded data, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conference system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination of the above. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform an audio decoding process.

[0188] For a terminal device having its own display, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can feed the decoded data to its display. For a terminal device not provided with a display, such as STB 3116, a video conference system 3118, or a video surveillance system 3120, an external display 3126 is connected thereto to receive and display the decoded data.

[0189] When each device in this system performs encoding or decoding, a video encoding device or a video decoding device may be used as shown in the above embodiments.

[0190] FIG. 8 is a diagram showing a configuration according to an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. This protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any combination of any of these types.

[0191] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to a demultiplexing unit 3204. The demultiplexing unit 3204 can separate multiplexed data into encoded audio data and encoded video data. In some practical scenarios described above, for example, in a video conference system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is transmitted to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0192] Through demultiplexing, a video elementary stream (ES), an audio ES, and optionally a subtitle are generated. The video decoder 3206, including the video decoder 30 as described in the above embodiment, decodes the video ES by the decoding method shown in the above embodiment to generate video frames and feeds this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames and feeds this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212.

[0193] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the representation of video and audio information. The information can be coded using syntax with timestamps on the representation of coded audio and visual data, and timestamps on the delivery of the data stream itself.

[0194] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and supplies the video / audio / subtitle to the video / audio / subtitle display 3216.

[0195] The present invention is not limited to the system described above, and either the video encoding device or the video decoding device in the above-described embodiment may be incorporated into other systems, such as vehicle systems.

[0196] Mathematical Operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real-valued division are defined. Numbering and counting conventions generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 2nd, and so on.

[0197] Arithmetic operators The following arithmetic operators are defined as follows: + Addition. - Subtraction (as a two-argument operator) or negation (as a unary prefix operator). * Includes multiplication and matrix multiplication. x y The exponent. Specifies x raised to the power of y. In other contexts, such notation is used as a superscript not intended to be interpreted as an exponential function. Integer division where the result is truncated toward zero. For example, 7 / 4 and -7 / -4 are truncated toward 1, and -7 / 4 and 7 / -4 are truncated toward -1. ÷ is used to indicate division in a mathematical expression where rounding or truncation is not intended.

number

number

[0198] Logical operators The following logical operators are defined as follows. x&&y Boolean logical "and" for x and y. x||y Boolean logical "or" for x and y. ! Boolean logical "not". x?y:z If x is true (TRUE) or non-zero, the expression evaluates to the value of y; otherwise, the expression evaluates to the value of z.

[0199] Relational operators The following relational operators are defined as follows. > Greater than. >= Greater than or equal to. < Less than. <= Less than or equal to. = Equal to. != Not equal to.

[0200] When a relational operator is applied to a syntax element or variable that has been assigned the value "na" (not applicable), the value "na" is treated as a distinct value of the syntax element or variable. The value "na" is considered not to be equal to any other value.

[0201] Bit-wise operators The following bit-wise operators are defined as follows. & Bit-wise "and". When operating on integer arguments, it operates on the two's complement representation of the integer value. When operating on a binary argument that contains fewer bits than another argument, the shorter argument is extended by adding higher-order bits equal to 0. | Bitwise "OR". When operating on integer arguments, it operates on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than another argument, the shorter argument is extended by adding more significant bits equal to 0. ^ Bitwise "exclusive OR". When operating on integer arguments, it operates on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than another argument, the shorter argument is extended by adding more significant bits equal to 0. x >> y Arithmetically right shift the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has a value equal to the MSB of x before the shift operation. x << y Arithmetically left shift the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.

[0202] Assignment operator The following arithmetic operators are defined as follows. = Assignment operator. ++ Increment, that is, x++ is equivalent to x = x + 1. When used in array indexing, it evaluates to the value of the variable before the increment operation. -- Decrement, that is, x-- is equivalent to x = x - 1. When used in array indexing, it evaluates to the value of the variable before the decrement operation. += Increment by a specified amount, that is, x += 3 is equivalent to x = x + 3, and x += (-3) is equivalent to x = x + (-3). -= Decrement by a specified amount, that is, x -= 3 is equivalent to x = x - 3, and x -= (-3) is equivalent to x = x - (-3).

[0203] Range notation The following notation is used to specify a range of values. x = y . . zx takes integer values ​​from y to z, where x, y, and z are integers, and z is greater than or equal to z.

[0204] mathematical function The following mathematical function is defined.

number

number

number

number

number

number

number

[0205] Order of operation precedence If the precedence in an expression is not explicitly indicated using parentheses, the following rules apply: - Higher-priority operations are evaluated before lower-priority operations. - Operations with the same priority are evaluated sequentially from left to right. The following table shows the order of operations from highest to lowest, with higher positions in the table indicating higher priority. In the C programming language, as well as with respect to the operators used, the precedence used in this specification is the same as the precedence used in the C programming language. [Table 13] Table: Priority of operations from highest (top of table) to lowest (bottom of table)

[0206] Text description of logical operations In text, logical operation statements are mathematically written in the following format:

number

[0207] Text description of logical operations In text, logical operation statements are mathematically written in the following format:

number

[0208] In text, logical operation statements are mathematically written in the following format: if(condition 0) statement 0 if(condition 1) statement 1 The above can be described in the following way: when condition 0, statement 0 When condition 1, statement 1

[0209] While embodiments of the present invention have been described primarily in relation to video coding, it should be noted that embodiments of the coding system 10, encoder 20, and decoder 30 (and corresponding systems 10), as well as other embodiments described herein, may also be configured for still image processing or coding, i.e., processing or coding of individual pictures independent of any preceding or consecutive pictures, as in video coding. Generally, when image processing coding is limited to a single picture 17, only the interpretation units 244 (encoder) and 344 (decoder) may not be available. All other functionalities (also referred to as tools or technologies) of the video encoder 20 and video decoder 30 can be equally used for still image processing, e.g., residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, and entropy coding 270 and entropy decoding 304.

[0210] For example, embodiments of the encoder 20 and decoder 30, and, for example, functions related to the encoder 20 and decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on a computer-readable medium or transmitted as one or more instructions or codes over a communication medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of computer programs from one location to another, for example, according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transient, tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this disclosure. A computer program product may include a computer-readable medium.

[0211] In particular, as shown in Figure 9, a method is provided for decoding an encoded video bitstream implemented in the decoder. This method includes: S901, obtaining a first syntax element (i.e., vps_independent_layer_flag[i]) from the encoded video bitstream that specifies whether the first layer uses interlayer prediction; S902, obtaining one or more second syntax elements (i.e., vps_direct_direct_depency_flag[i][j]) from the encoded video bitstream that relate to one or more second layers, each second syntax element specifying whether the second layer is a direct reference layer of the first layer, where if the value of the first syntax element specifies that the first layer is allowed to use interlayer prediction, at least one of the one or more second syntax elements has a value that specifies that the second layer is a direct reference layer of the first layer. Furthermore, S903, performing interlayer prediction on the picture of the first layer by using the picture of the second layer associated with at least one second syntax element as a reference picture.

[0212] Similarly, as shown in Figure 10, a method is provided for encoding a video bitstream containing coded data implemented in an encoder. This method includes: S1001, determining whether at least one second layer is a direct reference layer of the first layer; S1003, encoding a syntax element into a coded video bitstream, where the syntax element specifies whether the first layer uses interlayer prediction, where if none of the at least one second layer is a direct reference layer of the first layer, the value of the syntax element specifies that the first layer does not use interlayer prediction.

[0213] Figure 11 shows a decoder 1100 configured to decode a video bitstream containing coded data for multiple pictures. Decoder 1100 according to the example shown includes an acquisition unit 1110 and a prediction unit 1120. The acquisition unit 1110 is configured to acquire a first syntax element from the coded video bitstream that specifies whether a first layer uses interlayer prediction. The acquisition unit 1110 is further configured to acquire one or more second syntax elements related to one or more second layers, each second syntax element specifying whether the second layer is a direct reference layer of the first layer. Here, if the value of the first syntax element specifies that the first layer is allowed to use interlayer prediction, then at least one of the one or more second syntax elements has a value that specifies that the second layer is a direct reference layer of the first layer. The prediction unit 1120 is configured to perform interlayer prediction on the picture of the first layer by using the picture of the second layer associated with at least one second syntax element as a reference picture.

[0214] Here, the unit may be a software module for execution by a processor, or a processing circuit.

[0215] Here, the acquisition unit 1110 may be an entropy decoding unit 304. The prediction unit 1120 may be an interpretation unit 344. The decoder 1100 may be a destination device 14, a decoder 30, a device 500, a video decoder 3206, or a terminal device 3106.

[0216] Similarly, an encoder 1200 is provided, configured to encode a video bitstream containing encoded data for multiple pictures, as shown in Figure 12. The encoder 1200 includes a decision unit 1210 and an encoding unit 1220. The decision unit 1210 is configured to determine whether at least one second layer is a direct reference layer of the first layer. The encoding unit 1220 is configured to encode a syntax element into the encoded video bitstream, where the syntax element specifies whether the first layer uses interlayer prediction. If none of the at least one second layer is a direct reference layer of the first layer, the value of the syntax element specifies that the first layer does not use interlayer prediction.

[0217] Here, the unit may be a software module for execution by a processor, or a processing circuit.

[0218] The first coding unit 1210 and the second coding unit 1220 may be entropy coding units 270. The decision unit may be a mode selection unit 260. The encoder 1200 may be a source device 12, an encoder 20, or a device 500.

[0219] As an example, and not limited to, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium used to store desired program code in the form of instructions or data structures and which can be accessed by a computer. Also, any connection is appropriately called a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other transient media, but rather are directed toward non-transient, tangible storage media. Disks (disk and disc) as used herein include compact discs (CDs), laserdiscs, optical discs, digital multipurpose discs (DVDs), floppy disks (registered trademark), and Blu-ray discs. Here, a disc typically reproduces data magnetically, while another disc reproduces data optically using a laser. The above combinations should also be included within the scope of computer-readable media.

[0220] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term “processor” may refer to any of the aforementioned structures or any other structure suitable for implementing the technology described herein. In addition, in some embodiments, the functions described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the technology may be fully implemented in one or more circuits or logic elements.

[0221] The technology of this disclosure can be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight the functional aspects of devices configured to implement the disclosed technology, but implementation by different hardware units is not necessarily required. Rather, as described above, the various units may be combined within a codec hardware unit, or provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.

Claims

1. A method for decoding an encoded video bitstream, The steps include obtaining a first syntax element having index i from the encoded video bitstream, which specifies whether the first layer having index i uses interlayer prediction, The step is to obtain one or more second syntax elements associated with one or more second layers from the encoded video bitstream, wherein each second syntax element having index i and index j specifies whether the second layer having index j is a direct reference layer of the first layer having index i. If the value of the first syntax element is equal to 0, j is in the range from 0 to i-1, and at least one of the second syntax elements having index j is equal to 1. If all of the second syntax elements are 0 in the range of j from 0 to i-2, then the second syntax element for j being i-1 is presumed to be equal to 1 and is not signaled. Steps and A step of obtaining a third syntax element from the encoded video bitstream, wherein a third syntax element equal to 0 specifies that the Long-Term Reference Picture (LTRP) is not used for interpretation of any encoded picture in the encoded video sequence (CVS), and a third syntax element equal to 1 specifies that the LTRP may be used for interpretation of one or more encoded pictures in the CVS. Methods that include...

2. The first syntax element equal to 1 specifies that the first layer does not use interlayer prediction, or The first syntax element equal to 0 indicates that the first layer is permitted to use interlayer prediction. The method according to claim 1.

3. A second syntax element equal to 0 indicates that the second layer associated with the second syntax element is not a direct reference layer of the first layer, or The second syntax element equal to 1 specifies that the second layer associated with the second syntax element is a direct reference layer of the first layer. The method according to claim 1 or 2.

4. The step of obtaining one or more second syntax elements is performed if the value of the first syntax element indicates that the first layer is permitted to use interlayer prediction. The method according to any one of claims 1 to 3.

5. The first syntax element is represented as vps_independent_layer_flag[i], The method according to any one of claims 1 to 4.

6. The third syntax element is represented as long_term_ref_pics_flag, The method according to any one of claims 1 to 5.

7. The third syntax element is included in the sequence parameter set (SPS) of the encoded video bitstream. The method according to any one of claims 1 to 6.

8. A method for encoding an encoded video bitstream, wherein the method is The steps include encoding a first syntax element into the encoded video bitstream, wherein the first syntax element has an index i and specifies whether the first layer having the index i uses interlayer prediction. The step of encoding one or more second syntax elements associated with at least one second layer into the encoded video bitstream, wherein each second syntax element having index i and index j specifies whether the second layer having index j is a direct reference layer of the first layer having index i. If the value of the first syntax element is equal to 0, j is in the range from 0 to i-1, and at least one of the second syntax elements having index j is equal to 1. If all of the second syntax elements are 0 in the range of j from 0 to i-2, then the second syntax element for j being i-1 is presumed to be equal to 1 and is not signaled. Steps and A step of encoding a third syntax element into the encoded video bitstream, wherein a third syntax element equal to 0 specifies that the Long-Term Reference Picture (LTRP) is not used for interpretation of any encoded picture in the encoded video sequence (CVS), and a third syntax element equal to 1 specifies that the LTRP may be used for interpretation of one or more encoded pictures in the CVS. Methods that include...

9. The first syntax element equal to 1 specifies that the first layer does not use interlayer prediction, or The first syntax element equal to 0 indicates that the first layer is permitted to use interlayer prediction. The method according to claim 8.

10. A second syntax element equal to 0 indicates that the second layer associated with the second syntax element is not a direct reference layer of the first layer, or The second syntax element equal to 1 specifies that the second layer associated with the second syntax element is a direct reference layer of the first layer. The method according to claim 8 or 9.

11. The step of encoding one or more second syntax elements associated with at least one second layer into the encoded video bitstream is performed if the value of the first syntax element indicates that the first layer is permitted to use interlayer prediction. The method according to any one of claims 8 to 10.

12. The first syntax element is represented as vps_independent_layer_flag[i], The method according to any one of claims 8 to 11.

13. The third syntax element is represented as long_term_ref_pics_flag, The method according to any one of claims 8 to 12.

14. The third syntax element is included in the sequence parameter set (SPS) of the encoded video bitstream. The method according to any one of claims 8 to 13.

15. An encoder comprising a processing circuit for performing the method according to any one of claims 8 to 14.

16. A decoder comprising a processing circuit for performing the method according to any one of claims 1 to 7.

17. A computer program including program code, When the aforementioned program code is executed by the computer's processor, The computer is made to carry out the method according to any one of claims 1 to 7. Computer program.

18. A computer program including program code, When the aforementioned program code is executed by the computer's processor, The computer is made to carry out the method according to any one of claims 8 to 14. Computer program.

19. It is a decoder, One processor, A non-temporary computer-readable storage medium coupled to the processor and storing the programming executed by the processor, Includes, The programming described above is configured such that, when executed by the processor, it causes the decoder to perform the method according to any one of claims 1 to 7. decoder.

20. It is an encoder, One processor, A non-temporary computer-readable storage medium coupled to the processor and storing the programming executed by the processor, Includes, The programming described above is configured, when executed by the processor, to cause the encoder to perform the method according to any one of claims 8 to 14. Encoder.

21. A non-temporary computer-readable storage medium for transporting program code, When executed by a computer device, The computer device is made to carry out the method according to any one of claims 1 to 7. A non-temporary, computer-readable storage medium.

22. A non-temporary computer-readable storage medium for transporting program code, When executed by a computer device, The computer device is made to carry out the method according to any one of claims 8 to 14. A non-temporary, computer-readable storage medium.

23. A device for storing and transmitting a bitstream, The apparatus includes a receiver, a processor, a transmitter, and a storage medium, wherein the receiver is configured to receive a bitstream, the storage medium is configured to store the bitstream, and the transmitter is configured to transmit the bitstream. The bitstream includes a first syntax element, one or more second syntax elements, and a third syntax element. The first syntax element has an index i and specifies whether the first layer having the index i uses interlayer prediction. Each second syntax element having index i and index j specifies whether the second layer having index j is a direct reference layer of the first layer having index i. The third syntax element, which is equal to 0, specifies that the Long-Term Reference Picture (LTRP) is not used for interpretation of any coded picture in the coded video sequence (CVS). The third syntax element equal to 1 specifies that LTRP may be used for interpretation of one or more coded pictures in the CVS, If the value of the first syntax element is equal to 0, j is in the range from 0 to i-1, and at least one of the second syntax elements having index j is equal to 1. If all of the second syntax elements are 0 in the range of j from 0 to i-2, the second syntax element for j being i-1 is presumed to be equal to 1 and is not signaled. The processor analyzes the bitstream to obtain the first syntax element, the one or more second syntax elements, and the third syntax element. Device.

24. A system for a bitstream, comprising an encoding device, a decoding device, and one or more storage media, The encoding device is configured to acquire a video signal and encode the video signal to obtain one or more bitstreams. Each of the one or more bitstreams includes a first syntax element, one or more second syntax elements, and a third syntax element. The first syntax element has an index i and specifies whether the first layer having the index i uses interlayer prediction. Each second syntax element having index i and index j specifies whether the second layer having index j is a direct reference layer of the first layer having index i. The third syntax element, which is equal to 0, specifies that the Long-Term Reference Picture (LTRP) is not used for interpretation of any coded picture in the coded video sequence (CVS). The third syntax element equal to 1 specifies that LTRP may be used for interpretation of one or more coded pictures in the CVS, If the value of the first syntax element is equal to 0, j is in the range from 0 to i-1, and at least one of the second syntax elements having index j is equal to 1. If all of the second syntax elements are 0 in the range of j from 0 to i-2, the second syntax element for j being i-1 is presumed to be equal to 1 and is not signaled. The one or more storage media are used to store the one or more bitstreams. The decoding device is used to decode one or more bitstreams. system.