Multilayer video coding with constrained complexity

By designing an adaptive encoding configuration for multi-layer video, the problem of high computational complexity in multi-layer video encoding is solved, and parallel processing and resource efficiency of the decoder are achieved.

CN121128163APending Publication Date: 2025-12-12INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480032344.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-16
Filing Date
2024-05-13
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing video coding standards suffer from high computational complexity in multi-layer video coding, making it difficult to effectively control the decoding processing capabilities of each layer, resulting in uneven demand for decoder resources.

Method used

An adaptive coding configuration is adopted, with different coding configurations designed for each layer of multi-layer video. Operational constraints are introduced to reduce computational complexity and support parallel processing in the decoder.

Benefits of technology

It reduces computational complexity in multi-layer video coding, supports parallel processing of decoders, and improves resource utilization efficiency and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128163A_ABST
    Figure CN121128163A_ABST
Patent Text Reader

Abstract

Apparatuses and methods including techniques for encoding video data are disclosed. The disclosed techniques include obtaining video data, the video data including a multi-layer video. The techniques further include determining an encoding configuration that accommodates each layer of the multi-layer video, and encoding the layers into a bitstream based on the respective encoding configuration of the layers. Further, an apparatus and method including the technique for decoding video data are disclosed. The disclosed techniques include obtaining a bitstream encoding video data including a multi-layer video. The techniques further include decoding encoding configurations from the bitstream, which are determined to accommodate respective layers of the multi-layer video, and decoding the layers from the bitstream based on the respective encoding configurations of the layers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of European Application No. 23305785.0, filed on 16 May 2023, which is incorporated herein by reference in its entirety. Background Technology

[0002] Coding standards such as VVC define coding configurations through profiles, layers, and levels. These configurations impose operational constraints that limit the range of coding tools and parameters that the encoder can use for video compression. The profile (and associated layers and levels) upon which the encoder generates the encoded video bitstream is signaled to the decoder so that the decoder can verify whether computational resources can be allocated for decoding the bitstream. Similarly, multi-layer video can be encoded based on multi-layer profiles. Decoders conforming to multi-layer profiles need to have the processing power to decode multi-layer video. In principle, for each layer, the entire set of tools available in the encoder can be used, and therefore the processing power required by a multi-layer decoder can be the sum of the processing power required to decode each layer. However, in practice, and based on specific application requirements, different sets of tools can be applied to the encoding of each layer to control the processing power required to decode each layer while maintaining acceptable compression performance. Summary of the Invention

[0003] The aspects disclosed in this disclosure describe methods for encoding video data. The methods include obtaining video data, the video data comprising multiple layers of video. The methods further include determining encoding configurations adapted to each layer of the multiple-layer video; and encoding the layers into a bitstream based on the respective encoding configurations of the layers. The aspects disclosed in this disclosure also describe methods for decoding video data. The methods include obtaining a bitstream encoding video data comprising multiple layers of video. The methods further include decoding encoding configurations (determined to adapt to each layer of the multiple-layer video) from the bitstream; and decoding the layers from the bitstream based on the respective encoding configurations of the layers.

[0004] The aspects disclosed in this disclosure describe apparatus for encoding video data. The apparatus includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the apparatus to obtain video data, the video data comprising multiple layers of video. The instructions further cause the apparatus to determine an encoding configuration adapted to each layer of the multiple layers of video; and to encode the layers into a bitstream based on the respective encoding configurations of the layers. The aspects disclosed in this disclosure also describe apparatus for decoding video data. The apparatus includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the apparatus to obtain a bitstream encoding video data comprising multiple layers of video. The instructions further cause the apparatus to decode an encoding configuration (determined to adapt to each layer of the multiple layers of video) from the bitstream; and to decode the layers from the bitstream based on the respective encoding configurations of the layers.

[0005] Further aspects disclosed in this disclosure describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for encoding video data. The method includes obtaining video data, the video data comprising multiple layers of video. The method further includes determining an encoding configuration adapted to each layer of the multiple-layer video; and encoding the layers into a bitstream based on the respective encoding configuration of each layer. Aspects disclosed in this disclosure also describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for decoding video data. The method includes obtaining a bitstream encoding video data comprising multiple layers of video. The method further includes decoding an encoding configuration (determined to adapt to each layer of the multiple-layer video) from the bitstream; and decoding the layers from the bitstream based on the respective encoding configuration of each layer.

[0006] This summary is provided to introduce the chosen concepts in a simplified form, which are further described in the detailed embodiments below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to addressing any or all of the shortcomings mentioned in any part of this disclosure. Attached Figure Description

[0007] Figure 1 This is a block diagram of an example system, based on which various aspects of this embodiment can be implemented.

[0008] Figure 2 This is a functional block diagram of an example video encoder, based on which various aspects of this embodiment can be implemented.

[0009] Figure 3This is a functional block diagram of an example video decoder, based on which various aspects of this embodiment can be implemented.

[0010] Figure 4 This diagram illustrates a CU decoding pipeline, which enables various aspects of this embodiment.

[0011] Figure 5 This diagram illustrates a multi-layer encoding / decoding workflow, which allows for the implementation of various aspects of this embodiment.

[0012] Figure 6 This diagram illustrates a simplified CU decoding pipeline, which enables various aspects of this embodiment to be implemented.

[0013] Figure 7 The diagram illustrates a further simplified CU decoding pipeline, which enables various aspects of this embodiment to be implemented.

[0014] Figure 8 This diagram illustrates inter-layer image prediction, which enables various aspects of this embodiment to be achieved.

[0015] Figure 9 This is a flowchart of an example method for encoding multi-layer video, according to which various aspects of this embodiment can be implemented.

[0016] Figure 10 This is a flowchart of an example method for decoding multi-layer video, according to which various aspects of this embodiment can be implemented. Detailed Implementation

[0017] This paper presents a system and method for encoding and decoding multi-layer video. An encoding configuration adapted to each layer of the multi-layer video is designed based on various factors. The proposed adaptive encoding configuration introduces operational constraints that enable an overall reduction in computational complexity and enable parallel processing in the decoder. (See below for further details.) Figure 1-3 The traditional systems and methods for predictive video coding are described, followed by references. Figure 4-10 A description of various aspects of this disclosure.

[0018] Figure 1A block diagram of an example system 100 is illustrated. System 100 can be implemented as a device and can be configured to perform one or more aspects of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100 can be implemented individually or in combination in integrated circuits, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, the processing 110 and encoder / decoder 130 elements of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input ports and / or output ports.

[0019] System 100 includes at least one processor 110, which can be configured to execute instructions loaded therein for implementing various aspects, such as those described in this application. Processor 110 may include embedded memory, input and output interfaces, and various other circuitry as known in the art. System 100 includes at least one memory 120 (such as a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drive, and / or optical disk drive. For example, storage device 140 may be an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0020] System 100 includes an encoder / decoder module 130 configured to process data to provide encoded or decoded video data. The encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 may be implemented as a separate element of system 100, or it may be incorporated into processor 110 as a combination of hardware and / or software as known to those skilled in the art. Furthermore, the encoder / decoder module 130 represents one or more modules that can be implemented in separate devices to perform encoding and / or decoding functions.

[0021] Program code to be loaded into processor 110 or encoder / decoder 130 to execute the various aspects described herein may be stored in storage device 140 and subsequently loaded into memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more entries of various entries during execution of the processes described herein. Such stored entries may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, operational logic, and intermediate or final results from processing equations, formulas, or operations.

[0022] In several embodiments, the internal memory of processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing functions required during encoding or decoding. However, in other embodiments, external memory of the processing device (e.g., where the processing device may be processor 110 or encoder / decoder module 130) may be used for one or more of these functions. External memory may be memory 120 and / or storage device 140, which may include, for example, dynamically volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external dynamically volatile memory (such as RAM) is used as working memory for video encoding and decoding operations.

[0023] Input to the components of system 100 can be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to: (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a composite input terminal (COMP); (iii) a USB input terminal; and / or (iv) an HDMI input terminal.

[0024] In various embodiments, the input device of block 105 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select a signal band, which may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing some of these functions, such as down-converting the received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or down-converting it to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0025] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing integrated circuit or within processor 110. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface integrated circuit or within processor 110. Demodulation, error correction, and demultiplexing streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.

[0026] Various components of system 100 can be provided within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data therebetween using a suitable connection arrangement 115 (e.g., internal buses as known in the art, including I2C buses, wiring, and printed circuit boards).

[0027] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC). The communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0028] In various embodiments, a Wi-Fi network (such as IEEE 802.11) may be used to stream data to system 100. In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. In other embodiments, a set-top box may be used to stream data to system 100, delivering data via an HDMI connection to input block 105, or an RF connection to input block 105 may be used to stream data to system 100.

[0029] System 100 can provide output signals to various output devices, including display device 165, audio devices (e.g., speakers (one or more)) 175, and other peripheral devices 185. In various examples of embodiments, other peripheral devices 185 include one or more of a standalone DVR, disk player, stereo system, lighting system, and other devices that provide output functionality based on system 100. In various embodiments, control signals are transmitted between system 100 and display device 165, audio device 175, or other peripheral devices 185 using signaling such as AV.link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. In electronic devices (e.g., televisions), display device 165 and audio device 175 can be integrated into a single unit with other components of system 100. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0030] Alternatively, for example, if the RF portion of input 105 is part of a separate set-top box, the display device 165 and audio device 175 can be separate from one or more other components. In various embodiments where the display device 165 and audio device 175 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0031] Figure 2 The diagram illustrates a functional block diagram of a sample video encoder 200. This can be found in the reference... Figure 1 The system 100 described employs a video encoder 200. For example, the video encoder 200 may be an encoder that operates according to encoding standards such as Advanced Video Coding (AVC, H.264 / MPEG-4 | ISO / IEC 14496-10), High Efficiency Video Coding (HEVC, ITU-T H.265 | ISO / IEC 23008-2), or Universal Video Coding (VVC, standard ITU-T H.266, ISO / IEC 23090-3, 2020).

[0032] Before encoding, the video data can be preprocessed by a pre-encoding processor (not shown). Such preprocessing may include applying color model transformations (e.g., from RGB 4:4:4 to YCbCr 4:2:0) to the color components of the input video frame or mapping the color components of the input video frame to obtain a more compression-resistant signal distribution (e.g., applying a histogram equalizer and / or denoising filter to one or more of the color components of the video frame). Preprocessing may also include associating metadata with the video data that can be attached to the encoded video bitstream.

[0033] In encoder 200, video frames are encoded by encoder elements as generally described below. The image segmenter 202 segments the original video images (frames) to be encoded into coding units (i.e., original blocks). Typically, a coding unit (CU) contains a luma block and a corresponding chroma block, and therefore, the operations described herein as applied to CUs are typically applied to both the luma and chroma blocks. After segmentation 202, each CU can be encoded using either intra-frame prediction mode or inter-frame prediction mode. In intra-frame prediction mode, CU prediction is performed by intra-frame predictor 260. In intra-frame prediction mode, the content of a CU in a frame is predicted using reconstructed versions of one or more other CUs from the same frame (obtainable from the output of adder 255). In inter-frame prediction mode, motion estimation and motion compensation are performed by motion estimator 275 and motion compensator 270, respectively. In inter-frame prediction mode, the content of a CU in a frame is predicted using reconstructed versions of one or more other CUs from adjacent frames (obtainable from reference image buffer 280). The encoder 205 determines which prediction result (the prediction result obtained through the operation in intra-frame prediction mode 260 or the prediction result obtained through the operation in inter-frame prediction modes 270, 275) to use for encoding the CU, and indicates the selected prediction mode, for example, through a prediction mode flag. The selected prediction result can then be enhanced (e.g., filtered) by the prediction enhancer 285 to output the corresponding prediction block. Once a prediction block has been generated for each CU, the corresponding residual block is computed, for example, by subtracting the predicted CU (i.e., the prediction block) from the original CU (i.e., the original block).

[0034] The corresponding residual block of the CU, or its segmentation (i.e., transform block), is then transformed into a coefficient block by transformer 220; that is, the residual samples of the transform block are transformed into transform coefficients of the coefficient block. The resulting coefficient block is quantized by quantizer 230. Next, entropy encoder 245 entropy-encodes the quantized coefficient block and the corresponding encoding parameters (e.g., syntax elements including motion vectors and other control data). Thus, the entropy-encoded quantized coefficient block and the corresponding encoding parameters associated with each video frame of the original video are packed into a bitstream of encoded video data.

[0035] As described above, along with the encoding of the original block (CU), encoder 200 reconstructs the encoded original block to provide a reference for future predictions. Therefore, the quantized coefficient block (provided by quantizer 230) is dequantized by inverse quantizer 240 and then inverse transformed by inverse transformer 250 to reconstruct (decode) the residual block of the corresponding original block. The reconstructed residual block is added to the corresponding prediction block 255 to produce the corresponding reconstructed original block. Then, loop filter 265 can be applied to the reconstructed image (formed from the reconstructed original block) to perform, for example, deblocking filtering and / or sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered reconstructed image can then be stored in reference image buffer 280, which can be used for future predictions in inter-frame prediction modes. Therefore, encoder 200 also performs decoding operations 240, 250 to reconstruct the encoded image (frame). As explained above, the reconstructed image can then be stored in reference image buffer 280 and used to facilitate motion estimation 275 and compensation 270.

[0036] Figure 3 The diagram illustrates a functional block diagram of the example video decoder 300. See the reference... Figure 1 The described system 100 employs a video decoder 300. Typically, the operation of the video decoder 300 is the reverse of that of the video encoder 200. In the decoder 300, the bitstream of encoded video data generated by the video encoder 200 is first entropy-decoded by an entropy decoder 330 to decode quantized coefficient blocks and various coding parameters from the bitstream. The quantized coefficient blocks are dequantized by an inverse quantizer 340 and then inverse-transformed by an inverse transformer 350 to decode (reconstruct) the corresponding residual blocks. The reconstructed residual blocks are added 355 to the corresponding prediction blocks to produce the corresponding reconstructed original blocks. Depending on the selected prediction mode, the predicted original blocks can be obtained 370 from an intra-frame predictor 360 or from a motion compensator 375, and can then be enhanced (e.g., filtered) by a prediction enhancer 390 to generate prediction blocks. A loop filter 365 can be applied to the reconstructed picture (formed from the reconstructed original blocks) to output the reconstructed (decoded) video frame. The filtered reconstructed picture is also stored in a reference picture buffer 380 to facilitate motion compensation 375.

[0037] A post-decoding processor (not shown) can further process the reconstructed video. For example, post-decoding processing may include inverse color model transformation (e.g., a conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse mapping to reverse the mapping process performed by the pre-encoding processor. The post-decoding processor may use metadata derived by the pre-encoding processor and / or signaled in the video bitstream.

[0038] The aspects disclosed herein are described with reference to a CU; however, the described aspects are similarly applicable to any region (i.e., video data region) of a video frame to which encoding tools can be applied by an encoder 260 or a decoder 360. Generally, the aspects described herein can be applied to video data regions formed by video segments of any shape or size. The CU includes a luminance component. Y and chromaticity components Cr and Cb (Each of these is also referred to in this article) C ).

[0039] Figure 4 This diagram illustrates the CU decoding pipeline 400. It uses an encoding tool (developed by...). Figure 4 The building block representation shown) is used to extract from the encoded video bitstream (e.g., as per [reference]). Figure 2 The encoder generates the current CU as described above (e.g., as described in the encoder description). Figure 3 (As described in the decoder). Table 1 describes some of the encoding tools.

[0040] Table 1: Decoder Encoding Tools

[0041] exist Figure 4In the example, entropy decoding is first performed on the bitstream data associated with the current CU to extract information, including coding parameters and the coding residual of the current CU. A picture buffer stores a reconstructed picture that can be used as a reference for inter-frame prediction. Coding tools—including motion compensation (MC), BCW, BDOF, DMVR, and LMCS (luminance mapping)—perform operations associated with predicting the content of the current CU in inter-frame prediction mode 450. Coding tools—including intra-frame prediction and cross-component (CC)-based intra-frame prediction—perform operations associated with predicting the content of the current CU in intra-frame prediction mode 460. Therefore, the predicted CU 480 can be generated by intra-frame prediction 460 and / or by inter-frame prediction 450 (see adder 485). Note that the combined intra-frame-inter-frame prediction (CIIP) function in VVC can be used to create a prediction that is a combination of both intra-frame and inter-frame predictions, mixed with associated signaling weights. Encoding tools—including LFNST inverse, inverse quantization and transform (IQ / IT), and LMCS (chroma scaling)—perform operations associated with the residuals of reconstructing the current CU470. The residuals of the current CU470 are then added to the predicted current CU480 (see adder 495) to obtain the reconstructed current CU490. Loop filtering is then applied to the reconstructed current CU490 and the other reconstructed CUs of the current image to generate the reconstructed image. Therefore, the loop filtering operation can be performed by one or more encoding tools (including deblocking filters, SAO, and ALF tools) after applying LMCS (inverse luma mapping).

[0042] exist Figure 4 In the examples, the dependency on the data stream originating from the entropy decoder and the image buffer is illustrated with a solid line (e.g., 410); the dependency on the data stream associated with the reconstructed adjacent (one or more) CUs is illustrated with a double line (e.g., 420); and the dependency on the data stream associated with the current luminance component is illustrated with a double line. Y The dependencies of related data streams are illustrated using dashed lines (e.g., 430); and for the current chroma component... CThe dependencies of the associated data streams are illustrated using dotted lines (e.g., 440). These dependencies need to be satisfied to allow the correct operation of the decoding pipeline. Specifically, dependency 410 needs to be satisfied to decode coding parameters and to perform inter-frame prediction of the current CU by referencing previously decoded images from the image buffer. When decoding the current CU, dependencies associated with reconstructed CUs in the neighborhood of the current CU 420 need to be satisfied to perform intra-frame prediction of the current CU. For example, a CC-based intra-frame prediction tool depends on one or more reconstructed CUs 420 in its neighborhood to derive a model (e.g., CCLM) and depends on the reconstructed current luma component 430 to perform prediction. In the Enhanced Compression Model (ECM)—the software used by the Joint Video Experts Group (JVET) to study future video coding standards after VVC—dependencies associated with reconstructed CUs in the neighborhood of the current CU also need to be satisfied to apply Template Matching (TM) modes. TM modes use the reconstructed neighborhood of the current CU to predict the coding parameters of the current CU (such as motion vectors, intra-coding directions, or other coding modes).

[0043] Dependencies in the decoding pipeline also exist in multi-layer coding. VVC enables multi-layer coding using multi-layer profiles. For example, multi-layer coding can be very useful for use cases that require video scalability and / or multi-view (or 360°) video representation. In multi-layer coding, the layers of the video source are encoded separately or have inter-layer dependencies, and the resulting individual bitstreams are then merged into a single bitstream, which is delivered to the decoder for multi-layer decoding, as discussed in [the original text]. Figure 5 As further described.

[0044] Figure 5 This is a diagram illustrating the multi-layer encoding / decoding workflow 500. (Example) Figure 5 As shown in the example, encoder 510 and decoder 570 are connected via multiplexer 540, channel 550 (e.g., one or more wired or wireless communication links), and demultiplexer 560. In encoder 510, separate encoder instances can be used for each layer. For example, a separate encoder instance can be used. N Layer encoders, such as encoder layer 0 (530.0), encoder layer 1 (530.1), and encoder layer... N (530.N), they encode their respective inputs, which are: source layer 0 (520.0), source layer 1 (520.1), and source layer... N(520.N). For example, each source can correspond to different resolution levels of the same input image—from the lowest resolution level 520.0 to the intermediate resolution levels 520.1 to 520.N-1, to the highest resolution level 520.N. Additionally, a source layer (e.g., such as 520.1) can be encoded relative to a reference image provided by the next level encoder (e.g., encoder 530.0) (e.g., encoded by encoder 530.1). This creates interlayer dependencies due to interlayer predictions in both the encoder 510 end and the decoder 570 end (illustrated by dashed arrows).

[0045] Therefore, for each frame of the video sequence, layer encoder 530.0-N is initiated to encode the image of the corresponding layer, where a lower-layer encoder is used before the higher-layer encoder due to the dependency introduced by inter-layer prediction. For each frame, the output bitstream (i.e., sub-bitstream) of each layer encoder 530.0-N is multiplexed 540 into a single bitstream. Encoded in this bitstream is also the layer ID information (identifier), which associates the sub-bitstream with their respective layers. For example, the layer ID information can be signaled at the Network Abstraction Layer Unit (NALU) header (a NALU is a syntax structure that includes an indication of the type of data that follows, and bytes containing data in the form of a raw byte sequence payload, which, if necessary, is distributed with emulation prevention bytes). Layer encoders 530.0-N share layer-independent data structures, such as APS (Adaptive Parameter Set) and VPS (Video Parameter Set), which can also be encoded into the bitstream.

[0046] The bitstream generated by encoder 510 is delivered to decoder 570 via channel 550 for layered decoding. The channelized bitstream is first demultiplexed 560 into separate sub-bitstreams for each layer. Thus, the sub-bitstream of layer 0 can be independently decoded by decoder layer 0 (580.0) to generate the reconstructed source layer 0 (590.0), and the sub-bitstream of layer 1 can be decoded by decoder layer 1 (580.1) depending on the output from decoder layer 0 (580.0) (from which it receives a reference image) to generate the reconstructed source layer 1 (590.1), and the layer... N The sub-bitstream can be generated by the decoder layer. N (580.N) depends on the decoder layer. N The output of -1 (from which it receives a reference image) is decoded to generate the reconstructed source layer. N (590.N). Therefore, for each frame, decoding a multi-layer bitstream requires decoding the layers from lower layers before decoding the layers from higher layers. Figure 5 In the example, layer 0 is first decoded as 580.0, then layer 1 can be decoded as 580.1, and finally layer... NIt can be decoded as 580.N. Note that each layer decoder 580.0-N requires full decoding capabilities, including entropy decoding, inverse quantization, inverse transform, prediction, and loop filtering. This can result in high processing complexity at the decoder end, proportional to the number of layers that must be decoded for each frame.

[0047] The VVC standard defines the Output Layer Set (OLS) as a collection of layers for which one or more layers are designated as output layers. Therefore, a given layer only needs to be decoded if it is defined as an output layer or if it is required as a dependent layer of the output layer. OLSs are typically defined within VPSs. OLSs are signaled by indicating the output layers of the corresponding OLS. Other layers belonging to that OLS are derived from the layer dependencies indicated in the VPS. OLSs are used to (i) signal which layers can be output (if several layers in the same OLS are indicated as output layers, they are output together) and which layers will be decoded, and (ii) associate certain attributes (such as the PTL of each OLS) with each VPS. In practice, for spatially scalable video, the first OLS contains only the base layer as output, the second OLS contains the second layer as output and the base layer as a dependency, and the third OLS contains the third layer as output and the second and base layers as dependencies (the second layer is a direct dependency of the third layer, and the base layer is an indirect dependency required to decode the second layer). For each OLS, the complexity constraints can be different, as shown by the PTL associated with it.

[0048] To ensure practicality in implementing the full syntax defined by the VVC standard, a limited subset of the syntax is defined in VVC through “profiles,” “hierarchies,” and “levels” (i.e., PTLs). A profile and its associated hierarchies and levels represent a coding configuration that translates into different processing power requirements that the decoder must meet when decoding the individual bitstreams. The VVC PTLs are defined in Appendix A of the VVC specification (see Draft 10 (JVET-T2001) General Video Coding Editing Refinement, 20th Meeting, via teleconference, October 7-16, 2020).

[0049] A “profile” is a subset of the entire bitstream syntax specified in VVC. Within the bounds defined by a given profile, there is still considerable variation in encoder and decoder performance depending on the values ​​used for the syntax elements encoded into the bitstream (such as the specified size of the decoded images). In many applications, implementing a decoder capable of handling all hypothetical values ​​of the syntax elements in a particular profile is neither practical nor economical. To address this issue, “hierarchies” and “levels” are specified in each profile. Thus, a level is a set of specific constraints applied to the values ​​of syntax elements encoded into (or signaled in) the bitstream. Some of these constraints are expressed as simple restrictions on the values, while others take the form of constraints on arithmetic combinations of values ​​(e.g., image width multiplied by image height multiplied by the number of images decoded per second). Typically, levels specified for lower hierarchies are more constrained than those specified for higher hierarchies.

[0050] The `profile_tier_level()` syntax provides level information and optionally provides profiles, tiers, sub-profiles, and general constraint information for one or more OLSs. The OLS associated with a given `profile_tier_level()` syntax is defined as a scoped OLS (i.e., `OlsInScope`). When a `profile_tier_level()` syntax is included in a VPS, `OlsInScope` is one or more OLSs specified by the VPS. A list of PTLs is defined in the VPS, and for each OLS, an index indicating which PTL the current OLS references is signaled. Therefore, a given PTL can be associated with one or more OLSs. When a `profile_tier_level()` syntax is included in a Sequence Parameter Set (SPS), `OlsInScope` is an OLS that only includes the lowest-level layer in the layers referencing the SPS, and this lowest-level layer is an independent layer. Furthermore, a given level can be specified for a sub-layer of a layer, associated with a `TemporalId` value (suitable for time-scalable use cases).

[0051] The VVC standard also defines Network Abstraction Layer (NAL) units, Decoded Picture Buffers (DPBs), Picture Units (PUs), and Access Units (AUs). A NAL unit is a syntax structure that includes an indication of the type of data following it, and bytes containing data in the form of a Raw Byte Sequence Payload (RBSP), which may be scattered with emulation-prevention bytes. A DPB is a buffer (e.g., reference picture buffers 280 and 380) that holds a decoded picture for a hypothetical reference decoder, specifying a reference, output reordering, or output delay. A PU is a set of NAL units associated with each other according to specified classification rules, which are consecutive in decoding order and contain exactly one encoded picture. In other words, a PU represents a bitstream associated with a picture. As an example, a PU may contain several slices (in which case, the PU contains consecutive slice NAL units, as well as associated SPS, APS, and SEI NAL units). An AU is a set of PUs belonging to different layers and containing encoded pictures to be output from the DPB simultaneously. In other words, in a multi-layer context, an AU represents a bitstream associated with a single picture represented by multiple layers. When there is only one layer, AU contains only one image.

[0052] For the purpose of comparing the capabilities of different tiers, a tier with a general_tier_flag of 0 (i.e., the primary tier) is considered a lower tier than a tier with a general_tier_flag of 1 (i.e., a higher tier). For the purpose of comparing the capabilities of different levels, a specific level within a particular tier is considered a lower level than some other levels within the same tier when the value of general_level_idc or sublayer_level_idc[i] for that specific level is less than the values ​​for other levels.

[0053] For an OLS with an OLS index TargetOlsIdx, the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, and PicSizeMaxInSamplesY, along with the applicable dpb_parameters() syntax structure, are derived as follows: -If NumLayersInOls[TargetOlsIdx] equals 1 PicWidthMaxInSamplesY is set to be equal to sps_pic_width_max_in_luma_samples, PicHeightMaxInSamplesY is set to be equal to sps_pic_height_max_in_luma_samples, and PicSizeMaxInSamplesY is set to be equal to PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, where sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples exist in the SPS referenced by the layer in OLS, and the applicable dpb_parameters() syntax structure also exists in that SPS.

[0054] - Otherwise (NumLayersInOls[TargetOlsIdx] is greater than 1). PicWidthMaxInSamplesY is set to equal vps_ols_dpb_pic_width[MultiLayerOlsIdx[TargetOlsIdx]], PicHeightMaxInSamplesY is set to equal vps_ols_dpb_pic_height[MultiLayerOlsIdx[TargetOlsIdx]], PicSizeMaxInSamplesY is set to equal PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, and the applicable dpb_parameters() syntax structure is identified by vps_ols_dpb_params_idx[MultiLayerOlsIdx[TargetOlsIdx]] which exists in the VPS.

[0055] Table 2 specifies the limits for each level of each tier, excluding level 15.5 (i.e., Table 135 in the VVC standard). The tier and level of the bitstream conformance are indicated by the syntax elements general_tier_flag and general_level_idc, respectively; and the level of the sublayer conformance is indicated by the syntax element sublayer_level_idc[i], as follows: - If the specified level is not level 15.5, then according to the tier constraints specified in Table 2, a general_tier_flag of 0 indicates compliance with the primary tier, and a general_tier_flag of 1 indicates compliance with a higher tier. Furthermore, for levels below level 4 (corresponding to entries marked with "-" in Table 2), general_tier_flag should be equal to 0. Otherwise (if the specified level is level 15.5), the bitstream consistency requirement is that general_tier_flag should be equal to 1, and the value of general_tier_flag 0 is reserved for future use by ITU-T | ISO / IEC, and the decoder should ignore the value of general_tier_flag.

[0056] -general_level_idc and sublayer_level_idc[i] should be set to the value of general_level_idc equal to the level number specified in Table 2.

[0057] Table 2: General Hierarchy and Level Restrictions

[0058] The following describes the profile-specific level limitations. When the specified level is not level 15.5, the value of dpb_max_dec_pic_buffering_minus1[Htid]+1 should be less than or equal to MaxDpbSize, which is derived as follows: In Table 2, MaxLumaPs is specified, maxDpbPicBuf equals 8, and dpb_max_dec_pic_buffering_minus1[Htid] exists in or is derived from the applicable dpb_parameters() syntax structure. Assuming numDecPics is the number of images in AU n, the variable AuSizeMaxInSamplesY[n] is set to equal PicSizeMaxInSamplesY * numDecPics.

[0059] Table 3 shows the layer and level restrictions for video profiles (i.e., Table 136 in the VVC standard). According to Table 3, the restrictions for a given profile, level, and layer refer to all layers of a given OLS. That is, for example, if two layers must be decoded, the MaxLumaSr for one level will be doubled compared to a single level to set the correct level to use.

[0060] Table 3: Hierarchical and Level Restrictions for Video Profiles

[0061] Below is an example of a Multilayer Main 10 profile. Bitstreams conforming to Multilayer Main 10 should adhere to the following constraints: - The referenced SPS should have sps_chroma_format_idc (grey or 4:2:0 YUV video) equal to 0 or 1.

[0062] - The referenced SPS should have sps_bitdepth_minus8 in the range of 0 to 2, including (8 or 10 bits).

[0063] - The referenced SPS should have a sps_palette_enabled_flag equal to 0 (palette tool not used).

[0064] - In a bitstream conforming to the multi-level master 10 profile, the general_level_idc and sublayer_level_idc[i] values ​​of all i in the reference VPS (when available) and reference SPS should not be equal to 255 (which indicates level 15.5).

[0065] - The hierarchical and horizontal constraints specified in the multi-level master 10 profile should be met (if applicable).

[0066] The consistency of the bitstream with the multi-level master 10 profile is indicated by general_profile_idc, which is equal to 17.

[0067] A decoder conforming to a specific level of a multi-level master 10 profile should be able to decode all bitstreams that meet all of the following conditions: - The bitstream is indicated as conforming to multi-level master 10, master 10 or master 10 still image profile.

[0068] - The bitstream is indicated to conform to a level lower than or equal to the specified level.

[0069] - The bitstream is indicated to be a level that is not level 15.5 and is lower than or equal to the specified level.

[0070] In addition to PTL information, the PTL syntax structure may also include a General Constraint Information (GCI) syntax structure, which contains a list of constraint flags and non-flag syntax elements that indicate specific constraint attributes of the bitstream. The GCI definition is part of the VVC standard. When present, a GCI syntax element value greater than 0 indicates that the bitstream is constrained in a specific way, typically indicating that a specific encoding tool is not used in the bitstream; while a value of 0 indicates that the associated constraints may not apply, allowing (but not requiring) the use of a specific encoding tool in the bitstream (if its use is supported in the indicated profile).

[0071] GCI structures contain several types of constraint syntax elements, including: - Flags used for general bitstream constraints, such as flags indicating that only intra-frame coding is used, all layers are coded independently, or the bitstream contains only one AU; - Fields that constrain the bit depth and chroma format of the encoded image; - A flag indicating that certain NAL cell types are not allowed to appear in the bitstream; - Constraints are flags that indicate how an image can be segmented into slices, tiles, and sub-images in a bitstream; - Constrain the size of the CTU and the size and type of the segmentation tree; - A flag that restricts the use of specific intra-frame coding tools; - A flag that restricts the use of specific inter-frame coding tools; - Flags for constraint transformation, quantization, and residual coding tools; and - Markers of various aspects of the constrained loop filter.

[0072] A decoder conforming to a multi-layer profile needs to be able to decode several layers per frame. For each image in a layer, the encoder's complete toolset can be used. Therefore, the processing power required by the decoder can be the sum of the processing power required to decode each layer. There are no constraints defined by the PTL or profile / level, and there is a constant maximum total complexity for any number of layers. This limits the processing power overhead of the decoder when decoding multiple layers (compared to decoding a single layer) while preserving video properties such as image resolution or frame rate.

[0073] This disclosure describes encoding configurations for multi-layer coding that result in reduced decoding complexity without compromising compression performance. Typically, combinations of encoding tools described herein are used in a more constrained manner for encoding at least one layer of a multi-layer video. Low-complexity multi-layer profiles and low-latency multi-layer profiles are defined, wherein bitstreams conforming to these profiles can be decoded using a decoder conforming to a single-layer profile with only minor adjustments to the decoder.

[0074] As referenced above Figure 4 As described above, decoding of the current CU via the decoding pipeline 400 may be highly constrained by the dependencies of the encoding tools. These dependencies prevent the decoder from achieving high parallelism. Table 3 summarizes the dependencies of the encoding tools and their impact on the complexity of the decoding operation. Some tools are heavily burdened by high dependencies 420 in the core of the decoding loop, such as intra-frame prediction and LMCS tools. Other tools increase the latency of the CU decoding process (due to additional computation and memory access), such as BDOF and DMVR tools. The LFNST tool adds an additional step to the inverse transform. The deblocking filter, SAO, and ALF tools do not increase the CU decoding dependency, but additional computation for loop filtering is required after the CU is reconstructed 490 and before the reconstructed image is stored in the image buffer for future reference. In the hardware decoder, these tools are designed to process the maximum resolution and frame rate in real time with little margin. Processing multi-layer video adds additional frames (for each layer) to be processed.

[0075] Table 3: Dependency requirements of decoder encoding tools. .

[0077] The aspects described in this paper aim to reduce the computational power required to decode multi-layer bitstreams. This can be achieved by both reducing the complexity of each stage of the decoding pipeline and reducing the number of decoding stages. Figure 6 and Figure 7 Examples of reducing dependencies between CUs in a decoding pipeline are shown. In each of these examples, an encoding configuration is determined for which the constrained encoding tools have been deactivated, resulting in a more parallelizable and lower-latency pipeline, as further described below.

[0078] Figure 6 This is a simplified diagram of the CU decoding pipeline 600. Figure 6 In the example, the dependency on the data stream originating from the entropy decoder and the image buffer is illustrated with a solid line (e.g., 610), the dependency on the data stream associated with adjacent (one or more) reconstruction CUs is illustrated with a double line (e.g., 620), and the dependency on the data stream associated with the current luminance component is illustrated with a double line. Y The dependencies of related data streams are illustrated with dashed lines (e.g., 630), and are related to the current chroma component. C The dependencies between related data flows are illustrated using dotted lines (e.g., 640). Relative to... Figure 4The CU decoding pipeline 400 presented in the example has its encoding tools, which increase decoder complexity, deactivated in this pipeline 600. In this example of a simplified decoding pipeline 600, the encoding configuration is determined, where some encoding tools (e.g., BCW, BDOF, DMVR, LFNST, LMCS, deblocking filter, SAO, and ALF) are deactivated. While latency is reduced when using this decoding pipeline 600, the intra-frame prediction tool requires reconstructing neighboring CUs 620 to predict the current CU. This dependency reduces the parallelism achievable in the decoding pipeline.

[0079] As mentioned above, the encoding of layers (e.g., via...) Figure 5 The layer encoder 530.1 can be based on an encoder that encodes the layers below (e.g., Figure 5 The layer encoder 530.0 provides a reference layer to it. Similarly, the layer is decoded (e.g., via...). Figure 5 The layer decoder 580.1 can be based on the decoder of the layers below it (e.g., Figure 5 The reference layer provided by the layer decoder 580.0. At both ends 510, 570, when inter-layer prediction is enabled, the provided reference layer (illustrated by the dashed arrow) can be used by the corresponding layer encoder to predict the CU. In this case, it is not necessary to use intra-frame prediction tools (e.g., layer encoder 530.1-N and layer decoder 580.1-N) for all layers that may benefit from inter-layer prediction. Therefore, when the coding configuration is determined, further reduction in decoding pipeline dependency is possible, for which intra-frame prediction tools (which have a high impact on the decoding pipeline) are deactivated, such as the reference layer. Figure 7 As further described.

[0080] Figure 7 This is a diagram illustrating a further simplified CU decoding pipeline 700. Figure 7 In the example, the dependency on the data streams originating from the entropy decoder and the image buffer is illustrated with a solid line (e.g., 710), relative to the current luminance component. Y The dependencies of related data streams are illustrated with dashed lines (e.g., 730), and are related to the current chroma component. CThe dependencies of the associated data streams are illustrated using dotted lines (e.g., 740). In this second example of a streamlined CU decoding pipeline 700, it is possible to determine which coding configurations, including intra-prediction tools, are deactivated, resulting in a lower latency pipeline. However, in this case, deactivating intra-prediction tools also allows for parallel decoding of the layer's CUs (because the dependency of CU decoding on the reconstruction of pixels from neighboring CUs is removed). For example, in this coding configuration, TM mode (where the reconstruction neighborhood of the current CU is used to predict coding parameters) can be disabled to avoid dependence on the reconstruction of CUs above or to the left of the current CU.

[0081] Typically, and in accordance with the aspects described herein, to reduce latency and increase decoding parallelism, the encoding configuration for each layer can be determined accordingly. That is, the encoding toolset used can be adapted to each layer. Depending on the requirements of the application of interest, using a layer-adaptive encoding toolset can reduce the overall processing power required by the decoder while maintaining encoding performance.

[0082] In the first aspect, the toolset used to process the enhancement layer can be reduced (and / or constrained). If the base layer must be encoded with high quality, the bandwidth allocated to the enhancement layer can be reduced. In this case, the base layer can be encoded based on a high-performance encoding configuration, activating all tools (or a slightly streamlined toolset). On the other hand, there is no need to use a complex toolset to achieve high compression for the enhancement layer. For example, in spatially scalable use cases, the base layer has a lower resolution, which reduces decoding complexity. However, the enhancement layer can use a streamlined encoding toolset (such as... Figure 7 The toolset shown in the diagram is used for encoding. Furthermore, because the base layer has high quality, it can be used for inter-layer prediction when inter-frame prediction fails, for example, at the refinement layer encoder. Such an encoding scheme can significantly reduce the processing power required by the decoder.

[0083] Using loop filters only on the base layer (or only on one layer) further reduces latency and computational complexity. In variants, the use of loop filters in one or more layers is reduced. For example, loop filters can be used only on some layers, rather than all filters on a particular layer. Thus, a deblocking filter and ALF can be used on the base layer, while only a deblocking filter can be used on the enhancement layer(s); or, SAO and ALF can be used on the base layer, while a deblocking filter can be used on the enhancement layer(s). In this respect, decoding hardware components designed to decode a single layer at a given resolution can be used to decode multi-layer bitstreams with minimal (software) adjustments. The core decoding loop will be capable of decoding both the base layer and the refinement layer, as the former has lower resolution, while the latter is decoded with significantly reduced complexity using a streamlined decoding pipeline.

[0084] In the second aspect, the toolset for processing the base layer can be reduced. This is particularly useful for applications where low transmission latency is required, and therefore encoding and decoding must have very low latency. Besides the low bitrate constraint required for fast transmission, complexity has a significant impact on both the encoder and decoder. In this context, reducing the encoding and decoding complexity of the base layer can be advantageous. Therefore, constraints are applied to the encoding of the base layer in this respect. For example, enhancement layers are used when a more powerful decoder is available or when quality and / or resolution are more important than latency. In low-latency multicast or broadcast applications, some clients only decode the base layer. Other clients, capable of tolerating higher latency, decode the base layer and one or more enhancement layers, and thus can benefit from higher resolution video. In that case, a streamlined encoding toolset (such as...) can be used. Figure 6 The toolset illustrated in the diagram encodes the base layer. Intra-frame prediction can be used on the base layer to maintain sufficient compression performance. However, loop filters can be selectively disabled to reduce latency and complexity. In this regard, the enhancement layer can use a high-performance coding configuration, for example, that enables most tools.

[0085] In the third aspect, the encoding configuration of a layer can be determined so that the group of pictures (GOP) of that layer can be modified—for example, the temporal depth and length of the GOP can be reduced. Generally, longer GOPs with higher temporal depth 1) require more and larger memory buffers to be allocated in the DPB for timing prediction, and 2) result in higher latency because the pictures are not encoded in display order. Therefore, constrained layers may have shorter GOPs and / or GOPs with lower temporal depth (i.e., lower maximum time ID (TID) values) than other layers.

[0086] In this variant, the layer coding configuration can be determined such that inter-layer prediction can be used only for sub-layers of layers with a time ID value below a predetermined threshold (e.g., below the maximum time ID value). In VVC, sub-layers of a layer contain images of the layer at the same time depth. Therefore, in this variant, for example, sub-layers of enhancement layers with higher TID values ​​do not use inter-layer prediction. In such a case, the base layer can have a full frame rate, and sub-layers of the base layer with higher TID values ​​are not used for inter-layer prediction of the enhancement layer (i.e., there is no decoding dependency relative to sub-layers with higher TID values, such as...). Figure 8 As illustrated in the diagram.

[0087] Figure 8 This is an image showing the prediction of 800 between layers. Figure 8 In the example, the images with higher TID values—that is, the images with TID=3 in the base layer (layer 0) (images 4, 5, 7, 8, 12, 13, 15, and 16, in...) Figure 8 (Shown in gray) – These are not used for inter-layer prediction in the refinement layer (Layer 1). However, images with lower TID values ​​– namely, images of the base layer (Layer 0) with TID=0-2 (images 0, 3, 2, 6, 1, 11, 10, 14, and 9) – are used for inter-layer image prediction in the refinement layer (Layer 1), as illustrated by the arrows. Therefore, in Figure 8 In the example, to decode the enhancement layer (those images with TID=0-2), only a portion (half) of the base layer images need to be decoded. Therefore, only a portion of the base layer images need to be decoded before the corresponding images from the refinement layer are decoded, which reduces dependencies and latency.

[0088] In the fourth aspect, the coding configuration of the layer can be determined such that the coding of the residuals is constrained—that is, the use of residual data is limited when reconstructing the coding units of that layer. Constraining the coding of the residuals can reduce the complexity of the decoder. In this case, for example, CUs can be reconstructed based on the corresponding predicted CUs without adding 255 reconstructed residual data to them (see [link to documentation]). Figure 2 Therefore, decoding the CU without reconstructing the corresponding residual data can significantly accelerate the decoding process. In one aspect, constrained residual coding can be achieved by converting the transform coefficients of one or more CUs (e.g., by...) Figure 2 Some settings in the converter 220 (generated) are set to zero to achieve this.

[0089] In the fifth aspect, the encoding configuration of a layer can be determined for each sub-layer, allowing the sub-layer encoding configuration to be determined based on the sub-layer's time ID value. Therefore, constraints can be applied to the encoding of an image of a layer based on the sub-layer from which the image originates. For example, constraints can be applied only to images belonging to layers with the highest TID value. Figure 8 The diagram illustrates four sub-layers (layers 0 and 1) corresponding to four time ID values. The VVC standard only allows decoding of images (within a layer) with TID values ​​lower than or equal to a given TID. This capability provides temporal scalability when using a hierarchical GOP structure. One direct use of temporal scalability is to scale down when insufficient processing power is available (e.g., skipping decoding of images with the highest TID value). This approach results in low-frequency video that may cause scrambling. Therefore, an alternative approach could be to constrain the tools used to encode images with the highest TID values ​​so that they can be decoded with reduced processing power using the constraints described herein. Note that this aspect can also be applied to single-layer video (i.e., classic broadcast scenarios).

[0090] Therefore, in this respect, the encoding configuration of a layer can be determined such that each sublayer of that layer is subject to different constraints, for example, by encoding each sublayer using different toolsets. The constraints on the encoding of sublayers can be applied in a binary manner, i.e., a simplified toolset (or a constrained toolset) is applied when the TID value is above a threshold. Alternatively or in combination, toolsets can be applied incrementally, i.e., a further simplified (or more constrained) toolset is applied as the TID value increases. For example, complex encoding tools can be incrementally and selectively disabled as the TID value of an image increases. The list describing complex tools can be predefined or can be transmitted in the bitstream (e.g., in a VPS, SPS, or Picture Parameter Set (PPS)).

[0091] Therefore, the image of the augmentation layer with the highest time ID value can be constrained. For example, when encoding the image of the augmentation layer with the highest time ID value, the decoding of the residual data can be omitted. Thus, in this example, the CU can be reconstructed based on the corresponding predicted CU. Alternatively, the residual data can be used to reconstruct a finite number of CUs. This reduces the computational complexity of the inverse transformer 350. Constraining the encoding and decoding of the image of the augmentation layer with the highest time ID in this way is less likely to produce perceptible artifacts.

[0092] Constraints on the encoding of multi-layer video and its sub-layers can be signaled in the bitstream to allow the decoder to check for consistency. A new multi-layer profile is described in this paper, which contains constraints applied according to the aspects described herein. A decoder of a given level that conforms to the profile is capable of decoding a bitstream of a given level (and sub-levels) that conforms to the profile.

[0093] This article describes an example of a low-complexity multi-level Master 10 profile. Bitstreams conforming to this profile should also conform to the aforementioned multi-level Master 10 profile. Furthermore, bitstreams conforming to this profile should satisfy the following constraints:

[0094] The consistency of the bitstream with the low-complexity multi-layer master 10 profile is indicated by general_profile_idc equal to 19 (or another new number).

[0095] This document describes an example of a low-latency multi-level Master 10 profile. Bitstreams conforming to this profile should also conform to the aforementioned multi-level Master 10 profile. Furthermore, bitstreams conforming to this profile should satisfy the following constraints:

[0096] The consistency of the bitstream with the low-latency multilayer master 10 profile is indicated by general_profile_idc equal to 21 (or another new number).

[0097] GCI syntax can also be used. To apply constraints to a specific layer, the constraints can be integrated into a GCI syntax element, which is contained within the profile_tier_level structure defined in the SPS referenced by that specific layer.

[0098] Constraints can include, for example, setting existing flags to one: gci_no_cclm_constraint_flag, gci_no_bdof_constraint_flag, gci_no_dmvr_constraint_flag, gci_no_prof_constraint_flag, gci_no_bcw_constraint_flag, gci_no_ciip_constraint_flag, gci_no_lfnst_constraint_flag, gci_no_sao_constraint_flag, gci_no_alf_constraint_flag, and gci_no_lmcs_constraint_flag. Also, setting new flags added to the GCI structure to one: gci_inter_only_constraint_flag and gci_no_deblock_constraint_flag.

[0099] Figure 9 This is a flowchart of an example method 900 for encoding multi-layer video, according to which aspects of this embodiment can be implemented. Method 900 begins in step 910 by obtaining video data (including multi-layer video). In step 920, an encoding configuration is determined to suit the individual layers of the multi-layer video. Then, in step 930, the layers of the multi-layer video are encoded into a bitstream based on their respective encoding configurations. In one aspect, method 900 can be applied to: 1) determining an encoding configuration for a first layer of the multi-layer video, wherein inter-layer prediction is activated to enable the generation of a reference layer during the encoding of the first layer; and 2) determining an encoding configuration for a second layer of the multi-layer video, wherein enabling the use of the reference layer during the encoding of the second layer. Therefore, by using inter-layer prediction to predict coding units, encoding dependencies between coding units in the second layer can be eliminated. Thus, when encoding the second layer, the encoding of coding units in the second layer does not depend on reconstructing pixels from adjacent coding units. Advantageously, in this aspect, the decoding of coding units in the second layer can be implemented in parallel.

[0100] In another aspect, method 900 can be applied to: 1) determining the coding configuration for a first layer of a multi-layer video, defining a first toolset for coding the first layer; and 2) determining the coding configuration for a second layer of the multi-layer video, defining a second toolset for coding the second layer. One of the first and second toolsets is a simplified toolset relative to the other. In one aspect, the first layer may be a base layer, and the second layer may be a refinement layer. For example, the toolset can be simplified by deactivating one or more loop filters or by deactivating the coding of residual data of one or more coding units of the corresponding layer. In another example, method 900 can be applied to constrain the coding configuration of a layer (e.g., relative to the coding configurations of other layers in the multi-layer video) by reducing the temporal depth of the GOP of the layer or by reducing the length of the GOP of the layer.

[0101] A layer in a multi-layer video can contain sub-layers associated with their respective time ID values. In this case, method 900 can be further applied to determine the coding configuration for that layer, including the sub-layer coding configuration determined based on the respective time ID values ​​of the sub-layers. In one aspect, determining the sub-layer coding configuration may include determining a first toolset for a first sub-layer and a second toolset for a second sub-layer, wherein the time ID value of the second sub-layer is higher than that of the first layer, and wherein the second toolset is streamlined relative to the first toolset. In a further aspect, by using inter-layer prediction only for sub-layers of layers with time ID values ​​below a predetermined threshold, method 900 can be applied to constrain the coding configuration of the layer (e.g., the coding configuration relative to other layers in the multi-layer video) (as per [reference to...]). Figure 8 In another aspect, method 900 can be applied to constrain the coding configuration of a layer (e.g., the coding configuration relative to other layers of a multi-layer video) by limiting the use of residual data in the sub-layer reconstruction coding unit of a layer having a time ID value below a predetermined threshold.

[0102] Method 900 can be further configured to signal one or more syntax elements of the encoding configuration of the representation layer (the encoding configuration determined for each layer in step 920) (i.e., encode them into a bitstream). The signaling syntax elements can represent a profile, hierarchy, and level that define the encoding configuration for that layer, as illustrated above with examples of low-complexity multi-level main 10 profiles and low-latency multi-level main 10 profiles. Where the layer includes sub-layers, the profile, hierarchy, and level can further define the encoding configuration for each sub-layer (i.e., the sub-layer encoding configuration).

[0103] Figure 10This is a flowchart of an example method 1000 for decoding multi-layer video, according to which aspects of this embodiment can be implemented. This decoding method 1000 generally operates in reverse to the encoding method 900 described above. Method 1000 begins in step 1010 by obtaining a bitstream encoding video data comprising multi-layer video. In step 1020, method 1000 can be applied to decode an encoding configuration from the bitstream, which is determined (in step 920 of method 900) to suit the individual layers of the multi-layer video. Then, in step 1030, method 1000 can be applied to decode the layers from the bitstream based on their respective encoding configurations. Decoding the layer's encoding configuration (in step 1020) may include decoding one or more syntax elements representing the encoding configuration (signaled in the bitstream). The signaled syntax elements represent a profile, hierarchy, and level defining the encoding configuration, as illustrated above with examples of low-complexity multi-layer master 10 profiles and low-latency multi-layer master 10 profiles. In cases where a middle layer includes sub-layers, the profile, hierarchy, and level can further define the encoding configuration for each of the sub-layers (i.e., the sub-layer encoding configuration).

[0104] Several aspects and embodiments have been described in this disclosure. These aspects and embodiments provide at least the following outputs and results, including all combinations, across different claim classes and types: • Based on any of the aspects described herein, syntax elements that would enable the decoder to decode encoded video data would be encoded into the encoded video data.

[0105] • Includes a bitstream of one or more of the described syntax elements or their variants. The bitstream can be any dataset, whether it is transmitted, stored, or otherwise available.

[0106] • Create, transmit, receive, and / or decode bitstreams.

[0107] • An electronic device (e.g., a TV, set-top box, cellular phone, or tablet computer) tunes (e.g., using a tuner) a channel to receive a bitstream or receives a bitstream over the air (e.g., using an antenna). The electronic device decodes the syntax elements from the bitstream and optionally displays (e.g., using a monitor, screen, or any other type of display) the resulting image.

[0108] Throughout this disclosure, various other general and specific outputs, results, implementations, and claims are also supported and considered.

[0109] Various methods are described herein, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., e.g., "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a sequence of modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or within a time period overlapping with the second decoding.

[0110] The various methods and other aspects described in this application can be used to modify the module, for example, such as Figure 2 and Figure 3 The modules of the video encoder 200 and video decoder 300 shown are illustrated. Furthermore, this aspect is not limited to specific standards (such as VVC or HEVC) and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.

[0111] Various numerical values ​​are used in this application. Specific values ​​are used for illustrative purposes, and the aspects described are not limited to these specific values.

[0112] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process, such as performing a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more procedures typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase “decoding process” is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.

[0113] Various implementations involve encoding. In a manner similar to the discussion above regarding "decoding," the term "encoding," as used in this application, can encompass all or part of the process performed on input video data to produce an encoded bitstream. Furthermore, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "encoded" and "encoded," and the terms "image," "picture," and "frame." Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while the term "decoding" is used on the decoder side.

[0114] Note that the grammatical elements used in this article are descriptive terms. Therefore, they do not preclude the use of other grammatical element names.

[0115] This disclosure has described various fragments of information that can be transmitted or stored, such as, for example, syntax. This information can be packaged or arranged in a variety of ways, including those common in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other methods are also available, including those common in system-level or application-level standards, such as signaling the information in one or more of the following: a. SDP (Session Description Protocol) is used to describe the format of a multimedia communication session for the purpose of session announcement and session invitation, such as as described in RFCs and used in conjunction with RTP (Real-Time Transport Protocol) transmission.

[0116] b. DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and transmitted via HTTP. Descriptors are associated with a representation or set of representations to provide additional characteristics to the content representation.

[0117] c. RTP header extensions, such as those used during RTP streaming.

[0118] d. ISO basic media file format, such as the boxes used in OMAF, which are object-oriented building blocks (also called "atoms" in some specifications) defined by unique type identifiers and lengths.

[0119] e. An HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can be associated, for example, with a version of the content or a collection of versions to provide characteristics of the version or collection of versions.

[0120] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in an apparatus, such as a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (PDAs), and other devices that facilitate the transfer of information between end users.

[0121] References to “an aspect,” “an embodiment,” or “an implementation,” and their variations, mean that a particular feature, structure, characteristic, etc., described in connection with that aspect, embodiment, or implementation is included in at least one embodiment. Therefore, the phrases “in an aspect,” “in an embodiment,” or “in an implementation,” and any other variations appearing throughout this application, do not necessarily refer to the same embodiment.

[0122] Furthermore, this application may involve "determining" fragments of various information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory.

[0123] Furthermore, this application may relate to “accessing” fragments of various information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0124] Furthermore, this application may relate to "receiving" fragments of various information. As with "access," receiving is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., retrieving information from memory). Moreover, "receiving" is generally referred to in one or more ways during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0125] To be understood, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be clear to those skilled in the art and related fields, this can be extended to as many entries as possible listed.

[0126] Furthermore, among other things, as used herein, the term "signaling" also refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals the quantization parameters used for dequantization. Thus, in embodiments, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicitly signaling) to allow only the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual data. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, information is signaled to the corresponding decoder using one or more syntax elements, flags, etc. Although the verb form of the term "signaling" has been referred to above, the word "signal" can also be used as a noun herein.

[0127] As will be apparent to those skilled in the art, implementations can generate various signals that are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

Claims

1. A method comprising: Obtain video data, including multi-layered video; Determine the encoding configuration for each layer of a multi-layered video; as well as The layers are encoded into a bitstream based on their respective encoding configurations.

2. The method of claim 1, wherein the layer comprises a first layer and a second layer, and wherein determining the encoding configuration comprises: Determine the coding configuration for the first layer, where inter-layer prediction is activated to enable the generation of a reference layer during the coding of the first layer; as well as Determine the encoding configuration for the second layer, which enables the use of the reference layer during the encoding of the second layer.

3. The method according to claim 2, wherein the encoding of the layer further comprises: The second layer is encoded, where the encoding of the encoding units in the second layer does not depend on the reconstruction of pixels from adjacent encoding units.

4. The method of claim 1, wherein the layer comprises a first layer and a second layer, and wherein determining the encoding configuration comprises: Determine the encoding configuration for the first layer, which defines a first toolset for encoding the first layer; as well as Determine the encoding configuration for the second layer, which defines a second toolset for encoding the second layer. Among them, one of the first and second toolsets is a streamlined toolset relative to the other of the first and second toolsets.

5. The method according to claim 4, wherein the first layer is a base layer and the second layer is a refinement layer.

6. The method of claim 4, wherein, in a streamlined toolset, one or more loop filters are deactivated.

7. The method of claim 4, wherein in the simplified toolset, the encoding of one or more residual data in the encoding units of the corresponding layer is deactivated.

8. The method of claim 1, further comprising: The encoding configuration of the layers in a multi-layer video is constrained relative to the encoding configuration of other layers in the multi-layer video by reducing the temporal depth of the Group of Pictures (GOP) of the layers.

9. The method of claim 1, further comprising: The encoding configuration of the layers in a multi-layer video is constrained relative to the encoding configuration of other layers in the multi-layer video by reducing the length of the Group of Pictures (GOP) of each layer.

10. The method of any one of claims 1 to 9, wherein each layer of the multi-layer video comprises a sub-layer associated with a respective time ID value, and wherein determining the encoding configuration includes: The corresponding sub-layer encoding configuration for a given layer is determined based on the respective time ID value of each sub-layer.

11. The method of claim 10, wherein determining the sub-layer coding configuration further comprises: Determine a first toolset for the first sub-layer and a second toolset for the second sub-layer, wherein the time ID value of the second sub-layer is higher than the time ID value of the first sub-layer, and wherein the second toolset is simplified relative to the first toolset.

12. The method of any one of claims 1 to 11, wherein each layer of the multi-layer video comprises a sub-layer associated with a respective time ID value, and wherein determining the encoding configuration includes: The coding configuration of a layer is constrained by using inter-layer prediction only for sub-layers of layers with time ID values ​​below a predetermined threshold.

13. The method of any one of claims 1 to 12, wherein each layer of the multi-layer video comprises a sub-layer associated with a respective time ID value, and wherein determining the encoding configuration includes: The coding configuration of a layer is constrained by limiting the use of residual data in the reconstruction coding unit of a sublayer for a layer with a time ID value below a predetermined threshold.

14. A method comprising: Obtain the bitstream of video data encoded with multiple video layers; Decoding from the bitstream is determined to fit the encoding configuration of each layer of the multi-layer video; as well as The layer is decoded from the bitstream based on its respective encoding configuration.

15. The method of claim 14, wherein a layer of a multi-layer video comprises a sub-layer associated with a respective time ID value, and wherein the encoding configuration includes a sub-layer encoding configuration determined for the sub-layers of the layer based on the respective time ID values ​​of the sub-layers.

16. The method according to claim 14 or 15, further comprising: Decode one or more syntax elements from the bitstream that represent the encoding configuration in the encoding configuration, wherein one or more syntax elements represent the profile, hierarchy, and level that define the encoding configuration.

17. The method of claim 16, wherein the encoding configuration includes a sub-layer encoding configuration, and wherein the profile, hierarchy, and level further define the sub-layer encoding configuration.

18. An apparatus comprising: At least one processor; as well as A memory for storing instructions, which, when executed by at least one processor, cause the device to: Obtain video data, including multi-layered video; Determine the encoding configuration for each layer of a multi-layered video; and The layers are encoded into a bitstream based on their respective encoding configurations.

19. The apparatus of claim 18, wherein the layer comprises a first layer and a second layer, and wherein determining the encoding configuration comprises: Determine the coding configuration for the first layer, where inter-layer prediction is activated to enable the generation of a reference layer during the coding of the first layer; as well as Determine the encoding configuration for the second layer, which enables the use of the reference layer during the encoding of the second layer.

20. The apparatus of claim 18, wherein the layer comprises a first layer and a second layer, and wherein determining the encoding configuration comprises: Determine the encoding configuration for the first layer, which defines a first toolset for encoding the first layer; as well as Determine the encoding configuration for the second layer, which defines a second toolset for encoding the second layer. Among them, one of the first and second toolsets is a streamlined toolset relative to the other of the first and second toolsets.

21. An apparatus comprising: At least one processor; as well as A memory for storing instructions, which, when executed by at least one processor, cause the device to: Obtain the bitstream of video data encoded with multiple video layers; Decoding from the bitstream is determined to fit the encoding configuration of each layer of the multi-layered video; and The layer is decoded from the bitstream based on its respective encoding configuration.

22. The apparatus of claim 21, wherein a layer of multi-layer video comprises a sub-layer associated with a respective time ID value, and wherein the encoding configuration includes a sub-layer encoding configuration determined for the sub-layers of the layer based on the respective time ID values ​​of the sub-layers.

23. The apparatus according to claim 21 or 22, further comprising: Decode one or more syntax elements from the bitstream that represent the encoding configuration in the encoding configuration, wherein one or more syntax elements represent the profile, hierarchy, and level that define the encoding configuration.

24. The apparatus of claim 23, wherein the encoding configuration includes a sub-layer encoding configuration, and wherein the profile, hierarchy, and level further define the sub-layer encoding configuration.

25. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method comprising: Obtain video data, including multi-layered video; Determine the encoding configuration for each layer of a multi-layered video; as well as The layers are encoded into a bitstream based on their respective encoding configurations.

26. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method comprising: Obtain the bitstream of video data encoded with multiple video layers; Decoding from the bitstream is determined to fit the encoding configuration of each layer of the multi-layer video; as well as The layer is decoded from the bitstream based on its respective encoding configuration.