Apparatus, method and computer program for video encoding and decoding
By simplifying the tile and brick partition signaling in the video encoding standard, the problems of high bit rate and low encoding efficiency caused by complex signaling structure are solved, and a more efficient encoding process is achieved.
Patent Information
- Application Number
- CN202080040928.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-03
- Filing Date
- 2020-05-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-05-29
AI Technical Summary
In the existing video encoding standards, the signaling structure of tiles and brick partitions is complex, resulting in high bit rates and low encoding efficiency.
Partition signaling is simplified by determining the number of partitions and number of cells to be assigned and indicating or inferring the size and boundaries of partitions in a predefined scan sequence until all cells are assigned.
The bit rate required for signaling is reduced, encoding efficiency is improved, and partition signaling structure is simplified.
Smart Images

Figure CN113940084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to apparatus, methods and computer programs for video encoding and decoding. Background Art
[0002] Video coding standards and specifications generally allow an encoder to divide or partition an encoded picture into subsets. In video coding, a partition can be defined as dividing a picture or a sub-region of a picture into subsets (blocks) such that each element of the picture or the sub-region of the picture is exactly in one of the subsets (blocks). For example, H.265 / HEVC introduced the concept of coding tree units (CTUs) with a default size of 64×64 pixels. Based on a quadtree structure, a CTU can contain a single coding unit (CU), or can be recursively split into multiple smaller CUs, at least 8x8 pixels. H.265 / HEVC also recognizes tiles (which are rectangular and contain an integer number of CTUs) and slices (defined based on slice segments that contain an integer number of coding tree units, are sorted consecutively in a tile scan, and are contained in a single NAL unit).
[0003] Versatile Video Coding (VVC) (MPEG-I Part 3) (also known as ITU-T H.266) is a video compression standard (formal ISO / IEC JTC1 SC29 WG11) developed by the Joint Video Team (JVET) of the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU), which is the successor of HEVC / H.265. The VVC partitioning scheme includes not only tiles, but also bricks, which can include one or more CTU rows within a tile. The introduction of bricks also affects the definition of slices.
[0004] Therefore, a rather complex syntax structure has been created for signaling various options for tile and brick partitioning, which is sub-optimal in many respects, especially with respect to the bitrate required for such signaling. Summary of the Invention
[0005] Now, in order to at least mitigate the above problems, an enhanced coding method is introduced herein.
[0006] The scope of protection sought by the various embodiments of the present invention is set out in the independent claims. Embodiments and features described in this specification that do not fall within the scope of the independent claims (if any) will be construed as examples useful for understanding the various embodiments of the present invention.
[0007] The method according to the first aspect includes: determining the number of units to be assigned to partitions and initially set to unassigned; indicating or inferring the number of partitions of a specific size to be assigned; indicating the size of the partitions of the specific size, and accordingly marking the unassigned units as being assigned to the partitions in a predefined scan order; indicating the count of the units; repeatedly assigning the count of the units to the partitions, and accordingly marking the unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the count of the units; and if the number of unassigned units is greater than 0, assigning the unassigned units to the last partition.
[0008] According to an embodiment, the partition is one or more of the following: a tile column, a tile row, a brick row of one or more tile columns, a brick row of tiles, a grid column of a grid for indicating sub-graph partitioning, a grid row of a grid for indicating sub-graph partitioning.
[0009] According to an embodiment, the unit is one or more of the following: a rectangular block of samples of an image, a grid cell of a grid for indicating sub-graph partitioning.
[0010] The apparatus according to the second aspect includes: means for determining the number of units to be assigned to partitions and initially set to unassigned; means for indicating or inferring the number of partitions of a specific size to be assigned; means for indicating the size of the partitions of the specific size, and means for accordingly marking the unassigned units as being assigned to the partitions in a predefined scan order; means for indicating the count of the units; means for repeatedly assigning the count of the units to the partitions, and means for accordingly marking the unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the count of the units; and means for assigning the unassigned units to the last partition if the number of unassigned units is greater than 0.
[0011] The method according to the third aspect includes: determining the number of units to be assigned to partitions; determining the number of partitions of a specific size to be assigned; determining the size of the partitions of the specific size, and accordingly marking the unassigned units as being assigned to the partitions in a predefined scan order; determining the count of the units; repeatedly assigning the count of the units to the partitions, and accordingly marking the unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the count of the units; and if the number of unassigned units is greater than 0, assigning the unassigned units to the last partition.
[0012] According to an embodiment, determining the number of partitions of a specific size to be assigned includes decoding the number of partitions of the specific size from the syntactic structure; determining the size of the partitions of the specific size includes decoding the size of the partitions of the specific size from the syntactic structure; and determining the count of the units includes decoding the count of the units from the syntactic structure.
[0013] The apparatus according to the fourth aspect includes: a component for determining the number of units to be assigned to partitions; determining the number of partitions of a definite size to be assigned; a component for determining the size of the partitions of the definite size, and a component for correspondingly marking the unassigned units as being assigned to the partitions in a predefined scan order; a component for determining the count of the units; a component for repeatedly assigning the count of the units to the partitions, and a component for correspondingly marking the unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the count of the units; and a component for assigning the unassigned units to the last partition if the number of unassigned units is greater than 0.
[0014] Other aspects relate to an apparatus including: at least one processor and at least one memory, the at least one memory storing thereon code which, when executed by the at least one processor, causes the apparatus to perform at least one or more embodiments of the above methods and related embodiments thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To better understand the present invention, reference will now be made, by way of example, to the accompanying drawings, in which:
[0016] Figure 1 An electronic device adopting an embodiment of the present invention is schematically shown;
[0017] Figure 2 A user equipment suitable for adopting an embodiment of the present invention is schematically shown;
[0018] Figure 3 An electronic device adopting an embodiment of the present invention is also schematically shown, and the present invention is connected using wireless and wired network connections;
[0019] Figure 4 An encoder suitable for implementing an embodiment of the present invention is schematically shown;
[0020] Figure 5a 、 Figure 5b 、 Figure 5c Some examples of partitioning a picture into coding tree units (CTUs), tiles, bricks, and slices are shown;
[0021] Figure 6 The syntax structure of the signaling of slice, picture, and brick partitioning according to H.266 / VVC Draft 5 is shown;
[0022] Figure 7 A flowchart of an encoding method according to one aspect of the present invention is shown;
[0023] Figure 8Shows a flowchart of an encoding method according to another aspect of the present invention;
[0024] Figure 9 Shows a flowchart of an encoding method according to an embodiment of the present invention;
[0025] Figure 10 a, Figure 10 b, Figure 10 c shows some examples of tile and brick partitioning;
[0026] Figure 11 Shows a schematic diagram of a decoder suitable for implementing an embodiment of the present invention;
[0027] Figure 12 Shows a flowchart of a decoding method according to an embodiment of the present invention;
[0028] Figure 13 Shows a flowchart of a decoding method according to another embodiment of the present invention;
[0029] Figure 14a and Figure 14b Shows a flowchart of an encoding and decoding method according to yet another embodiment of the present invention; and
[0030] Figure 15 Shows a schematic diagram of an example multimedia communication system in which various embodiments can be implemented. Detailed Description of the Invention
[0031] The following further describes in detail the suitable devices and possible mechanisms for initiating a viewpoint switch. In this regard, first refer to Figure 1 and Figure 2 , in which Figure 1 Shows a block diagram of a video encoding system according to an example embodiment as a schematic block diagram of an exemplary device or electronic device 50, which may include a codec according to an embodiment of the present invention. Figure 2 Shows the layout of the device according to the example embodiment. Figure 1 and Figure 2 The elements of will be explained next.
[0032] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system. However, it should be understood that embodiments of the present invention can be implemented within any electronic device or apparatus that may require encoding and decoding or encoding or decoding of video images.
[0033] The device 50 may include a housing 30 for containing and protecting the equipment. The device 50 may also include a display 32 in the form of a liquid crystal display. In other embodiments of the present invention, the display may be any suitable display technology suitable for displaying images or videos. The device 50 may also include a keypad 34. In other embodiments of the present invention, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0034] The device may include a microphone 36 or any suitable audio input that may be a digital or analog signal input. The device 50 may also include an audio output device, which in embodiments of the present invention may be an earpiece 38, a speaker, or any one of analog audio or digital audio output connections. The device 50 may also include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device such as a solar cell, a fuel cell, or a wind-up generator). The device may also include a camera capable of recording or capturing images and / or videos. The device 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may also include any suitable short-range communication solution, such as for example a Bluetooth wireless connection or a USB / FireWire wired connection.
[0035] The device 50 may include a controller 56, a processor, or processor circuitry for controlling the device 50. The controller 56 may be connected to a memory 58, which in embodiments of the present invention may store data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may also be connected to codec circuitry 54, which is suitable for performing encoding and decoding of audio and / or video data or assisting in the encoding and decoding performed by the controller.
[0036] The device 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC reader, to provide user information and suitable for providing authentication information for authenticating and authorizing the user at the network.
[0037] The device 50 may include radio interface circuitry 52, which is connected to the controller and suitable for generating wireless communication signals for communicating with, for example, a cellular communication network, a wireless communication system, or a wireless local area network. The device 50 may also include an antenna 44, which is connected to the radio interface circuitry 52 to transmit radio frequency signals generated at the radio interface circuitry 52 to (a) other device(s), and receive radio frequency signals from (a) other device(s).
[0038] Device 50 may include a camera capable of recording or detecting individual frames, which are then passed to codec 54 or a controller for processing. The device may receive video image data from another device for processing before transmission and / or storage. Device 50 may also receive images wirelessly or via a wired connection for encoding and decoding. The structural elements of device 50 described above represent examples of components for performing corresponding functions.
[0039] With respect to Figure 3 , examples of systems in which embodiments of the present invention can be used are shown. System 10 includes a plurality of communication devices capable of communicating via one or more networks. System 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, CDMA networks, etc.), wireless local area networks (WLAN) defined by any IEEE802.x standard, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.
[0040] System 10 may include wired and wireless communication devices and / or device 50 suitable for implementing embodiments of the present invention.
[0041] For example, Figure 3 the system shown depicts a representation of a mobile telephone network 11 and the Internet 28. Connectivity to the Internet 28 may include but is not limited to long-range wireless connections, short-range wireless connections, and various wired connections, including but not limited to telephone lines, cable lines, power lines, and similar communication paths.
[0042] Example communication devices shown in system 10 may include but are not limited to a combination of an electronic device or device 50, a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, and a laptop computer 22. When carried by a mobile individual, device 50 may be stationary or mobile. Device 50 may also be positioned in a transportation mode, including but not limited to an automobile, a truck, a taxi, a bus, a train, a ship, an airplane, a bicycle, a motorcycle, or any similar suitable transportation mode.
[0043] Embodiments may also be implemented in a set-top box (i.e., a digital TV receiver that may / may not have a display or wireless capabilities), in a tablet computer or a (laptop) personal computer (PC) (with a combination of hardware or software or encoder / decoder implementations), in various operating systems, and in a chipset, a processor, a DSP, and / or an embedded system that provides hardware / software-based encoding.
[0044] Some or other device can send and receive calls and messages and communicate with a service provider via a wireless connection 25 to a base station 24. The base station 24 can be connected to a network server 26 that allows communication between a mobile telephone network 11 and the Internet 28. The system can include additional communication devices and various types of communication devices.
[0045] Communication devices can communicate using various transmission technologies, including but not limited to Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol - Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Messaging Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, and any similar wireless communication technology. Communication devices involved in various embodiments of the present invention can communicate using various media, including but not limited to radio, infrared, laser, cable connection, and any suitable connection.
[0046] In telecommunication and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection on a multiplexed medium capable of carrying multiple logical channels. A channel can be used to convey an information signal (e.g., a bit stream) from one or more transmitters (or transmitters) to one or more receivers.
[0047] The MPEG-2 Transport Stream (TS), specified equivalently in ISO / IEC 13818-1 or in ITU-T Recommendation H.222.0, is a format for carrying audio, video, and other media, as well as program metadata or other metadata, in a multiplexed stream. Packet Identifiers (PIDs) are used to identify elementary streams (also known as packetized elementary streams) within the TS. Thus, logical channels within the MPEG-2 TS can be considered to correspond to specific PID values.
[0048] Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOBMFF) and the file format for NAL unit structured video (ISO / IEC 14496-15), which is derived from ISOBMFF.
[0049] A video codec consists of an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a visual form. The video encoder and / or the video decoder can also be separated from each other, i.e., it is not necessary to form a codec. Typically, the encoder discards some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate).
[0050] Typical hybrid video encoders (e.g., many encoder implementations of ITU-T H.263 and H.264) encode video information in two stages. First, pixel values in a certain picture region (or “block”) are predicted, e.g., by motion compensation (finding and indicating a region in a previously encoded video frame that closely corresponds to the block being encoded) or spatially (using pixel values around the block to be encoded in a specified manner). Second, the prediction error (i.e., the difference between the predicted pixel block and the original pixel block) is encoded. This is typically done by transforming the difference in pixel values using a specified transform (such as the discrete cosine transform (DCT) or a variant thereof), quantifying the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bitrate).
[0051] In temporal prediction, the prediction source is a previously decoded picture (also known as a reference picture). In intra-block copy (IBC; also known as intra-block copy prediction), the prediction is applied similarly to temporal prediction, but the reference picture is the current picture, and only previously decoded samples can be referenced during the prediction process. Inter-layer or inter-view prediction can be applied similarly to temporal prediction, but the reference pictures are decoded pictures from another scalable layer or from another view, respectively. In some cases, inter-frame prediction can refer only to temporal prediction, while in other cases, inter-frame prediction can collectively refer to temporal prediction and any one of intra-block copy, inter-layer prediction, and inter-view prediction, provided that they are performed in the same or a similar process as temporal prediction. Inter-frame prediction or temporal prediction is sometimes referred to as motion compensation or motion-compensated prediction.
[0052] Inter-frame prediction (which can also be referred to as temporal prediction, motion compensation, or motion-compensated prediction) reduces temporal redundancy. In inter-frame prediction, the prediction source is a previously decoded picture. Intra-frame prediction exploits the fact that adjacent pixels within the same picture are likely to be correlated. Intra-frame prediction can be performed in the spatial or transform domain, i.e., sample values or transform coefficients can be predicted. Intra-frame prediction is typically exploited in intra-frame coding where inter-frame prediction is not applied.
[0053] One result of the encoding process is a set of encoding parameters, such as motion vectors and quantized transform coefficients. If many parameters are first predicted from spatially or temporally neighboring parameters, many parameters can be entropy encoded more efficiently. For example, a motion vector can be predicted from a spatially adjacent motion vector, and only the difference relative to the motion vector predictor can be encoded. The prediction of encoding parameters and intra prediction can be collectively referred to as intra-picture prediction.
[0054] Figure 4 FIG. shows a block diagram of a video encoder suitable for adopting an embodiment of the present invention. Figure 4 An encoder for two layers is proposed, but it should be understood that the proposed encoder can be similarly extended to encode more than two layers. Figure 4 FIG. illustrates an embodiment of a video encoder, which includes a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 may include similar elements for encoding an incoming picture. The encoder sections 500, 502 may include pixel predictors 302, 402, prediction error encoders 303, 403, and prediction error decoders 304, 404. Figure 4 Embodiments of the pixel predictors 302, 402 are also shown, which include an inter-frame predictor 306, 406, an intra-frame predictor 308, 408, a mode selector 310, 410, filters 316, 416, and reference frame memories 318, 418. The pixel predictor 302 of the first encoder section 500 receives 300 base layer images of the video stream for encoding at the inter-frame predictor 306 (determining the difference between the image and the motion compensated reference frame 318) and the intra-frame predictor 308 (determining the prediction of an image block based only on the processed part of the current frame or picture). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 310. The intra-frame predictor 308 may have more than one intra-frame prediction mode. Thus, each mode may perform intra-frame prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer picture 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives 400 enhancement layer images of the video stream for encoding at both the inter-frame predictor 406 (determining the difference between the image and the motion compensated reference frame 418) and the intra-frame predictor 408 (determining the prediction of an image block based only on the processed part of the current frame or picture). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 410. The intra-frame predictor 408 may have more than one intra-frame prediction mode. Thus, each mode may perform intra-frame prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer picture 400.
[0055] Depending on which coding mode is selected to encode the current block, the output of the inter-frame predictors 306, 406, or the output of one of the optional intra-frame predictor modes, or the output of the surface encoder within the mode selector, is passed to the output of the mode selectors 310, 410. The output of the mode selector is passed to the first summing device 321, 421. The first summing device may subtract the output of the pixel predictors 302, 402 from the base layer picture 300 / enhancement layer picture 400 to generate a first prediction error signal 320, 420 that is input to the prediction error encoders 303, 403.
[0056] The pixel predictors 302, 402 also receive a combination of the predicted representation of the image blocks 312, 412 and the outputs 338, 438 of the prediction error decoders 304, 404 from the preliminary reconstructors 339, 439. The preliminary reconstructed images 314, 414 may be passed to the intra-frame predictors 308, 408 and the filters 316, 416. The filters 316, 416 that receive the preliminary representation may filter the preliminary representation, and the output may be stored as the final reconstructed images 340, 440 in the reference frame memories 318, 418. The reference frame memory 318 may be connected to the inter-frame predictor 306 to be used as a reference image for comparison with future base layer pictures 300 in inter-frame prediction operations. According to some embodiments, in the case where the base layer is selected and indicated as the source for inter-layer sample prediction and / or inter-layer motion information prediction for the enhancement layer, the reference frame memory 318 may also be connected to the inter-frame predictor 406 to be used as a reference image for comparison with future enhancement layer pictures 400 in inter-frame prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-frame predictor 406 to be used as a reference image for comparison with future enhancement layer pictures 400 in inter-frame prediction operations.
[0057] According to some embodiments, in the case where the base layer is selected and indicated as the source for predicting the filtering parameters of the enhancement layer, the filtering parameters of the filter 316 from the first encoder section 500 may be provided to the second encoder section 502.
[0058] The prediction error encoders 303, 403 include transform units 342, 442 and quantizers 344, 444. The transform units 342, 442 transform the first prediction error signals 320, 420 into the transform domain. The transform is, for example, a DCT transform. The quantizers 344, 444 quantize the transform domain signals (e.g., DCT coefficients) to form quantized coefficients.
[0059] The prediction error decoders 304, 404 receive the outputs from the prediction error encoders 303, 403, and perform the reverse processes of the prediction error encoders 303, 403 to generate decoded prediction error signals 338, 438, which, when combined with the predicted representations of the image blocks 312, 412 at the second summing devices 339, 439, generate the preliminary reconstructed images 314, 414. The prediction error decoders can be considered to include dequantizers 361, 461 that dequantize quantization coefficient values (e.g., DCT coefficients) to reconstruct the transform signals and inverse transform units 363, 463 that perform an inverse transform on the reconstructed transform signals, where the output of the inverse transform units 363, 463 contains the (multiple) reconstructed blocks. The prediction error decoders can further include block filters that can filter the (multiple) reconstructed blocks based on further decoded information and filter parameters.
[0060] The entropy encoders 330, 430 receive the outputs from the prediction error encoders 303, 403, and can perform appropriate entropy coding / variable length coding on the signals to provide error detection and correction capabilities. The outputs of the entropy encoders 330, 430 can be inserted into the bitstream, for example, via the multiplexer 508.
[0061] Entropy coding / decoding can be performed in many ways. For example, context-based coding / decoding can be applied, where both the encoder and the decoder modify the context state of the coding parameters based on previously coded / decoded coding parameters. Context-based coding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC) or any similar entropy coding. Entropy coding / decoding can alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding / decoding or Exp-Golomb coding / decoding, etc. The decoding of the coding parameters from the bitstream or codewords of the entropy coding can be referred to as parsing.
[0062] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard was published by two parent standardization organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There are multiple versions of the H.264 / AVC standard, with new extensions or features integrated in the specification. These extensions include scalable video coding (SVC) and multi-view video coding (MVC).
[0063] Version 1 of the High Efficiency Video Coding (H.265 / HEVC, also known as HEVC) standard was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of VCEG and MPEG. This standard was published by two parent standardization organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Later versions of H.265 / HEVC include scalable, multi-view, fidelity range, 3D, and screen content coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0064] SHVC, MV-HEVC, and 3D-HEVC use the common base specifications specified in Annex F of Version 2 of the HEVC standard. This common base includes, for example, advanced syntax and semantics (such as inter-layer dependencies) that specify some characteristics of each layer of the bitstream, as well as decoding processes, such as reference picture list construction, including inter-layer reference pictures and picture order count derivation for multi-layer bitstreams. Annex F can also be used in potential subsequent multi-layer extensions of HEVC. It is to be understood that even though video encoders, video decoders, encoding methods, decoding methods, bitstream structures, and / or embodiments may be described below with reference to specific extensions such as SHVC and / or MV-HEVC, they generally apply to any multi-layer extension of HEVC and, even more generally, to any multi-layer video coding scheme.
[0065] Versatile Video Coding (VVC) (MPEG-I Part 3) (also known as ITU H.266) is a video compression standard developed by the Joint Video Exploration Team (JVET) of the MPEG consortium and ITU as the successor to HEVC / H.265.
[0066] Some key definitions, bitstreams, and coding structures, as well as concepts and some of their extensions of H.264 / AVC, HEVC, and VVC are described in this chapter as examples of video encoders, decoders, encoding methods, decoding methods, and bitstream structures, where embodiments can be implemented. The various aspects of the present invention are not limited to H.264 / AVC or HEVC or VCC or their extensions. Instead, the description is given for a possible basis on which the embodiments can be implemented partially or fully. Whenever VVC or any draft version thereof is cited below, it is to be understood that the description matches the VVC draft specification, and later draft versions and the (multiple) final versions of VVC may change, and the description and embodiments can be adjusted to match the (multiple) final versions of VVC.
[0067] Video coding standards can specify the bitstream syntax and semantics as well as the decoding process of an error-free bitstream, while the encoding process may not be specified, but the encoder may only need to generate a compliant bitstream. Bitstream and decoder compliance can be verified using a Hypothetical Reference Decoder (HRD). These standards can include coding tools that help handle transmission errors and losses, but the use of such tools in encoding can be optional, and the decoding process for an erroneous bitstream may not have been specified.
[0068] Syntax elements can be defined as data elements represented in the bitstream. A syntax structure can be defined as zero or more syntax elements that appear together in the bitstream in a specified order.
[0069] Each syntax element can be described by its name and a descriptor for its encoding representation method. A convention where the syntax element name consists of all lowercase letters with an underscore character can be used. The decoding process of a video decoder can behave according to the value of the syntax element and the values of previously decoded syntax elements.
[0070] When describing H.264 / AVC, HEVC, VCC, and example embodiments, the following descriptors and / or descriptions can be used to specify the parsing process of each syntax element.
[0071] -u(n): Unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies in a way that depends on the value of other syntax elements. The pairing process for this descriptor is specified by the next n bits from the bitstream, interpreted as the binary representation of an unsigned integer, where the most significant bit is written first.
[0072] -ue(v): Unsigned integer exponential Golomb coding (also known as exp Golomb coding) syntax element, with the left bit first.
[0073] The exponential Golomb bitstring can be converted to a code number (codeNum), for example using the following table:
[0074] bit string codeNum 1 0 0 1 0 1 0 1 1 2 0 0 1 0 0 3 0 0 1 0 1 4 0 0 1 1 0 5 0 0 1 1 1 6 0 0 0 1 0 0 0 7 0 0 0 1 0 0 1 8 0 0 0 1 0 1 0 9 … …
[0075] In some cases, the syntax table can use the values of other variables derived from the syntax element values. A variable naming convention can be used where lowercase and uppercase letters are mixed and there are no underscore characters. Variables starting with an uppercase letter can be derived for use in decoding the current syntax structure and all related syntax structures. Variables starting with an uppercase letter can be used in later syntax structures during the decoding process without referring to the original syntax structure of the variable. A convention where variables starting with a lowercase letter can only be used within the context in which they are derived can be used.
[0076] In some cases, the "mnemonic" name of a syntax element value or variable value can be used interchangeably with its numerical value. Sometimes the "mnemonic" name is used without any associated numerical value.
[0077] A flag can be defined as a variable or a single-bit syntax element that can take on one of two possible values: 0 and 1.
[0078] An array can be a syntax element or a variable. Square brackets can be used for array indexing. A one-dimensional array can be called a list. A two-dimensional array can be called a matrix.
[0079] A function can be described by its name. A convention can be used where the function name starts with an uppercase letter, contains a mix of lowercase and uppercase letters without any underscore characters, and ends with a left parenthesis and a right parenthesis, including zero or more variable names (for definition) or values (for use) separated by commas (if there are more than one variable).
[0080] The function Ceil(x) can be defined as returning the smallest integer greater than or equal to x. The function Log2(x) can be defined as returning the base-2 logarithm of x.
[0081] The process for describing the decoding of syntax elements can be specified. The process can have separate specifications and calls. It can be specified that all syntax elements and uppercase variables belonging to the current syntax structure and related syntax structures are available in the process specification and call, and it can also be specified that the process specification can also have lowercase variables explicitly specified as inputs. Each process specification can explicitly specify one or more outputs, and each output can be a variable that can be an uppercase variable or a lowercase variable.
[0082] Syntax, semantics, and processes can be described using arithmetic, logical, relational, bitwise, and assignment operators similar to those used in the C programming language. Specifically, the operator / is used to indicate integer division (with truncation), and the operator % is used to indicate the modulus (i.e., the remainder of the division).
[0083] Numbering and counting conventions can start from 0, e.g., "the first" corresponds to the 0th, "the second" corresponds to the 1st, etc.
[0084] The basic units for encoder input and decoder output are usually pictures. The picture given as input to the encoder can also be called the source picture, and the picture decoded through decoding can be called the decoded picture or the reconstructed picture.
[0085] The source picture and the decoded picture are each composed of one or more arrays of samples, such as one of the sets of sample arrays below:
[0086] - Luminance (Y) only (monochrome).
[0087] - Luminance and two chrominances (YCbCr or YCgCo).
[0088] - Green, blue, and red (GBR, also known as RGB).
[0089] - An array representing other unspecified monochrome or tristimulus color samplings (e.g., YZX, also known as XYZ).
[0090] Hereinafter, these arrays may be referred to as luminance (or L or Y) and chrominance, where the two chrominance arrays may be referred to as Cb and Cr; regardless of the actual color representation method used. The actual color representation method used can be indicated, for example, in the coded bitstream using the video usability information (VUI) syntax such as that of HEVC. A component can be defined as an array or a single sample from one of the three sample arrays (luminance and two chrominances) or an array or a single sample of an array constituting a monochrome format picture.
[0091] A picture can be defined as a frame or a field. A frame includes a matrix of luminance samples and possibly corresponding chrominance samples. A field is a set of alternating sample rows of a frame and can be used as an encoder input when the source signal is interlaced. When compared to the luminance sample array, the chrominance sample array may be absent (and thus monochrome sampling may be used), or the chrominance sample array may be subsampled.
[0092] Some chrominance formats can be summarized as follows:
[0093] - In monochrome sampling, there is only one sample array, which nominally can be considered as the luminance array.
[0094] - In 4:2:0 sampling, each of the two chrominance arrays has half the height and half the width of the luminance array.
[0095] - In 4:2:2 sampling, each of the two chrominance arrays has the same height as the luminance array and half the width.
[0096] - In 4:4:4 sampling without using separate color planes, each of the two chrominance arrays has the same height and width as the luminance array.
[0097] An encoding format or standard may allow encoding the sample arrays as separate color planes into the bitstream and decoding the separately encoded color planes from the bitstream. When separate color planes are used, each of them is processed separately as a picture (by the encoder and / or decoder) using monochrome sampling.
[0098] A partition can be defined as dividing a set into subsets such that each element of the set is in exactly one of the subsets. In video coding, a partition can be defined as dividing a picture or a sub-region of a picture into subsets such that each element of the picture or the sub-region of the picture is in exactly one of the subsets. For example, in partitions related to HEVC encoding and / or decoding and / or VVC encoding and / or decoding, the following terms can be used. A coding block can be defined as an NxN block of samples for some value of N such that dividing a coding tree block into coding blocks is a partition. A coding tree block (CTB) can be defined as an NxN block of samples for some value of N such that dividing a component into coding tree blocks is a partition. A coding tree unit (CTU) can be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture with three sample arrays, or a coding tree block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures (used to encode the samples). A coding unit (CU) can be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture with three sample arrays, or a coding block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures (used to encode the samples). A CU with the maximum allowed size can be named an LCU (largest coding unit) or a coding tree unit (CTU), and a video picture is divided into non-overlapping LCUs.
[0099] In HEVC, a CU consists of one or more prediction units (PUs) that define the prediction process for the samples within the CU and one or more transform units (TUs) that define the prediction error coding process for the samples in the CU. Typically, a CU consists of a square block of samples, the size of which can be selected from a predefined set of possible CU sizes. Each PU and TU can also be split into smaller PUs and TUs to increase the granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it that defines which prediction will be applied to the pixels within the PU (e.g., motion vector information for a PU used for inter prediction and intra prediction directionality information for a PU used for intra prediction).
[0100] Each TU can be associated with information that describes the prediction error decoding process for the samples within the TU (including, for example, DCT coefficient information). Typically, it is signaled at the CU level whether prediction error coding is applied to each CU. In the absence of a prediction error residual associated with a CU, it can be considered that there is no TU for that CU. Dividing an image into CUs and dividing a CU into PUs and TUs is typically signaled in the bitstream, allowing the decoder to reproduce the expected structure of these units.
[0101] In the draft versions of H.266 / VVC, the following partitioning applies. Note that what is described here may still evolve in later draft versions of H.266 / VVC before the standard is finalized. Although the maximum CTU size has been increased to 128x128, similar to HEVC, pictures are partitioned into CTUs. Coding tree units (CTUs) are first partitioned by a quadtree (also known as a four - tree) structure. Then the leaf nodes of the quadtree can be further partitioned by a multi - type tree structure. There are four split types in the multi - type tree structure, namely, vertical binary split, horizontal binary split, vertical ternary split, and horizontal ternary split. The leaf nodes of the multi - type tree are called coding units (CUs). CUs, PUs, and TUs have the same block size, unless the CU is too large for the maximum transform length. The segmentation structure of the CTU is a quadtree with a nested multi - type tree that uses binary and ternary splits, i.e., there is no separate CU, PU, and TU concept, except for CUs that need to be too large for the maximum transform length. A CU can have a square or rectangular shape.
[0102] The basic unit of the encoder output for some coding formats (such as VVC, v) and the input to the decoder for some coding formats (such as VVC) is the network abstraction layer (NAL) unit. For transmission over packet - oriented networks or storage in structured files, NAL units can be encapsulated into packets or similar structures.
[0103] A byte - stream format can be specified for the NAL unit stream for transmission or storage environments that do not provide a frame structure. The byte - stream format separates NAL units from each other by appending a start code in front of each NAL unit. To avoid misdetecting NAL unit boundaries, the encoder runs a byte - oriented start - code emulation prevention algorithm that adds emulation prevention bytes to the NAL unit payload if the start code would otherwise occur. To enable direct gateway operation between packet - oriented and stream - oriented systems, start - code emulation prevention can always be performed regardless of whether the byte - stream format is in use.
[0104] An NAL unit can be defined as a syntax structure that contains an indication of the data type to follow and the bytes that contain the data, which are in the form of an RBSP and are scattered with emulation prevention bytes if necessary. The raw byte sequence payload (RBSP) can be defined as a syntax structure that contains an integral number of bytes that are encapsulated in an NAL unit. The RBSP is either empty or has the form of a data bit - string that contains syntax elements, followed by an RBSP stop bit, and then zero or more subsequent bits equal to 0.
[0105] An NAL unit consists of a header and a payload. The NAL unit header indicates the type of the NAL unit, etc.
[0106] NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0107] Non-VCL NAL units can be, for example, one of the following types: Sequence Parameter Set, Picture Parameter Set, Supplemental Enhancement Information (SEI) NAL unit, Access Unit Delimiter, end of sequence NAL unit, end of bitstream NAL unit, or filler data NAL unit. Parameter sets may be required for reconstructing decoded pictures, while many other non-VCL NAL units are not required for reconstructing decoded sample values.
[0108] Some coding formats specify parameter sets that can carry parameter values required for decoding or reconstructing decoded pictures. Parameters can be defined as syntax elements of a parameter set. A parameter set can be defined as a syntax structure that contains parameters and can be, for example, referenced from or activated by another syntax structure using an identifier.
[0109] Some types of parameter sets are briefly described below, but it is to be understood that other types of parameter sets may exist and embodiments can be applied but are not limited to the types of parameter sets described. Parameters that remain constant throughout an encoded video sequence can be included in the Sequence Parameter Set (SPS). In addition to the parameters that may be required for the decoding process, the Sequence Parameter Set may optionally contain Video Usability Information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. The Picture Parameter Set (PPS) contains such parameters that may be constant across multiple encoded pictures. The Picture Parameter Set may include parameters that can be referenced by the coded picture segments of one or more encoded pictures. The Header Parameter Set (HPS) has been proposed to contain such parameters that may vary based on pictures.
[0110] A bitstream can be defined as a sequence of bits that, in some coding formats or standards, can take the form of a NAL unit stream or a byte stream, which forms a representation of encoded pictures and associated data (forming one or more encoded video sequences). In the same logical channel, such as in the same file or in the same connection of a communication protocol, a second bitstream can follow a first bitstream. An elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. In some coding formats or standards, the end of a first bitstream can be indicated by a specific NAL unit, which can be referred to as the end of bitstream (EOB) NAL unit and is the last NAL unit of the bitstream.
[0111] A bitstream portion may be defined as a contiguous subset of a bitstream. In some contexts, it may be required that a bitstream portion consists of one or more complete syntax structures and no incomplete syntax structures. In other contexts, a bitstream portion may include any contiguous segment of the bitstream and may contain incomplete (multiple) syntax structures.
[0112] Phrases such as along the bitstream (e.g., along the bitstream indication) or along the coding units of the bitstream (e.g., along the coded tile indication) may be used in the claims and the described embodiments to refer to transmission, signaling, or storage in a way that associates "out-of-band" data with the bitstream or the coding units respectively but is not included within the bitstream or the coding units. Decoding of phrases such as along the bitstream or along the coding units of the bitstream etc. may refer to decoding of reference out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) associated with the bitstream or the coding units respectively. For example, when the bitstream is contained in a container file, such as a file conforming to the ISO base media file format, and certain file metadata is stored in the file in a way that associates the metadata with the bitstream, such as boxes in the sample entry of the track containing the bitstream, sample groups of the track containing the bitstream, or the timing metadata track associated with the track containing the bitstream.
[0113] A coded video sequence (CVS) may be defined as such a sequence of coded pictures in decoding order that can be decoded independently and is followed by the end of another coded video sequence or bitstream. When a specific NAL unit (which may be referred to as an end-of-sequence (EOS) NAL unit) appears in the bitstream, the coded video sequence may be additionally or alternatively specified as ended.
[0114] An image can be split into independently encodable and decodable image segments (e.g., slices and / or tiles and / or tile groups). Such image segments can enable parallel processing. A "slice" in this description may refer to an image segment consisting of a certain number of base coding units processed in default encoding or decoding order, while a "tile" may refer to an image segment that has been defined as a rectangular image region along a tile grid. A tile group may be defined as a group of one or more tiles. Image segments can be encoded as separate units in the bitstream, such as VCL NAL units in H.264 / AVC and HEVC and VVC. The encoded image segment may include a header and a payload, where the header contains parameter values required to decode the payload. The payload of a slice may be referred to as slice data.
[0115] In HEVC, a picture can be partitioned into rectangular tiles, and contains an integer number of LCU. In HEVC, the partitioning of tiles forms a regular grid, where the height and width of tiles differ by at most one LCU from each other. In HEVC, a slice is defined as an integer number of coding tree units contained in a single slice segment and all subsequent dependent slice segments (if any) before the next single slice segment (if any) within the same access unit. In HEVC, a slice segment is defined as an integer number of coding tree units that are consecutively ordered in a tile scan and contained in a single NAL unit. Dividing each picture into slice segments is partitioning. In HEVC, a single slice segment is defined as a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values of the previous slice segment, and a dependent slice segment is defined as a slice segment for which the values of some of the syntax elements of the slice segment header are inferred from the values of the previous single slice segment in decoding order. In HEVC, a slice header is defined as the slice segment header of a single slice segment that is the current slice segment or the single slice segment before the current dependent slice segment, and a slice segment header is defined as a part of the coded slice segment that contains data elements related to the first or all coding tree units represented in the slice segment. If tiles are not used, CUs are scanned in raster scan order of the LCU within the tile or within the picture. Within an LCU, the CUs have a specific scan order.
[0116] Thus, video coding standards and specifications can allow an encoder to partition a coded picture into coded slices, etc. Intra-picture prediction is typically disabled across slice boundaries. Thus, slices can be regarded as a way to split a coded picture into independently decodable segments. In H.264 / AVC and HEVC, intra-picture prediction can be disabled across slice boundaries. Thus, slices can be regarded as a way to split a coded picture into independently decodable segments, and thus slices are typically regarded as the basic unit for transmission. In many cases, an encoder can indicate in the bitstream which types of intra-picture prediction are turned off across slice boundaries, and a decoder operation takes this information into account, for example, when figuring out which prediction sources are available. For example, if neighboring CUs reside in different slices, samples from neighboring CUs can be considered unavailable for intra-frame prediction.
[0117] In a draft version of VVC (i.e., VVC draft 5), partitioning a picture into slices, tiles, and bricks is defined as follows. Other draft versions of VVC may define partitioning a picture into slices, tiles, and bricks similarly.
[0118] The picture is divided into one or more tile rows and one or more tile columns. Partitioning the picture into tiles forms a tile grid, which may be characterized by a list of tile column widths (in CTUs) and a list of tile row heights (in CTUs).
[0119] A tile is a sequence of coding tree units (CTUs) that covers a "cell" in the tile grid, i.e., a rectangular region of the picture. A tile is divided into one or more bricks, each brick consisting of a number of CTU rows within the tile. A tile that is not partitioned into multiple bricks is also referred to as a brick. However, a brick, which is a proper subset of a tile, is not referred to as a tile.
[0120] A slice contains multiple tiles of the picture or multiple bricks of a tile. A slice is a VCL NAL unit.
[0121] Two slice modes are supported, namely, the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a sequence of tiles in the tile raster scan of the picture. In the rectangular slice mode, a slice contains multiple bricks that together form a rectangular region of the picture. The bricks within a rectangular slice are in the brick raster scan order of the slice.
[0122] Brick scan can be defined as a specific sequential ordering of the CTUs that partition the picture, where the CTUs are sequentially ordered in the CTU raster scan within a brick, the bricks within a tile are sequentially ordered in the brick raster scan of the tile, and the tiles in the picture are sequentially ordered in the tile raster scan of the picture. It may be required, for example, in an encoding standard that for the first CTU of each encoded slice NAL unit, the encoded slice NAL units should be in the order of increasing CTU address in the brick scan order, where the CTU address can be defined as increasing in the CTU raster scan within the picture. Raster scan can be defined as mapping a rectangular two-dimensional pattern to a one-dimensional pattern such that the first entry in the one-dimensional pattern comes from the first row of the two-dimensional pattern scanned from left to right, followed by the second row, third row, etc. of the pattern scanned similarly from left to right (downward).
[0123] Figure 5a An example of raster scan slice partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan slices. Figure 5b An example of rectangular slice partitioning of a picture (with 18x12 CTUs) is shown, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices. Figure 5cAn example of a picture partitioned into tiles, bricks, and rectangular slices is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows), 11 bricks (the upper left tile contains 1 brick, the upper right tile contains 5 bricks, the lower left tile contains 2 bricks, and the lower right tile contains 3 bricks), and 4 rectangular slices.
[0124] The partitioning of tiles, bricks, and rectangular slices is specified in the Picture Parameter Set (PPS). Figure 6 The syntax for indicating the partitioning of tiles and bricks is shown, which is performed in two stages: the tile grid (i.e., tile column width and tile row height) is provided as the first stage, and the indicated tiles are then further partitioned into bricks.
[0125] There are two modes for indicating the tile grid: uniform (indicated by the syntax element uniform_tile_spacing_flag with a value equal to 1) and explicit. In uniform tile spacing, the tiles have equal widths (except possibly for the rightmost tile column) and equal heights (except possibly for the bottommost tile row). In explicit tile spacing, the widths and heights of the tile columns and rows (respectively) are indicated (in units of CTUs), except for the rightmost column and bottommost row (respectively).
[0126] Similar to how the tile grid is indicated, there are two modes for indicating how a tile is split into bricks, i.e., each tile can indicate uniform or explicit brick spacing. The signaling is similar to that of the tile rows.
[0127] When rectangular slices are used, the following is provided for each slice using a syntax structure that includes and follows the syntax element num_slices_in_pic_minus1:
[0128] - The upper left brick index (except for the first slice which is inferred to have index 0)
[0129] - The differential brick index of the lower right brick of the slice relative to the upper left brick index.
[0130] Figure 6 The semantics of the syntax elements have been specified as follows in the VVC draft 5:
[0131] single_tile_in_pic_flag being equal to 1 specifies that there is only one tile reference PPS in each picture. single_tile_in_pic_flag being equal to 0 specifies that there are more than one tile reference PPSs in each picture. Note – If there is no further tile partitioning within a tile, the whole tile is called a brick. When a picture contains only a single tile without further tile partitioning, it is called a single brick. The requirement for bitstream conformance is that the value of single_tile_in_pic_flag should be the same for all PPSs activated within the CVS.
[0132] uniform_tile_spacing_flag being equal to 1 specifies that the tile column boundaries and the same tile row boundaries are uniformly distributed over the picture and are signaled using the syntax elements tile_cols_width_minus1 and tile_rows_height_minus1. uniform_tile_spacing_flag being equal to 0 specifies that the tile column boundaries and the same tile row boundaries may or may not be uniformly distributed over the picture and are signaled using the syntax elements num_tile_columns_minus1 and num_tile_rows_minus1 and a list of syntax element pairs tile_column_width_minus1[i] and tile_row_height_minus1[i]. When absent, the value of uniform_tile_spacing_flag is inferred to be equal to 1.
[0133] tile_cols_width_minus1 plus 1 specifies the width in CTBs of the tile columns (excluding the rightmost tile column of the picture) when uniform_tile_spacing_flag is equal to 1. The value of tile_cols_width_minus1 should be in the range (inclusive) of 0 to PicWidthInCtbsY - 1. When absent, the value of tile_cols_width_minus1 is inferred to be equal to PicWidthInCtbsY – 1.
[0134] tile_rows_height_minus1 plus 1 specifies the height of the tile rows in CTB units (excluding the bottom tile row of the picture) when uniform_tile_spacing_flag is equal to 1. The value of tile_rows_height_minus1 should be in the range of 0 to PicHeightInCtbsY - 1 (inclusive). When not present, the value of tile_rows_height_minus1 is inferred to be equal to PicHeightInCtbsY – 1.
[0135] num_tile_columns_minus1 plus 1 specifies the number of tile columns into which the picture is partitioned when uniform_tile_spacing_flag is equal to 0. The value of num_tile_columns_minus1 should be in the range of 0 to PicWidthInCtbsY - 1 (inclusive). If single_tile_in_pic_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be specified during CTB raster scan, tile scan, and brick scan.
[0136] num_tile_rows_minus1 plus 1 specifies the number of tile rows into which the picture is partitioned when uniform_tile_spacing_flag is equal to 0. The value of num_tile_rows_minus1 should be in the range of 0 to PicHeightInCtbsY - 1 (inclusive). If single_tile_in_pic_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be specified during CTB raster scan, tile scan, and brick scan. The variable NumTilesInPic is set to be equal to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1). When single_tile_in_pic_flag is equal to 0, NumTilesInPic should be greater than 1.
[0137] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in CTB units.
[0138] tile_row_height_minus1[i] + 1 specifies the height of the i-th tile row in CTB units.
[0139] brick_splitting_present_flag being equal to 1 specifies that one or more tiles of a picture that references PPS can be divided into two or more bricks. brick_splitting_present_flag being equal to 0 specifies that no picture tiles that reference PPS are divided into two or more bricks.
[0140] brick_split_flag[i] being equal to 1 specifies that the i-th tile is divided into two or more bricks. brick_split_flag[i] being equal to 0 specifies that the i-th tile is not divided into two or more bricks. When not present, the value of brick_split_flag[i] is inferred to be equal to 0.
[0141] uniform_brick_spacing_flag[i] being equal to 1 specifies that horizontal brick boundaries are evenly distributed over the i-th tile and are signaled using the syntax element brick_height_minus1[i]. uniform_brick_spacing_flag[i] being equal to 0 specifies that horizontal brick boundaries may or may not be evenly distributed over the i-th tile and are signaled using the list of syntax elements num_brick_rows_minus1[i] and syntax element brick_row_height_minus1[i][j]. When not present, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1.
[0142] brick_height_minus1[i] + 1 specifies the height of the brick rows in CTB units (excluding the bottom brick in the i-th tile) when uniform_brick_spacing_flag[i] is equal to 1. When present, the value of brick_height_minus1 should be in the range (inclusive) of 0 to RowHeight[i] - 2. When not present, the value of brick_height_minus1[i] is inferred to be equal to RowHeight[i] – 1.
[0143] num_brick_rows_minus1[i] + 1 specifies the number of bricks for partitioning the i-th tile when uniform_brick_spacing_flag[i] equals 0. When present, the value of num_brick_rows_minus1[i] shall be in the range (inclusive) of 1 to RowHeight[i] - 1. If brick_split_flag[i] equals 0, the value of num_brick_rows_minus1[i] is inferred to be equal to 0. Otherwise, when uniform_brick_spacing_flag[i] equals 1, the value of num_brick_rows_minus1[i] is inferred during CTB raster scan, tile scan, and brick scan processes.
[0144] brick_row_height_minus1[i][j] + 1 specifies the height of the j-th brick in the i-th tile in CTB units when uniform_tile_spacing_flag equals 0.
[0145] The following variables are derived, and when uniform_tile_spacing_flag equals 1, the values of num_tile_columns_minus1 and num_tile_rows_minus1 are inferred, and for each i in the range from 0 to NumTilesInPic - 1 (inclusive), when uniform_brick_spacing_flag[i] equals 1, the value of num_brick_rows_minus1[i] is inferred by calling CTB raster scan, tile scan, and brick scan processes:
[0146] - The list RowHeight[j] for j in the range from 0 to num_tile_rows_minus1 (inclusive) specifies the height of the j-th tile row in CTB units,
[0147] - The list CtbAddrRsToBs[ctbAddrRs] for ctbAddrRs in the range from 0 to PicSizeInCtbsY - 1 (inclusive) specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the brick scan,
[0148] - The list CtbAddrBsToBs[ctbAddrBs] for ctbAddrBs in the range from 0 to PicSizeInCtbsY - 1 (inclusive) specifies the conversion from the CTB address in the brick scan to the CTB address in the CTB raster scan of the picture,
[0149] - The list BrickId[ctbAddrBs] specifies the conversion from the CTB address in the brick scan to the brick ID for the range of ctbAddrBs from 0 to PicSizeInCtbsY-1 (inclusive).
[0150] - The list NumCtusInBrick[brickIdx] specifies the conversion from the brick index to the number of CTUs in the brick for the range of brickIdx from 0 to NumBricksInPic-1 (inclusive).
[0151] - The list FirstCtbAddrBs[brickIdx] specifies the conversion from the brick ID to the CTB address in the brick scan of the first CTB in the brick for the range of brickIdx from 0 to NumBricksInPic-1 (inclusive).
[0152] single_brick_per_slice_flag being equal to 1 specifies that each slice referring to this PPS contains one brick. single_brick_per_slice_flag being equal to 0 specifies that the slices referring to this PPS may contain more than one brick. When absent, the value of single_brick_per_slice_flag is inferred to be equal to 1.
[0153] rect_slice_flag being equal to 0 specifies that the bricks within each slice are in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag being equal to 1 specifies that the bricks within each slice cover a rectangular region of the picture and the slice information is signaled in the PPS. When single_brick_per_slice_flag is equal to 1, rect_slice_flag is inferred to be equal to 1.
[0154] num_slices_in_pic_minus1 plus 1 specifies the number of slices in each picture referring to the PPS. The value of num_slices_in_pic_minus1 should be in the range from 0 to NumBricksInPic-1 (inclusive). When absent and single_brick_per_slice_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to NumBricksInPic-1.
[0155] top_left_brick_idx[i] specifies the brick index of the brick located at the upper left corner of the i-th slice. For any i not equal to j, the value of top_left_brick_idx[i] should not be equal to the value of top_left_brick_idx[j]. When it does not exist, the value of top_left_brick_idx[i] is inferred to be equal to i. The length of the top_left_brick_idx[i] syntax element is Ceil(Log2(NumBricksInPic)) bits.
[0156] bottom_right_brick_idx_delta[i] specifies the difference between the brick index of the brick located at the lower right corner of the i-th slice and top_left_brick_idx[i]. When single_brick_per_slice_flag is equal to 1, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 0. The length of the bottom_right_brick_idx_delta[i] syntax element is Ceil(Log2(NumBricksInPic - top_left_brick_idx[i])) bits.
[0157] The requirement for bitstream conformance is that a slice should consist of multiple complete tiles or a contiguous sequence of complete bricks of only one tile. The variables NumBricksInSlice[i] and BricksToSliceMap[j] that specify the number of bricks in the i-th slice and the mapping of bricks to slices are derived as follows:
[0158]
[0159] Thus, a rather complex syntax structure is designed to signal tile and brick partitioning. It is suboptimal in many respects, such as in terms of the number of syntax elements, lines in the syntax, number of operation modes (i.e., the separate-unified and explicit modes for both tiles and bricks, and the separate-indicated tile and brick boundaries), and bit counting when signaling.
[0160] VVC Draft 6 supports sub - pictures (also known as sub - images). A sub - picture can be defined as a rectangular region of one or more slices within a picture, where one or more slices are complete. Thus, a sub - picture consists of one or more slices that jointly cover a rectangular region of the picture. The slices of a sub - picture may need to be rectangular slices. The partitioning of the picture into sub - pictures can be indicated in the SPS and / or decoded from the SPS. One or more of the following features can be indicated (e.g., by the encoder) or decoded (e.g., by the decoder) or inferred (e.g., by the encoder and / or decoder) jointly for sub - pictures or separately for each sub - picture: i) whether a sub - picture is considered a picture during the decoding process; in some cases, this feature excludes loop filter operations, which may be indicated / decoded / inferred separately; ii) whether loop filter operations are performed across sub - picture boundaries.
[0161] An improved method for signaling tile and brick partitioning is now introduced.
[0162] Figure 7 The method for encoding according to the first aspect as shown includes: encoding a bitstream (700) at a time, the bitstream including an indication of tile columns and an indication of the brick height of one or more tile columns, or encoding an indication of tile columns and an indication of the brick height of one or more tile columns at a time in the bitstream or along the bitstream; inferring (702) potential tile rows after detecting brick rows aligned through the picture; inferring (704) whether the boundaries of the potential tile rows are the boundaries of tile rows or indicating whether the boundaries of the potential tile rows are the boundaries of tile rows; and encoding one or more pictures (706) into the bitstream using the indicated tile columns, the indicated or inferred tile rows, and the indicated brick height, where the one or more pictures are partitioned into a tile grid along the indicated tile columns and the indicated or inferred tile rows, the tiles in the tile grid include an integer number of coding tree units and are partitioned into one or more bricks, and where a brick includes an integer number of rows of coding tree units within a tile.
[0163] The method for decoding according to the first aspect includes: decoding an indication of tile columns and an indication of the brick height of one or more tile columns from the bitstream or along the bitstream at a time; inferring potential tile rows when detecting brick rows aligned through the picture; inferring whether the boundaries of the potential tile rows are the boundaries of tile rows or decoding whether the boundaries of the potential tile rows are the boundaries of tile rows; and decoding one or more pictures from the bitstream using the indicated tile columns, the indicated or inferred tile rows, and the indicated brick height, where the one or more pictures are partitioned into a tile grid along the indicated tile columns and the indicated or inferred tile rows, the tiles in the tile grid include an integer number of coding tree units and are partitioned into one or more bricks, and where a brick includes an integer number of rows of coding tree units within a tile.
[0164] Accordingly, the tile column and the brick height in the tile column direction are indicated by the encoder and / or decoded by the decoder, excluding the tile row height. Potential tile row boundaries are inferred as those tile row boundaries where the bricks are aligned (horizontally) through the picture. Accordingly, for a certain potential tile row boundary, it can be inferred that the potential tile row boundary is a tile row boundary. Inferring that a potential tile row boundary is a tile row boundary can be based on conclusions made from other information available in the syntax structure or the absence of certain information in the syntax structure (e.g., the absence of a specific flag).
[0165] Alternatively, it can be indicated by the encoder and / or decoded by the decoder whether a potential tile row boundary is a tile row boundary. This indication can be based on, for example, one or more flags present in the syntax structure.
[0166] Accordingly, by only indicating the tile column and the brick height in the tile column direction and inferring potential tile rows, partitioning can be signaled without signaling the tile row height. Thus, the coding efficiency is increased and the bit rate required for the signaling is reduced.
[0167] In the following, a number of exemplary embodiments of the first aspect are provided, namely, the syntax and semantics for indicating the tile column and the brick height in the tile column direction and excluding the tile row height. The embodiments are equally applicable to the encoding of generating a bitstream portion conforming to the syntax and semantics and the decoding of decoding the bitstream portion according to the syntax and semantics.
[0168] Example 1 of syntax and semantics:
[0169]
[0170] A value of uniform_tile_col_spacing_flag equal to 1 specifies that the tile column boundaries are evenly distributed over the picture and is signaled using the syntax element tile_cols_width_minus1. A value of uniform_tile_spacing_flag equal to 0 specifies that the tile column boundaries may or may not be evenly distributed over the picture and is signaled using the syntax element num_tile_columns_minus1 and a list of the syntax elements tile_column_width_minus1[i]. When absent, the value of uniform_tile_col_spacing_flag is inferred to be equal to 1.
[0171] The semantics of tile_cols_width_minus1, num_tile_columns_minus1, and tile_column_width_minus1[i] are specified to be the same as those of the syntax elements with the same names in VVC Draft 5.
[0172] If uniform_tile_col_spacing_flag equals 1, NumTileColsInPic is set to be equal to PicWidthInCtbsY / (tile_cols_width_minus1 + 1) + PicWidthInCtbsY % (tile_cols_width_minus1 + 1). Otherwise, NumTileColsInPic is set to be equal to num_tile_columns_minus1 + 1.
[0173] uniform_brick_spacing_flag[i] being equal to 1 specifies that the horizontal brick boundaries are evenly distributed over the i-th tile column and is signaled using the syntax element brick_height_minus1[i]. uniform_brick_spacing_flag[i] being equal to 0 specifies that the horizontal brick boundaries may or may not be evenly distributed over the i-th tile column and is signaled using the syntax element num_brick_rows_minus1[i] and a list of syntax elements brick_row_height_minus1[i][j]. When absent, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1.
[0174] brick_height_minus1[i] plus 1 specifies the height of the brick rows in CTBs (excluding the bottom brick in the i-th tile column) when uniform_brick_spacing_flag[i] equals 1.
[0175] num_brick_rows_minus1[i] plus 1 specifies the number of bricks partitioning the i-th tile column when uniform_brick_spacing_flag[i] equals 0.
[0176] brick_row_height_minus1[i][j] plus 1 specifies the height of the j-th brick in the i-th tile column in CTBs when uniform_tile_spacing_flag equals 0.
[0177] Example 2 of syntax and semantics
[0178] Example 2 is similar to Example 1, but additionally indicates by the encoder and / or decodes whether the partitioning of the current tile column into bricks is the same as that of the previous tile column.
[0179]
[0180] The semantics of the syntax elements are the same as in Example 1, but the following semantics of copy_previous_col_flag[i] are added:
[0181] copy_previous_col_flag[i] being equal to 1 specifies all of the following:
[0182] - uniform_brick_spacing_flag[i] is inferred to be equal to uniform_brick_spacing_flag[i–1].
[0183] When brick_height_minus1[i–1] exists, brick_height_minus1[i] is inferred to be equal to brick_height_minus1[i–1].
[0184] When num_brick_rows_minus1[i–1] exists, num_brick_rows_minus1[i] is inferred to be equal to num_brick_rows_minus1[i–1], and for each value of j in the range from 0 to num_brick_rows_minus1[i]–1 (inclusive), brick_row_height_minus1[i][j] is inferred to be equal to brick_row_height_minus1[i-1][j].
[0185] On the other hand, the sub-optimal syntax structure problem of VVC Draft 5 can be alleviated by a method in which the tile column width, tile row height, or brick height can be indicated by the encoder and / or decoded by the decoder in a certain predefined scan order until the remaining tile columns, tile rows, or bricks (respectively) are indicated or decoded (respectively) to have equal dimensions.
[0186] The method for encoding according to this second aspect is in Figure 8is shown, wherein the method comprises: step a) indicating (800) the number of partitions to be assigned; b) determining (802) the number of units to be assigned to the partitions; c) indicating (804) whether the plurality of units to be assigned are evenly assigned to the plurality of partitions; and if not, then d) indicating (806) the number of units to be assigned to the next partition, and e) repeating (808) steps c) and d) until all units have been assigned to partitions.
[0187] A method for decoding according to a second aspect comprises: step a) decoding the number of partitions to be assigned; b) determining the number of units to be assigned to the partitions; c) decoding whether the plurality of units to be assigned are evenly assigned to the plurality of partitions; and if not, then d) decoding the number of units to be assigned to the next partition, and e) repeating steps c) and d) until all units have been assigned to partitions.
[0188] According to an embodiment applicable to encoding and / or decoding, the number of units to be assigned to a partition is one of the following: the picture width in CTUs (e.g., when the partition is a tile column), the picture height in CTUs (e.g., when the partition is a tile row, or when the partition is a brick row indicated for one or more complete tile columns at a time), the number of CTU rows in a tile (e.g., when the partition is a brick of a tile).
[0189] According to an embodiment applicable to encoding and / or decoding, the partition is one or more of the following: a tile column, a tile row, a brick row.
[0190] According to an embodiment applicable to encoding and / or decoding, the unit is a rectangular block of samples of a picture. For example, Figure 8 the unit in the second aspect shown may be a coding tree block.
[0191] Thus, by indicating the number of units to be assigned to partitions in a predefined scan order, a significant saving in the number of required syntax elements and the bitrate required for such signaling can be achieved, especially if the plurality of unassigned units should be evenly assigned to the remaining partitions.
[0192] Figure 9 is shown Figure 8Examples of how the method can be implemented according to an embodiment. Thus, a plurality of partitions to be assigned, such as tile columns and / or tile rows, are first indicated (900), and a plurality of units NU to be assigned to the partitions, such as units of coding tree blocks (CTBs), are determined (902). To create a loop for checking that all units have been assigned to a partition, it is checked (904) whether the number of partitions NP to be assigned is greater than 1. If not, i.e., NP = 1, it is inferred or indicated (910) that all remaining units that have not been assigned should be assigned to the remaining partition.
[0193] However, if NP > 1, it is checked (906) whether the number of units NU to be assigned is divisible by the number of partitions NP. If so, it is determined (908) whether the plurality of units NU should be evenly assigned to the remaining partitions. If so, it is inferred or indicated (910) that all remaining units NU that have not been assigned should be evenly assigned to the remaining partition(s). If it is noted that the number of units NU to be assigned is not divisible by the number of partitions NP (906), or it is determined (908) that the plurality of units NU should not be evenly assigned to the remaining partitions, the number of units to be assigned to the next partition in a predefined scan order is indicated (912). The number of units NU to be assigned is decreased (914) by the number of units indicated to be assigned to the next partition, and the number of partitions NP to be assigned is decremented by 1 (916). Then, if the number of partitions NP to be assigned is greater than 1, the process loops back to the check (904).
[0194] Figure 9 The method can be similarly implemented for decoding according to the embodiments described hereinafter. First, a plurality of partitions to be assigned, such as tile columns and / or tile rows, are decoded from or along the bitstream, and a plurality of units NU to be assigned to the partitions, such as units of coding tree blocks (CTBs), are determined. To create a loop for checking that all units have been assigned to a partition, it is checked whether the number of partitions NP to be assigned is greater than 1. If not, i.e., NP = 1, it is inferred or decoded from the bitstream or along the bitstream that all remaining units that have not been assigned should be assigned to the remaining partition.
[0195] However, if NP > 1, check whether the number NU of units to be assigned is divisible by the number NP of partitions. If so, decode from the bitstream or along the bitstream whether the multiple units NU should be evenly assigned to the remaining partitions. If it is noted that the number NU of units to be assigned is not divisible by the number NP of partitions, or it is decoded from the bitstream or along the bitstream that the multiple units NU should not be evenly assigned to the remaining partitions, then the number of units to be assigned to the next partition in a predefined scan order is decoded from the bitstream or along the bitstream. The number NU of units to be assigned is reduced by the indicated number of units to be assigned to the next partition, and the number NP of partitions to be assigned is decremented by 1. Then, if the number NP of partitions to be assigned is greater than 1, loop back to the check.
[0196] In the following, exemplary embodiments of a second aspect are provided, namely, the syntax and semantics of unified signaling for explicit / unified tile / brick partitioning. The embodiments are equally applicable to the encoding of a bitstream portion that conforms to the syntax and semantics and the decoding of the bitstream portion according to the syntax and semantics.
[0197] In this example, unified signaling is used to specify the tile column width and tile row height, and compared with VVC draft 5, the signaling of the bricks remains unchanged.
[0198]
[0199]
[0200] rem_tile_col_equal_flag[i] being equal to 0 specifies that the tile columns with indices in the range from 0 to i (inclusive) may or may not have equal widths in terms of CTBs. rem_tile_col_equal_flag[i] being equal to 1 specifies that the tile columns with indices in the range from 0 to i (inclusive) have equal widths in terms of CTBs, and tile_column_width_minus1[j] is inferred to be equal to remWidthInCtbsY / (i + 1) for each value of j in the range from 0 to i (inclusive).
[0201] rem_tile_row_equal_flag[i] being equal to 0 specifies that the tile rows within the range from 0 to i (inclusive) of the index may or may not have equal height in terms of CTBs. rem_tile_row_equal_flag[i] being equal to 1 specifies that the tile rows within the range from 0 to i (inclusive) of the index have equal height in terms of CTBs, and tile_row_width_minus1[j] is inferred to be equal to remHeightInCtbsY / (i + 1) for each value of j within the range from 0 to i (inclusive).
[0202] The semantics of other syntax elements can be specified in the same way as the semantics of the syntax elements with the same name in VVC Draft 5.
[0203] In the following, an exemplary embodiment of another aspect is provided, that is, the syntax and semantics for both the brick height indicating the tile column and the tile column direction and the unified signaling for explicit / unified tile / brick partitioning. The embodiment is equally applicable to the encoding for generating a bitstream portion conforming to the syntax and semantics and the decoding for decoding the bitstream portion according to the syntax and semantics.
[0204] The exemplary embodiment for encoding can be summarized as follows, and the exemplary embodiment can be applied to decoding by replacing the term "indicating" with the term "decoding".
[0205] The tile columns are indicated as follows:
[0206] - The number of tile columns is indicated (num_tile_columns_minus1).
[0207] - The following is indicated in a loop that traverses the tile columns from right to left until all tile columns have been traversed or until the remaining tile columns are indicated to have equal width:
[0208] ○ If the remaining width (in terms of CTBs) is divisible by the number of tile columns that have not been specified, it is indicated whether the remaining tile columns have equal width (rem_tile_col_equal_flag[i]).
[0209] ○ If the widths of the remaining tile columns are not equal, the tile column widths are indicated (tile_column_width_minus1[i]).
[0210] The following is indicated in a loop that traverses the tile columns from left to right for a rectangular slice and contains only one loop entry that specifies the tile row height of the raster scan slice:
[0211] - If the brick partition of a tile column is the same as the brick partition of the previous tile column, it is a flag. This flag does not exist for the leftmost tile column (copy_previous_col_flag[i]). Note that this flag can be omitted in this exemplary embodiment, or there are other alternatives discussed further below that can be used to implement functionality similar to that implemented by this flag.
[0212] - When the tile brick partition of a tile column is different from that of the previous tile column, the bricks of the tile column are indicated as follows:
[0213] ○ The number of bricks in the tile column is indicated (num_bricks_minus1[i]).
[0214] ○ The following is indicated in a loop that traverses the bricks from right to left until all the bricks of the tile column have been traversed or until the remaining bricks of the tile column are indicated to have equal height:
[0215] ■ If the remaining height (in CTBs) is divisible by the number of bricks that have not been specified, it is indicated whether the remaining bricks have equal height (rem_brick_height_equal_flag[i][j]).
[0216] ■ If the heights of the remaining bricks are not equal, the brick heights are indicated (brick_height_minus1[i][j]).
[0217] The following syntax can be used in this exemplary embodiment:
[0218]
[0219]
[0220] num_tile_columns_minus1 plus 1 specifies the number of tile columns for partitioning the picture when uniform_tile_spacing_flag is equal to 0. The value of num_tile_columns_minus1 should be in the range (inclusive) of 0 to PicWidthInCtbsY - 1. When single_tile_in_pic_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be equal to 0.
[0221] rem_tile_col_equal_flag[i] being equal to 0 specifies that the tile columns with indices in the range from 0 to i (inclusive) may or may not have equal widths in CTB units. rem_tile_col_equal_flag[i] being equal to 1 specifies that the tile columns with indices in the range from 0 to i (inclusive) are inferred to have equal widths in CTB units. When absent, the value of rem_tile_equal_flag[i] is inferred to be equal to 0.
[0222] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in CTB units.
[0223] copy_previous_col_flag[i] being equal to 0 specifies that num_bricks_minus1[i] exists. copy_previous_col_flag[i] being equal to 1 specifies all of the following:
[0224] - num_bricks_minus1[i] is inferred to be equal to num_bricks_minus1[i–1],
[0225] - For all such values of j in the range from 1 to num_bricks_minus1[i] (inclusive), rem_brick_height_equal_flag[i][j] is inferred to be equal to rem_brick_height_equal_flag[i-1][j], where the value of rem_brick_height_equal_flag[i-1][j] exists or is inferred.
[0226] - For all such values of j in the range from 1 to num_bricks_minus1[i] (inclusive), brick_height_minus1[i][j] is inferred to be equal to brick_height_minus1[i-1][j], where the value of brick_height_minus[i-1][j] exists.
[0227] copy_previous_col_flag[i] being equal to 0 specifies that num_bricks_minus1[i] exists. copy_previous_col_flag[i] being equal to 1 specifies all of the following:
[0228] - num_bricks_minus1[i] is inferred to be equal to num_bricks_minus1[i–1],
[0229] - For all such values of j in the range (inclusive) from 1 to num_bricks_minus1[i], rem_brick_height_equal_flag[i][j] is inferred to be equal to rem_brick_height_equal_flag[i - 1][j], where the value of rem_brick_height_equal_flag[i - 1][j] exists or is inferred.
[0230] - For all such values of j in the range (inclusive) from 1 to num_bricks_minus1[i], brick_height_minus1[i][j] is inferred to be equal to brick_height_minus1[i - 1][j], where the value of brick_height_minus[i - 1][j] exists.
[0231] rem_brick_height_equal_flag[i][j] being equal to 0 specifies that the bricks within the range (inclusive) of indices from 0 to j in the i-th tile column may or may not have equal heights in terms of CTB. rem_brick_height_equal_flag[i][j] being equal to 1 specifies that the bricks within the range (inclusive) of indices from 0 to j in the i-th tile column are inferred to have equal heights in terms of CTB. When it does not exist, the value of rem_brick_height_equal_flag[i][j] is inferred to be equal to 0.
[0232] brick_height_minus1[i][j] plus 1 specifies the height of the j-th brick in the i-th tile column in terms of CTB.
[0233] The semantics of other syntax elements can be specified in the same way as the semantics of the syntax elements with the same names in VVC Draft 5.
[0234] The decoding process may use variables defined as follows or similar to the following definitions:
[0235] For i ranging from 0 to num_tile_columns_minus1 (inclusive), the list colWidth[i] specifying the width of the i-th tile column in terms of CTB is derived as follows:
[0236]
[0237] i ranges from 0 to num_tile_columns_minus1 (inclusive), and j ranges from 0 to num_bricks_minus1[i] (inclusive), specifying the list colBrickHeight[i][j] of the height of the j-th brick row within the i-th tile column in CTB units; j ranges from 0 to NumTileRows-1 (inclusive), specifying the list RowHeight[j] of the height of the j-th tile row in CTB units; j ranges from 0 to NumTileRows (inclusive), specifying the list tileRowBd[j] of the position of the boundary of the j-th tile row in CTB units; the value of NumTileRows and the value of NumTilesInPic are derived as follows:
[0238]
[0239]
[0240]
[0241] When single_tile_in_pic_flag is equal to 0, NumTilesInPic should be greater than 1. The list tileColBd[i] specifying the position of the boundary of the i-th tile column in CTB units is derived as follows: for (tileColBd[0] = 0, i = 0; i <= num_tile_columns_minus1; i++)
[0242] tileColBd[i+1] = tileColBd[i] + colWidth[i]
[0243] A variable NumBricksInPic that specifies the number of bricks in the picture referenced by PPS and a brickIdx ranging from 0 to NumBricksInPic - 1 (inclusive) are defined. A list BrickColBd[brickIdx], BrickRowBd[brickIdx], BrickWidth[brickIdx], and BrickHeight[brickIdx] that specifies the positions of the vertical brick boundaries in CTB units, the positions of the horizontal brick boundaries in CTB units, the width of the brick columns in CTB units, and the height of the brick columns in CTB units is derived. And for each i from 0 to NumTilesInPic - 1 (inclusive), when uniform_brick_spacing_flag[i] is equal to 1, the value of num_brick_rows_minus1[i] is inferred as follows:
[0244]
[0245] According to an embodiment, a method for encoding includes: step a) determining the number of units to be assigned to a partition; b) indicating or inferring the number of partitions of a specific size to be assigned; c) indicating the size of the partitions of a specific size or the number of units; and d) indicating or inferring the number of partitions of a uniform size to be assigned.
[0246] According to an embodiment for encoding, step d) includes: step d1) indicating a count of units; d2) repeatedly assigning the count of units to partitions until the number of unassigned units is less than the count of units; and d3) if the number of unassigned units is greater than 0, assigning the unassigned units to the last partition in a predefined scan order.
[0247] Therefore, the encoding method according to the above embodiments is illustrated in Figure 14a and can be implemented independently or in combination with one or more of the embodiments described herein. The method includes: determining (1400) the number of units to be assigned to a partition and initially marking them as unassigned; indicating or inferring (1402) the number of partitions of a specific size to be assigned; indicating (1404) the size of the partitions of a specific size and accordingly marking the unassigned units as being assigned to the partitions in a predefined scan order; indicating (1406) a count of units; repeatedly assigning (1408) the count of units to partitions and accordingly marking the unassigned units as being assigned in a predefined scan order until the number of unassigned units is less than the count of units; and if the number of unassigned units is greater than 0, assigning (1410) the unassigned units to the last partition.
[0248] According to an embodiment, a method for decoding includes: step a) determining the number of units to be assigned to partitions; b) decoding or inferring the number of partitions of a specific size to be assigned; c) decoding the size of the partitions of the specific size or the number of units; and d) decoding or inferring the number of partitions of a uniform size to be assigned.
[0249] According to an embodiment for decoding, step d) includes: step d1) decoding the count of units; d2) repeatedly assigning the count of a single unit to a partition until the number of unassigned units is less than the count of units; d3) if the number of unassigned units is greater than 0, assigning the unassigned units to the last partition in a predefined scan order.
[0250] Thus, the decoding method according to the above embodiments is illustrated in Figure 14b and can be implemented independently or in combination with one or more of the embodiments described herein. The method includes: determining (1450) the number of units to be assigned to partitions; determining (1452) the number of partitions of a specific size to be assigned; determining (1454) the size of the partitions of the specific size and correspondingly marking the unassigned units to be assigned to the partitions in a predefined scan order; determining (1456) the count of units; repeatedly assigning (1458) the unit count to partitions and correspondingly marking the unassigned units to be assigned in a predefined scan order until the number of unassigned units is less than the unit count; and if the number of unassigned units is greater than 0, assigning (1460) the unassigned units to the last partition.
[0251] According to an embodiment for encoding and / or decoding, the method further includes:
[0252] - Initially marking a plurality of units as unassigned, or equivalently initially marking a plurality of units as unassigned, for example as part of or in connection with step a);
[0253] - Marking the unassigned units to be assigned according to the number of partitions of a specific size to be assigned and the size of the partitions of the specific size, for example as part of and / or in connection with steps b) and / or c);
[0254] - Marking the unassigned units to be assigned whenever a counted unit is assigned to a partition, for example as part of or in connection with step d2).
[0255] Assigning units to partitions and / or marking unassigned units as assigned can be done according to a scan order. In an embodiment, the scan order is predefined, for example, in an encoding standard. The scan order can be, for example, from left to right (e.g., for assigning CTU columns to tile columns), or from top to bottom (e.g., for assigning CTU rows to tile rows, or for assigning CTU rows within a tile to bricks). In an embodiment, the encoder selects a scan order from a list of predefined scan orders and indicates the selected scan order in or along the bitstream, for example, as an index into the list of predefined scan orders. In an embodiment, the decoder decodes the scan order from the bitstream or along the bitstream, for example, an index into the list of predefined scan orders.
[0256] According to an embodiment for encoding and / or decoding, step c) further includes or is followed by the step of assigning a size or a plurality of units to a partition of a defined size.
[0257] The size of the partition of a defined size can be indicated as the number of units to be assigned.
[0258] According to an embodiment applicable to encoding and / or decoding, the unit is one of the following: CTB, CTU, CTU row, CTU column, grid unit (for a grid used to indicate subpicture partitioning), grid row (for a grid used to indicate subpicture partitioning), grid column (for a grid used to indicate subpicture partitioning).
[0259] In an embodiment for encoding and / or decoding, information about units that have not been assigned to a partition can be maintained. After determining the number of units to be assigned to a partition, all units can be immediately marked as unassigned. When a set of units is assigned to a partition, the set of units can be marked as assigned, or the marking of the set of units as "unassigned" can be removed or cancelled. Marking units as assigned or unassigned can be represented, for example, by an array variable, etc., where each unit to be assigned is represented by an entry in the array, and the value of the entry in the array indicates whether the corresponding unit has been assigned. In another example, the number of unassigned units (i.e., the number of remaining unassigned units) is maintained by the steps of the embodiment.
[0260] According to an embodiment applicable to encoding and / or decoding, the number of units to be assigned to a partition is one of the following: the picture width in CTUs (e.g., when the partition is a tile column), the picture height in CTUs (e.g., when the partition is a tile row, or when the partition is a brick row indicated for one or more complete tile columns at a time), the number of CTU rows in a tile (e.g., when the partition is a brick of a tile), the picture width in grid columns (for a grid used to indicate subpicture partitioning), the picture height in grid rows (for a grid used to indicate subpicture partitioning).
[0261] According to embodiments applicable to encoding and / or decoding, the partitioning is one or more of the following: tile columns, tile rows, brick rows, grid columns (for a grid used to indicate subpicture partitioning), grid rows (for a grid used to indicate subpicture partitioning).
[0262] According to embodiments applicable to encoding and / or decoding, where step d includes steps d1, d2, and d3 above, the following syntax etc. can be used:
[0263]
[0264]
[0265] num_exp_tile_columns_minus1 plus 1 specifies the number of explicitly provided tile column widths.
[0266] num_exp_tile_rows_minus1 plus 1 specifies the number of explicitly provided tile row heights.
[0267] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of CTB, where i ranges from 0 to num_exp_tile_columns_minus1 - 1 (inclusive). tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with an index greater than or equal to num_exp_tile_columns_minus1.
[0268] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of CTB, where i ranges from 0 to num_exp_tile_rows_minus1 - 1 (inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with an index greater than or equal to num_exp_tile_rows_minus1.
[0269] The brick_splitting_present_flag and brick_spit_flag[i] can be specified as described earlier.
[0270] NumTilesInPic can be inferred to be equal to the number of tiles in the picture. NumTileColumns can be inferred to be equal to the number of tile columns in the picture. RowHeight[tileY] can be inferred to be equal to the number of CTUs in the tileY-th tile row.
[0271] num_exp_brick_rows_minus1[i] plus 1 specifies the number of brick row heights explicitly provided in the i-th tile. When not present, the value of num_exp_brick_rows_minus[i] can be inferred to be equal to -1.
[0272] brick_row_height_minus1[i][j] plus 1 specifies the height of the j-th brick in the i-th tile in units of CTB, where j ranges from 0 to num_exp_brick_rows_minus1[i] - 1 (inclusive). brick_row_height_minus1[i][num_exp_brick_rows_minus1] is used to derive the height of brick rows in the i-th tile with indices greater than or equal to num_exp_brick_rows_minus1[i].
[0273] According to an example embodiment encoded (or separately decoded as indicated by the following parentheses) using the above syntax, tile columns are specified as follows:
[0274] - The number of explicitly provided (or decoded) tile column widths is indicated (or decoded) (num_exp_tile_columns_minus1)
[0275] - The tile column widths are explicitly provided (or decoded and assigned) from left to right (tile_column_width_minus1[i]).
[0276] - The last explicitly provided (or decoded) tile column width (tile_column_width_minus1[num_exp_tile_columns_minus1]) is repeated until no other tile columns of that width fit within the picture boundaries.
[0277] - The remaining CTUs (if any) not assigned to any tile column are assigned to the rightmost tile column.
[0278] Tile row heights and brick rows are specified similarly to tile rows.
[0279] According to embodiments applicable to the above syntax and encoding and / or decoding, for a variable numTileColumns that specifies the number of tile columns and an i ranging from 0 to numTileColumns-1 (inclusive), a list colWidth[i] that specifies the width of the i-th tile column in units of CTB is derived as follows:
[0280]
[0281] According to embodiments applicable to the above syntax and encoding and / or decoding, for a variable numTileRows that specifies the number of tile rows and a j ranging from 0 to numTileRows-1 (inclusive), a list RowHeight[j] that specifies the height of the j-th tile row in units of CTB is derived as follows:
[0282]
[0283] In some embodiments, the following may apply:
[0284] - The variable NumTilesInPic is set to be equal to numTileColumns * numTileRows.
[0285] - For an i ranging from 0 to numTileColumns (inclusive), a list tileColBd[i] that specifies the position of the boundary of the i-th tile column in units of CTB is derived as follows: for (tileColBd[0]=0, i = 0; i < numTileColumns; i++)
[0286] tileColBd[i + 1] = tileColBd[i] + colWidth[i]
[0287] - For a j ranging from 0 to numTileRows (inclusive), a list tileRowBd[j] that specifies the position of the boundary of the j-th tile row in units of CTB is derived as follows:
[0288] for (tileRowBd[0]=0, j = 0; j < numTileRows; j++)
[0289] tileRowBd[j + 1] = tileRowBd[j] + RowHeight[j]
[0290] According to an embodiment applicable to the above syntax and encoding and / or decoding, a variable NumBricksInPic specifying the number of bricks in a picture referring to the PPS, and a brickIdx ranging from 0 to NumBricksInPic-1 (inclusive), a list BrickColBd[brickIdx], BrickRowBd[brickIdx], BrickWidth[brickIdx], and BrickHeight[brickIdx] specifying the position of the vertical brick boundary in units of CTB, the position of the horizontal brick boundary in units of CTB, the brick width in units of CTB, and the brick height in units of CTB are derived as follows:
[0291]
[0292] According to an embodiment applicable to encoding and / or decoding, step d includes: determining the number of units not yet assigned to a partition by decrementing the number of units in a partition of explicit size from the number of units to be assigned to the partition, and the method further includes:
[0293] - Assigning a partition to a partition of explicit size according to the size or number of units in the partition of explicit size and according to a predefined or indicated / decoded scan order;
[0294] - Assigning a partition to a partition of uniform size by dividing the number of units not yet assigned to a partition by the number of partitions of uniform size and according to a predefined or indicated / decoded scan order.
[0295] According to an embodiment applicable to encoding and / or decoding, the number of partitions of explicit size is indicated and / or decoded in a higher-level syntax structure (such as SPS), while the size of the partitions of explicit size and / or the number of partitions of uniform size to be assigned can be indicated and / or decoded in a lower-level syntax structure (such as PPS).
[0296] According to an embodiment, it is inferred (e.g., predefined in an encoding standard) that the number of partitions of explicit size is equal to 1.
[0297] According to an embodiment applicable to encoding and / or decoding, step d includes:
[0298] - Determining a set or list of partition sizes by which the number of unassigned units is divisible;
[0299] - If the number of items in the set or list is equal to 1, inferring that the number of partitions of uniform size is equal to 1;
[0300] - Indicating and / or decoding an index (or the like) corresponding to an item in a set or list, where the index indicates the number of uniformly sized partitions to be assigned.
[0301] According to embodiments applicable to encoding and / or decoding, the set or list of partition sizes by which the number of unassigned units is divisible is constrained by excluding partition sizes smaller than a threshold, where the threshold may be predefined in a coding standard or indicated / decoded. For example, the minimum tile column width in a CTU may be predefined or indicated / decoded.
[0302] According to an embodiment, the index corresponding to an item in a set or list is encoded using a fixed-length codeword, such as u(v), where the length of the codeword is determined by the number of items in the set or list.
[0303] Indicates whether the boundary of the tile row is the boundary of a potential tile row
[0304] In some embodiments, when the horizontal tile boundary aligns across pictures, the tile row boundary is inferred. This section presents embodiments for signaling the tile row boundary. This embodiment can be applied together with any embodiment in which the tile boundary is indicated in the syntax before the tile row boundary.
[0305] This embodiment may include one or more of the following steps (some of which have been described earlier):
[0306] - Potential tile row boundaries are inferred as those tile row boundaries where the tile boundaries are aligned through the picture (horizontally).
[0307] - The encoder infers in the bitstream or along the bitstream and / or the decoder decodes from the bitstream or along the bitstream whether all aligned tile boundaries form tile row boundaries. For example, a flag in the bitstream syntax can be used.
[0308] - If all aligned tile boundaries do not form tile row boundaries, the encoder indicates in the bitstream or along the bitstream and / or the decoder decodes from the bitstream or along the bitstream for each aligned tile boundary whether that boundary is a tile row boundary. For example, there may be a flag in the bitstream syntax for each aligned tile boundary (excluding the aligned tile boundaries that are picture boundaries).
[0309] For example, the following syntax can be used:
[0310]
[0311] In another embodiment, NumAlignedBrickRows can be derived like NumTileRows.
[0312] The semantics of the proposed syntax elements can be specified as follows:
[0313] explicit_tile_rows_flag being equal to 0 specifies that tile row boundaries are inferred whenever horizontal tile boundaries align across pictures. explicit_tile_rows_flag being equal to 1 specifies the existence of the tile_row_flag[i] syntax element.
[0314] tile_row_flag[i] being equal to 0 specifies that the i-th such horizontal boundary where the horizontal tile boundary aligns across pictures is not a tile row boundary. tile_row_flag[i] being equal to 1 specifies that the i-th such horizontal boundary where the horizontal tile boundary aligns across pictures is a tile row boundary. The 0-th such horizontal boundary where the horizontal tile boundary aligns on the picture is the top boundary of the picture.
[0315] Indicates the tile columns partitioned in the same way as the bricks
[0316] In some example embodiments, the syntax includes the following indication: the tile columns are partitioned into bricks in the same order as the previous tile columns in the loop entry order (e.g., scanning tile columns from left to right). It is to be understood that the embodiments are similarly applicable to cases without an indication or any similar indication. For example, the scanning order of the tile columns can be from right to left, so it can be indicated that the brick partitioning of the current tile column is copied from the tile column on the right. In another example, it is indicated that all tile columns of equal width have the same brick partitioning. In yet another example, the index of the brick column from which the brick partitioning is copied is indicated. It is also to be understood that the embodiments are similarly applicable when another way of partitioning tile columns into bricks is derived based on an earlier indication. Some related embodiments are presented in this section.
[0317] In an embodiment, the encoder indicates in the bitstream or along the bitstream and / or the decoder decodes from the bitstream or along the bitstream whether all tile columns of the same width (e.g., in CTB) have the same brick partitioning. For example, a syntax element called same_brick_spacing_in_equally_wide_tile_cols_flag can be used.
[0318] In an embodiment, the encoder indicates in the bitstream or along the bitstream and / or the decoder decodes from the bitstream or along the bitstream the number of adjacent tile columns in the loop entry order (e.g., scanning tile columns from left to right) that have the same brick partitioning. For example, a syntax element that can be called num_tile_cols_with_same_brick_partitioning_minus1[i] can be u(v)-encoded, where v is determined by the remaining tile columns for which the brick partitioning has not been indicated.
[0319] In an embodiment, an encoder may indicate in or along a bitstream and / or a decoder may decode from or along a bitstream whether there exists a (plural) syntax element (e.g., copy_previous_col_flag[i]) related to indicating that a tile column is partitioned into bricks in the same way. In an embodiment, the indication is in a sequence-level syntax structure such as an SPS. In another embodiment, the indication is in a picture-level syntax structure such as a PPS. In an embodiment, it is predefined in an encoding standard, for example, and the absence of the (plural) syntax element related to indicating that a tile column is partitioned into bricks in the same way causes the brick partitioning to be indicated and / or decoded for each tile column one by one. In an embodiment, it is predefined in an encoding standard, for example, and the absence of the (plural) syntax element related to indicating that a tile column is partitioned into bricks in the same way causes the brick partitioning to be indicated and / or decoded for one tile column and inferred to be the same for all tile columns. In an embodiment, the method for handling the absence of the (plural) syntax element related to indicating that a tile column is partitioned into bricks in the same way is indicated by an encoder in or along a bitstream, for example, in an SPS, and / or decoded by a decoder from or along a bitstream, for example, from an SPS. The method may be indicated and / or decoded in a predefined set of procedures, which may include but may not be limited to i) the brick partitioning is to be indicated and / or decoded for each tile column one by one, and ii) the brick partitioning is to be indicated and / or decoded for one tile column and inferred to be the same for all tile columns.
[0320] In some embodiments, the number of units to be assigned to a partition is determined or inferred. The number of units may be, for example, the number of CTU columns in a picture, the number of CTU rows in a picture, or the number of CTU rows in a tile. In decoding, the number of units to be assigned to a partition may be decoded from an SPS and / or a PPS. In some embodiments, it may be desirable to avoid parsing dependencies between syntax structures, and all syntax elements sufficient to infer the number of units are included in the same syntax structure, such as a PPS. For example, the picture width in luma samples, the picture height in luma samples, and the CTU size may be included in the PPS, enabling the number of CTU columns in the picture and the number of CTU rows in the picture to be inferred.
[0321] For example, the syntax element log2_pps_ctu_size_minus5 encoded by u(2) may be included in the PPS. log2_pps_ctu_size_minus5 plus 5 specifies the luma coding tree block size of each CTU. log2_pps_ctu_size_minus5 being equal to 0, 1, or 2 specifies that the luma coding tree block size of each CTU is equal to 32×32, 64×64, or 128×128 luma samples, respectively. It may be desirable for log2_pps_ctu_size_minus5 to be less than or equal to 2. It may be required that log2_pps_ctu_size_minus5 be equal to log2_ctu_size_minus5 specified in the SPS.
[0322] Indicates the sub - graph partition
[0323] It should be noted that many of the described embodiments for indicating slice, tile, and / or brick partitioning are independent of how sub-pictures are indicated. This section presents embodiments for indicating and / or decoding the partitioning of a picture into sub-pictures. Many of the embodiments in this section can, but need not, be used independently of other embodiments, such as those for indicating slice, tile, and / or brick partitioning.
[0324] According to an embodiment, the grid used in signaling sub-picture partitioning is indicated in the bitstream or along the bitstream (e.g., in the SPS) or decoded from the bitstream or along the bitstream (e.g., from the SPS). The cells of the grid specify the units at which sub-picture boundaries are signaled. In other words, sub-picture boundaries can only be located at the boundaries of the indicated grid.
[0325] According to an embodiment, the grid is indicated and / or decoded in units of CTUs (or similar basic coding blocks of the codec).
[0326] According to an embodiment, the signaling includes information indicating the number of grid columns, the number of grid rows, the width of a grid column (e.g., the count of CTUs), and the height of a grid row (e.g., the count of CTUs). For example, the following syntax may be used:
[0327]
[0328] num_id_columns_minus1 specifies the number of columns for partitioning the picture minus one. num_id_rows_minus1 specifies the number of rows for partitioning the picture minus one. id_column_width_minus1[i]+1 specifies the width of the i - th column in units of coding tree blocks id_row_height_minus1[i]+1 specifies the height of the i - th row in units of coding tree blocks. The right - most grid column can be inferred to consist of the columns of coding tree blocks not assigned to other grid columns. The bottom - most grid row can be inferred to consist of coding tree blocks not assigned to other grid rows. According to another example, the following syntax can be used:
[0329] Indicates the split of the tile partition as a sub - graph partition
[0330]
[0331] When using the above signaling, all grid columns have equal width, except possibly the rightmost grid column, which can be inferred to consist of columns of coding tree blocks not assigned to other grid columns. Similarly, all grid rows have equal height, except possibly the bottommost grid row, which can be inferred to consist of coding tree blocks not assigned to other grid rows. id_column_width_minus1 plus 1 specifies the width of the grid column in units of coding tree blocks. id_row_height_minus1 plus 1 specifies the height of the grid row in units of coding tree blocks.
[0332] According to an embodiment, assigning grid units to sub - pictures is indicated and / or decoded in the bitstream or along the bitstream (e.g., in the SPS).
[0333] In an embodiment, assigning grid units to sub - pictures is indicated and / or decoded by the index of the bottom - right grid unit of each sub - picture (possibly separate from the last sub - picture where the bottom - right grid unit can be inferred to be the bottom - right grid unit of the grid). For example, the following syntax etc. can be used (and is considered appended after the syntax etc. presented above):
[0334]
[0335] num_subpics_minus2 plus 2 specifies the number of sub - pictures within a picture. bottom_right_grid_idx_length - minus1 plus 1 specifies the number of bits used to represent the syntax element bottom_right_grid_idx_delta[i]. When i is greater than 0, bottom_right_grid_idx_delta[i] specifies the difference between the grid index of the grid cell at the bottom - right corner of the i - th sub - picture and the grid index of the grid cell at the bottom - right corner of the (i – 1)-th sub - picture. bottom_right_grid_idx_delta[0] specifies the grid index of the grid cell at the bottom - right corner of the 0 - th sub - picture. The bottom - right corner of the last sub - picture is inferred to be the bottom - right cell of the grid. grid_idx_delta_sign_flag[i] being equal to 1 indicates the positive sign of bottom_right_grid_idx_delta[i]. sign_bottom_right_grid_idx_delta[i] being equal to 0 indicates the negative sign of bottom_right_grid_idx_delta[i]. The variables NumGridRows, NumGridCols, and NumGridCellsInPic can be inferred to be equal to the count of grid rows, grid columns, and grid cells in the picture, respectively. The variables TopLeftGridIdx[i], BottomRightGridIdx[i], NumGridCellsInSubpic[i], and GridCellsToSubpicMap[j] that specify the grid index of the grid cell at the top - left corner of the i - th sub - picture, the grid index of the grid cell at the bottom - right corner of the i - th sub - picture, the number of grid cells in the i - th sub - picture, and the mapping of grid cells to sub - pictures are derived as follows:
[0336]
[0337] Figure 9
[0338] According to an embodiment that can be used in conjunction with or independent of other embodiments, the picture - to - sub - picture partitioning is indicated in the bitstream or along the bitstream (e.g., in the SPS) and / or decoded from the bitstream or along the bitstream (e.g., from the SPS), and based on the sub - picture partitioning, the partitioning of tile columns, tile rows, and / or slices is indicated in the bitstream or along the bitstream (e.g., in the PPS) and / or decoded from the bitstream or along the bitstream (e.g., from the PPS).
[0339] According to an embodiment, the encoder indicates in the bitstream or along the bitstream the tile column boundaries and / or the decoder decodes the tile column boundaries from the bitstream or along the bitstream as follows. Since the vertical boundaries of the rectangular slices match the tile column boundaries, and since the sub-pictures are composed of one or more complete rectangular slices, the tile column boundaries can be derived to coincide with the vertical boundaries of the sub-pictures. Thus, the vertical sub-picture boundaries are obtained (e.g., by the encoder) or decoded (e.g., by the decoder) and are derived as the tile column boundaries. Additionally, the encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether there are additional tile column boundaries (in addition to those derived from the vertical sub-picture boundaries). For example, the tile_col_splitting_present_flag can be indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). The tile_col_splitting_present_flag being equal to 0 specifies that the tile column boundaries are the same as the obtained or decoded vertical sub-picture boundaries. The tile_col_splitting_present_flag being equal to 1 specifies that there are additional tile columns in addition to those columns derived from the vertical sub-picture boundaries.
[0340] In an embodiment, when the encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) that there are additional tile column boundaries (in addition to those derived from the vertical sub-picture boundaries), the following applies. The vertical sub-picture boundaries that are not picture boundaries can be sorted, e.g., from left to right, and indexed from 0 to NumSubPicCols - 2 (inclusive), where NumSubPicCols is the number of different vertical sub-picture boundaries (not picture boundaries). The i-th sub-picture column can be defined as bounded on the left by the right boundary of the (i - 1)-th sub-picture column when i > 0 or by the left picture boundary when i = 0, and bounded on the right by the i-th sub-picture boundary when i < NumSubPicCols - 1 or by the right picture boundary when i = NumSubPicCols - 1. The encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether the sub-picture column is split into more than one tile column. For example, for each sub-picture column, tile_col_split_flag[i] can be indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). tile_col_split_flag[i] equal to 0 specifies that the i-th sub-picture column consists of exactly one tile column. tile_col_split_flag[i] equal to 1 specifies that the i-th sub-picture column consists of more than one tile column.
[0341] In an embodiment, when the encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) that the sub-picture column is split into more than one tile column, the tile columns within the sub-picture column are indicated similarly to the VVC draft standard or any proposed embodiment. For example, the encoder can indicate in the bitstream or along the bitstream and / or the decoder can decode from the bitstream or along the bitstream a syntax element indicating the number of tile columns within the sub-picture column. For example, the num_tile_cols_minus2[i] syntax element can be used, where num_tile_cols_minus2[i] plus 2 specifies the number of tile columns in the i-th sub-picture column. The width of each tile column within the sub-picture column is indicated in the bitstream or along the bitstream and / or decoded from the bitstream or along the bitstream (except for the last tile column of the sub-picture column, whose width can be inferred to encompass all remaining but unallocated sub-picture regions). For example, tile_col_width_minus1[i][j] can be used, where tile_col_width_minus1[i][j] plus 1 specifies the width of the j-th tile column within the i-th sub-picture column in terms of CTUs.
[0342] The above embodiments may use, for example, the following syntax, etc.:
[0343]
[0344] According to an embodiment, the encoder indicates in the bitstream or along the bitstream the tile row boundary and / or the decoder decodes the tile row boundary from the bitstream or along the bitstream as follows. The horizontal sub-picture boundary may be juxtaposed with the tile row boundary or the brick boundary within the tile. Thus, the horizontal sub-picture boundary is obtained (e.g., by the encoder) or decoded (e.g., by the decoder), and is derived as the tile row boundary or the brick boundary.
[0345] In an embodiment, the encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether the tile row boundary matches one-to-one with the horizontal sub-picture boundary (and vice versa), or whether there are additional tile row boundaries or whether some of the horizontal sub-picture boundaries are the brick boundaries within the (multiple) tiles. For example, the tile_row_splitting_present_flag may be indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). The tile_row_splitting_present_flag being equal to 0 specifies that the tile row boundary is the same as the obtained or decoded horizontal sub-picture boundary. The tile_col_splitting_present_flag being equal to 1 specifies that there are additional tile rows in addition to those rows derived from the horizontal sub-picture boundary, or that some of the horizontal sub-picture boundaries are the brick boundaries within the (multiple) tiles.
[0346] In an embodiment, when the encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) that there are additional tile row boundaries (in addition to those derived from the horizontal sub-picture boundaries) or some of the horizontal sub-picture boundaries are brick boundaries within a (multiple) tile, the following applies. The horizontal sub-picture boundaries that are not picture boundaries can be sorted, e.g., from top to bottom, and indexed from 0 to NumSubPicRows-2 (inclusive), where NumSubPicRows is the number of different horizontal sub-picture boundaries (not picture boundaries). The i-th sub-picture row can be defined as having its top bounded by the bottom boundary of the (i-1)-th sub-picture row when i>0 or by the top picture boundary when i = 0, and its bottom bounded by the i-th sub-picture boundary when i<NumSubPicRows-1 or by the bottom picture boundary when i = NumSubPicRows-1. The encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether the sub-picture row boundary is a tile row boundary or a brick boundary. For example, for each sub-picture row, tile_row_flag[i] can be indicated in or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). tile_row_flag[i] equal to 0 specifies that the i-th sub-picture row boundary is a brick boundary within a (multiple) tile. tile_row_flag[i] equal to 1 specifies that the i-th sub-picture row boundary is a tile row boundary. The encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether the sub-picture row contains data for more than one tile row. For example, for each sub-picture column, tile_row_split_flag[i] can be indicated in or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). tile_row_split_flag[i] equal to 0 specifies that the i-th sub-picture row contains data for one tile row. tile_row_split_flag[i] equal to 1 specifies that the i-th sub-picture row contains data for more than one tile row.
[0347] In an embodiment, when the encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from or along the bitstream (e.g., from the PPS) that a subpicture row contains data of more than one tile row, the tile rows within the subpicture row are indicated similarly to the draft VVC standard or any proposed embodiment. For example, the encoder can indicate in or along the bitstream and / or the decoder can decode from or along the bitstream a syntax element indicating the number of tile row boundaries within the subpicture row. For example, the num_tile_rows_minus2[i] syntax element can be used, where num_tile_rows_minus2[i] plus 2 specifies the number of tile row boundaries in the i-th subpicture row. The height of each tile row within the subpicture row is indicated in or along the bitstream and / or decoded from or along the bitstream (except for the last tile row of the subpicture row, whose height can be inferred to enclose all remaining but unallocated regions until the next tile row boundary at or within the next (multiple) subpicture row, depending on the value of tile_row_flag[j], where j is greater than i). For example, tile_row_height_minus1[i][j] can be used, where tile_row_height_minus1[i][j] plus 1 specifies the height of the j-th tile row within the i-th subpicture row in terms of CTUs.
[0348] The above embodiments can use, for example, the following syntax, etc.:
[0349]
[0350] According to an embodiment for encoding and / or decoding, when the subpicture boundary is not collocated with the tile row boundary, the brick boundary is inferred to be collocated with the subpicture boundary. The encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether there are additional bricks (in addition to those derived from the subpicture boundary). For example, the brick_splitting_present_flag can be indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). The variable NumInferredBricksInPic is set according to the number of bricks inferred from the number of tiles of the indicated tile grid, and the tiles are further split along the inferred brick boundaries. The brick_splitting_present_flag being equal to 0 specifies that there are no brick boundaries other than those inferred from the subpicture boundary. The brick_splitting_present_flag being equal to 1 specifies that there are brick boundaries not inferred from the subpicture boundary (i.e., the number of bricks is greater than NumInferredBricksInPic).
[0351] In an embodiment, when the encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) that there are brick boundaries not inferred from the subpicture boundary, the following applies. The inferred bricks can be indexed in a predefined scan order. The encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether the inferred candidate bricks are split into more than one brick. For example, for each inferred candidate brick, the brick_split_flag[i] can be indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder). The brick_split_flag[i] being equal to 0 specifies that the i-th inferred candidate brick consists of one brick. The brick_split_flag[i] being equal to 1 specifies that the i-th inferred candidate brick consists of more than one brick.
[0352] The above embodiments can use, for example, the following syntax, etc.:
[0353]
[0354] In an embodiment, when the encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) that an inferred candidate tile is split into more than one tile, the further splitting of the inferred candidate tile is indicated in the bitstream or along the bitstream (e.g., by the encoder) and / or decoded from the bitstream or along the bitstream (e.g., by the decoder), similar to the draft VVC standard or any proposed embodiment. For example, the encoder can indicate in the bitstream or along the bitstream and / or the decoder can decode from the bitstream or along the bitstream a syntax element for the number of tiles within the inferred candidate tile. For example, the num_bricks_minus2[i] syntax element can be used, where num_bricks_minus2[i] plus 2 specifies the number of tiles in the i-th inferred candidate tile. The height of each tile within the inferred candidate tile is indicated in the bitstream or along the bitstream and / or decoded from the bitstream or along the bitstream (except for the last tile of the inferred candidate tile, whose width can be inferred to enclose all the remaining but unallocated regions of the inferred candidate tile). For example, brick_row_height_minus1[i][j] can be used, where brick_row_height_minus1[i][j] plus 1 specifies the height of the j-th tile within the i-th inferred candidate tile in terms of CTUs.
[0355] Thus, the syntax and semantics for signaling tile and brick partitioning according to an embodiment provide a significant saving in the number of syntax elements and syntax lines required for signaling. Thus, a significant saving is achieved in terms of the number of bits required to indicate tile and brick partitioning.
[0356] These benefits are illustrated by the following examples, where Figure 9 a, Figure 9 b and Figure 10 c shown three different tile and brick partitions are used to compare the performance of tile and brick partitioning according to VVC draft 5 and tile and brick partitioning according to some embodiments.
[0357] Figure 10 a and 10b present (respectively) tile and brick partitions for implementing 6K equirectangular projection (ERP) and cubemap projection (CMP) schemes, which are described (respectively) in clauses D.6.3 and D.6.4 of the omnidirectional media format (OMAF, ISO / IEC 23090-2). These schemes have been recommended in the VR industry forum guidelines. Figure 10 The scheme presented in c is otherwise equivalent to Figure 10 the scheme in b, but uses a different picture aspect ratio.
[0358] The following features are derived from the VVC Draft 5 and embodiments of both the combined indication of tile column and tile column direction, the tile height, and the unified signaling of explicit / unified tile / brick partitioning:
[0359] - The number of syntax elements for indicating tile and brick partitioning
[0360] - The number of syntax rows for indicating tile and brick partitioning
[0361] - The number of bits required for tile and brick partitioning representing the scenarios included in the following figure
[0362] - The percentage of bit count savings provided by the embodiments when compared to the VVC Draft 5 for the scenarios proposed in the following figure.
[0363] The features derived using a 128×128 luma CTB size are presented in the following table.
[0364]
[0365]
[0366] Therefore, it can be seen that the number of syntax elements is reduced from 13 to 7, and the number of required syntax rows is reduced from 26 to 20. Compared to the VVC Draft 5, the embodiments provide a bit count savings of more than 50% for each of the tile and brick partitions from Figure 10 a to Indicates an uncoded tile or brick c. It should also be noted that, as proposed, the semantics and derivation processes also become shorter.
[0367] Figure 11
[0368] In some applications, it may be reasonable to assign the content to be encoded and / or decoded to tiles and / or bricks in such a way that only a subset of the tiles and / or bricks is occupied. For example, in viewport-dependent streaming of 360-degree video, only a subset of the independently encoded picture regions (such as tiles) may be received. In another example, patch-based coding of volumetric or point cloud video is applied, and the patches only occupy a subset of the tiles and / or bricks of the picture.
[0369] According to an embodiment for encoding, the unencoded tiles or bricks are indicated in the bitstream or in the syntax structure above the slice data along the bitstream. The syntax elements of the unencoded tiles or bricks are not encoded as slice data. The unencoded tiles or bricks are reconstructed (e.g., reconstructed as decoded reference pictures) using a predefined or indicated method (such as setting the reconstructed sample values to 0 in the sample array).
[0370] According to an embodiment for decoding, an uncoded tile or brick is decoded from the bitstream or from the syntax structure above the slice data along the bitstream. The syntax elements of the uncoded tile or brick are not decoded from the slice data. The uncoded tile or brick is decoded (e.g., decoded as a decoded reference picture) using a predefined or indicated method (such as setting the reconstructed sample values to 0 in a sample array).
[0371] According to an embodiment applicable to encoding and / or decoding, the number of uncoded bricks is indicated and / or decoded for a tile column, e.g., using a variable length codeword such as ue(v). The bricks in the tile column are traversed in a predefined scan order (e.g., from bottom to top). For each traversed brick, a flag is indicated and / or decoded to determine whether the brick is uncoded. If the brick is uncoded, the remaining number of bricks to be assigned as uncoded is decremented by 1. This process is continued until there are no remaining bricks to be assigned as uncoded.
[0372] Figure 11 A block diagram of a video decoder suitable for embodiments of the present invention is shown. Figure 12 The structure of a two-layer decoder is depicted, but it should be understood that the decoding operations can be similarly applied in a single-layer decoder.
[0373] Video decoder 550 includes a first decoder section 552 for the base layer and a second decoder section 554 for the prediction layer. Block 556 illustrates a demultiplexer that is used to deliver information about the base layer picture to the first decoder section 552 and information about the prediction layer picture to the second decoder section 554. Reference numeral P’n represents the predicted representation of an image block. Reference numeral D’n represents the reconstructed prediction error signal. Blocks 704, 804 illustrate the preliminarily reconstructed image (I’n). Reference numeral R’n represents the final reconstructed image. Blocks 703, 803 illustrate the inverse transform (T -1 ). Blocks 702, 802 illustrate the inverse quantization (Q -1 ). Blocks 701, 801 illustrate the entropy decoding (E -1 ). Blocks 705, 805 illustrate the reference frame memory (RFM). Blocks 706, 806 illustrate the prediction (P) (inter-frame prediction or intra-frame prediction). Blocks 707, 807 illustrate the filtering (F). Blocks 708, 808 can be used to combine the decoded prediction error information with the predicted base layer / prediction layer image to obtain the preliminarily reconstructed image (I’n). The preliminarily reconstructed and filtered base layer image can be output from the first decoder section 552 at 709, and the preliminarily reconstructed and filtered base layer image can be output from the first decoder section 554 at 809.
[0374] In this document, a decoder should be interpreted to cover any operating unit capable of performing a decoding operation, such as a player, a receiver, a gateway, a demultiplexer, and / or a decoder.
[0375] Figure 13 A flowchart showing the operation of a decoder according to an embodiment of the present invention is shown. Except that the decoder decodes the indications from them, the decoding operation of the embodiment is similar to the encoding operation in other aspects. Accordingly, the decoding method includes: decoding (1200) a bitstream portion at a time, the bitstream portion including an indication of a tile column and an indication of the brick height of one or more tile columns; inferring (1202) potential tile rows when detecting brick rows aligned by a picture; inferring (1204) whether the boundary of the potential tile rows is the boundary of a tile row or an indication of whether the boundary of decoding the potential tile rows is the boundary of a tile row; and decoding (1206) one or more pictures from the bitstream using the indicated tile columns, the indicated or inferred tile rows, and the indicated brick height, wherein the one or more pictures are partitioned into a tile grid along the indicated tile columns and the indicated or inferred tile rows, the tiles in the tile grid include an integer number of coding tree units and are partitioned into one or more bricks, and the bricks include an integer number of coding tree unit rows within the tiles.
[0376] Indicates an uncoded rectangular slice A flowchart showing the operation of a decoder according to another embodiment of the present invention is shown. The decoding method includes: step a) decoding (1300) an indication of the number of partitions to be assigned; b) determining (1302) the number of units to be assigned to the partitions; c) determining (1304) whether the plurality of units to be assigned are evenly assigned to the plurality of partitions; and if not, then d) determining (1306) the number of units to be assigned to the next partition, and e) repeating (1308) steps c) and d) until all units have been assigned to partitions. Indicating the partitioning of a picture into rectangular slices
[0377] Embodiments of an improved method for signaling for encoding and / or decoding rectangular slices are presented in the following paragraphs. These embodiments can be applied together with or independently of the embodiments for tile and brick partitioning. The embodiments are based on the definitions and characteristics of tiles, bricks, and rectangular slices specified in VVC draft 5. With the embodiments, the bit count required to indicate a rectangular slice (e.g., indicating or deriving the position, width, and height of a rectangular slice) is reduced.
[0378] An encoding method according to a first aspect includes: indicating or inferring the position of the upper left brick of a rectangular slice in or along a bitstream; deriving from the position whether the rectangular slice includes one or more bricks of a tile; and if the rectangular slice includes one or more bricks of a tile, indicating or inferring the number of bricks in the rectangular slice in or along the bitstream.
[0379] The decoding method according to the first aspect includes: decoding or inferring the position of the upper left tile of a rectangular slice from a bitstream or along the bitstream; deriving from this position whether the rectangular slice includes one or more tiles of a picture; if the rectangular slice includes one or more tiles of a picture, decoding or inferring the number of tiles in the rectangular slice from the bitstream or along the bitstream.
[0380] The position of a tile can be, for example, the tile index in the tile scan of a picture.
[0381] In an embodiment applicable to encoding and / or decoding, when the position of the upper left tile of a rectangular slice is not the upper left tile of any picture, it is derived that the rectangular slice includes one or more tiles of a picture.
[0382] In an embodiment applicable to encoding and / or decoding, when the position of the upper left tile of a rectangular slice is the upper left tile of a picture and the picture includes multiple tiles, it is derived that the rectangular slice can include the tiles of the picture or the complete picture. In an embodiment applicable to encoding, if it is derived that the rectangular slice can include the tiles of the picture or the complete picture, it is indicated in the bitstream or along the bitstream whether the rectangular slice includes the tiles of the picture or the complete picture. In an embodiment applicable to decoding, if it is derived that the rectangular slice can include the tiles of the picture or the complete picture, it is decoded from the bitstream or along the bitstream whether the rectangular slice includes the tiles of the picture or the complete picture.
[0383] In an embodiment applicable to encoding and / or decoding, if it is inferred or indicated (as part of encoding) or decoded that a rectangular slice includes the tiles of a picture, a variable (e.g., called numDeltaValues) is set to be equal to the number of tiles within the same picture after the position of the upper left tile of the rectangular slice. If numDeltaValues is equal to 1, it is inferred that the rectangular slice exactly contains one tile. Otherwise, the variable numDeltaValues is used to derive the syntax element length of a first syntax element indicating the lower right tile of the rectangular slice or a second syntax element indicating the number of tiles in the rectangular slice or a third syntax element indicating the height (in tiles) of the rectangular slice or any similar syntax element. For example, the first syntax element or the second syntax element or the third syntax element or any similar syntax element can be u(v)-encoded, and its length is Ceil(Log2(numDeltaValues)) bits.
[0384] In an exemplary embodiment, the following syntax is used:
[0385]
[0386] The semantics of the syntax element can be specified as described earlier, except that bottom_right_brick_idx_delta[i] specifies the difference between the brick index of the brick at the bottom-right corner of the i-th slice and top_left_brick_idx[i]. When single_brick_per_slice_flag is equal to 1, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 0. When it is absent, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 0. The variable numDeltaValues that specifies the number of values that bottom_right_brick_idx_delta[i] can have is derived as follows:
[0387]
[0388] The length of the bottom_right_brick_idx_delta[i] syntax element is Ceil(Log2(numDeltaValues)) bits. The variable NumBricksInTile[brickIdx] specifies the number of bricks in the tile that contains the brick with index brickIdx in the brick scan of the picture. When the brick is identified by the index brickIdx in the brick scan of the picture, the variable BrickIdxInTile[brickIdx] specifies the index of the brick within the tile that contains the brick.
[0389] In embodiments applicable to encoding and / or decoding, the position of the top-left brick of the rectangular slice is inferred. Initially, all brick positions are marked as empty. The loop that assigns bricks to the rectangular slice is included in or along the bitstream, or decoded from or along the bitstream. For each loop entry, the top-left brick of the rectangular slice is inferred to be the next empty brick position in the predefined, indicated, or decoded scan order. For example, it can be predefined in the coding standard in which the brick scan order within the picture is used. The bottom-right brick of the rectangular slice can be inferred, indicated, or decoded, for example as described above, and the bricks that form the rectangle bounded by the top-left and bottom-right bricks are marked as assigned. The same or a similar process is repeated for each loop entry.
[0390] In embodiments applicable to encoding and / or decoding, when it is derived (as part of encoding or decoding) or indicated (as part of encoding) or decoded (as part of decoding) that the rectangular slice includes a complete tile, the syntax element that indicates the bottom-right brick of the rectangular slice is derived from one or more of the following:
[0391] - A set of possible lower-right tile positions is derived. This set can include only those tile positions that are the last tile positions within a tile, are at or below the tile row of the upper-left tile containing the rectangular slice, are at or to the right of the tile column of the upper-left tile containing the rectangular slice, and enclose a rectangular set of empty tile positions (not yet assigned to any rectangular slice).
[0392] - Entries in the set of possible lower-right tile positions are indexed or enumerated.
[0393] - The length of the syntax element indicating the u(v) coding of the lower-right tile of the rectangular slice is derived from the number of entries in the set of possible lower-right tile positions. If the number of entries numEnt is equal to 1, the lower-right tile index does not need to be indicated or decoded. Otherwise, the length of the syntax element is equal to Ceil(Log2(numEnt)) bits.
[0394] - The syntax element indicating the lower-right tile of the rectangular slice is an index into the enumerated set of possible lower-right tile positions.
[0395] The encoding method according to the second aspect includes
[0396] - indicating or inferring the position of the upper-left tile of the rectangular slice in the bitstream or along the bitstream;
[0397] - determining from this position whether the rectangular slice includes one or more tiles of the tile;
[0398] - if the rectangular slice includes one or more tiles of the tile, inferring that the width of the rectangular slice is equal to one tile column, otherwise if the upper-left tile is on the rightmost tile column, inferring that the width of the rectangular slice is equal to one tile column, and otherwise indicating the width of the rectangular slice in the tile column in the bitstream or along the bitstream;
[0399] - if the rectangular slice includes one or more tiles of the tile, indicating or inferring the number of tiles in the rectangular slice in the bitstream or along the bitstream; otherwise if the upper-left tile is on the bottommost tile row, inferring that the height of the rectangular slice is equal to one tile row, and otherwise indicating the height of the rectangular slice in the tile row in the bitstream or along the bitstream.
[0400] The decoding method according to the second aspect includes
[0401] - decoding or inferring the position of the upper-left tile of the rectangular slice from the bitstream or along the bitstream;
[0402] - determining from this position whether the rectangular slice includes one or more tiles of the tile;
[0403] - If the rectangular slice includes one or more bricks of a tile, infer that the width of the rectangular slice is equal to one tile column; otherwise, if the top-left brick is on the rightmost tile column, infer that the width of the rectangular slice is equal to one tile column, and otherwise decode the width of the rectangular slice in the tile column from the bitstream or along the bitstream;
[0404] - If the rectangular slice includes one or more bricks of a tile, decode or infer the number of bricks in the rectangular slice from the bitstream or along the bitstream; otherwise, if the top-left brick is on the bottommost tile row, infer that the height of the rectangular slice is equal to one tile row, and otherwise decode the height of the rectangular slice in the tile row from the bitstream or along the bitstream.
[0405] In embodiments applicable to encoding and / or decoding and applicable to the first aspect and / or the second aspect, when it is derived or indicated or decoded that the rectangular slice contains bricks of a tile and the tile contains two bricks and the current brick (i.e., the top-left brick of the rectangular slice) is the top brick of the tile, or the current brick is the bottommost brick of the tile, the number of bricks in the rectangular slice is inferred to be equal to 1 (in terms of bricks).
[0406] In an example embodiment, the following syntax is used:
[0407]
[0408] The variables and semantics of the syntax elements can be specified as described earlier, and the following are added:
[0409] - tlBrickIdx can be specified as the next empty brick position in a predefined scan order, such as the brick scan in a picture, as described earlier. tlBrickIdx is re-derived for each value of i, i.e., for each loop entry.
[0410] - numFreeColumnsOnTheRight[brickIdx] is a variable indicating the number of tile columns to the right of the brick, where the index is brickIdx.
[0411] - slice_width_minus1[i] plus 1 specifies the width of the i-th rectangular slice in the tile column. When it does not exist, slice_width_minus[i] is inferred to be equal to 0.
[0412] - When full_tiles_in_slice_flag[i] is equal to 0, it specifies that the i-th rectangular slice contains one or more bricks of a single tile. When full_tiles_in_slice_flag[i] is equal to 1, it specifies that the i-th rectangular slice contains one or more complete slices. When BrickIdxInTile[tlBrickIdx] is greater than 0 (i.e., when the top-left brick of the rectangular slice is not the top-left brick of any tile), full_tiles_in_slice_flag[i] is inferred to be equal to 0. When slice_width_minus1[i] is greater than 0 or when BrickIdxInTile[tlBrickIdx] is equal to 0 and NumBricksInTile[tlBrickIdx] is equal to 1, full_tiles_in_slice_flag[i] is inferred to be equal to 1.
[0413] - If full_tiles_in_slice_flag[i] is equal to 0, then numFreeRowsBelow[brickIdx] is a variable indicating the number of bricks in the tile below the brick with index brickIdx in the same tile. Otherwise, numFreeRowsBelow[brickIdx] is a variable indicating the number of tile rows below the tile containing the brick with index brickIdx.
[0414] - If full_tiles_in_slice_flag[i] is equal to 0, then slice_height_minus1[i] + 1 specifies the height of the i-th rectangular slice in the brick. Otherwise, slice_height_minus1[i] + 1 specifies the height of the i-th rectangular slice in the tile row. When it does not exist, slice_height_minus[i] is inferred to be equal to 0.
[0415] In embodiments applicable to encoding and / or decoding, based on the position of the top-left brick of the rectangular slice, the length of the slice width syntax element (e.g., slice_width_minus1) is derived from a plurality of possible values. Using the above variables and syntax elements, the length of slice_width_minus1 is equal to Ceil(Log2(numFreeColumnsOnTheRight[tlBrickIdx]+1)).
[0416] In embodiments applicable to encoding and / or decoding, the length of the slice height syntax element (e.g., slice_height_minus1) is derived from a number of possible values based on the top-left tile position of the rectangular slice and whether the rectangular slice includes tiles of a single tile or a complete tile. Using the above variables and syntax elements, the length of slice_height_minus1 is equal to Ceil(Log2(numDeltaValues)), where numDeltaValues is derived as follows:
[0417]
[0418] Embodiments of an improved method for signaling rectangular slices for encoding and / or decoding are presented in the following paragraphs. These embodiments can be applied together with or independently of embodiments for tile and block partitioning. The embodiments are based on the definitions and characteristics of one or more of tiles, blocks, rectangular slices, and sub-pictures specified in VVC draft 6. Using the embodiments, the bit count required to indicate a rectangular slice (e.g., to indicate or derive the position, width, and height of the rectangular slice) is reduced.
[0419] According to an embodiment, one and only one rectangular slice of each sub-picture is encoded. The encoder indicates in the bitstream or along the bitstream (e.g., in the PPS) that for pictures in the indicated range, one and only one rectangular slice of each sub-picture is encoded. The indicated range can be, for example, a picture that references the PPS and includes an indication that one and only one rectangular slice of each sub-picture is used. The encoder omits explicit signaling of the rectangular slice boundaries. The encoder infers that the boundaries of the rectangular slice are the same as the sub-picture boundaries.
[0420] According to an embodiment, the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) an indication that each sub-picture includes one and only one rectangular slice in the pictures in the indicated range. The indicated range can be, for example, a picture that references the PPS that includes the indication. The decoder omits explicit signaling of decoding the rectangular slice boundaries. The decoder infers that the boundaries of the rectangular slice are the same as the sub-picture boundaries.
[0421] The following syntax, etc., can be used together with the above embodiments, where single_slice_per_subpic_flag equal to 0 specifies that a sub-picture can include any number of rectangular slices, and single_slice_per_subpic_flag equal to 1 specifies that each sub-picture includes one and only one rectangular slice:
[0422]
[0423] According to an embodiment, the encoder indicates in or along the bitstream (e.g., in the PPS) and / or the decoder decodes from the bitstream or along the bitstream (e.g., from the PPS) whether a subpicture consists of one and only one rectangular slice, or whether it contains more than one rectangular slice. For example, a syntax element encoded with u(1) (e.g., called subpic_split_flag[i]) can indicate (when equal to 0) that the i-th subpicture consists of exactly one rectangular slice or (when equal to 1) that the i-th subpicture consists of more than one rectangular slice.
[0424] According to an embodiment, the encoder indicates and / or the decoder uses tile indices within a subpicture to partition the subpicture into rectangular slices. The tiles within a subpicture are indexed, e.g., starting from 0, and incremented by 1 in a predefined scan order (such as the tile scan order within the subpicture), i.e., the tile raster scan within the subpicture is in major order, and the brick raster scan within a tile is in minor order. Since the value range of the brick indices within a subpicture is smaller than the value range of the brick indices within a picture, the syntax element indicating the brick index within a subpicture is likely to be shorter than the corresponding syntax element for the brick index within a picture, and thus the signaling can be made more compact.
[0425] According to an embodiment, the following syntax etc. can be used.
[0426]
[0427]
[0428] A rect_slice_splitting_present_flag equal to 0 specifies that there is exactly one rectangular slice in each sub-picture. A rect_slice_splitting_present_flag equal to 1 specifies that each sub-picture may contain one or more rectangular slices. bottom_right_brick_idx_length_minus1 plus 1 specifies the length of the bottom_right_brick_idx_delta[i] syntax element. The NumSubPics variable is derived to be equal to the number of sub-pictures within a picture, e.g., based on sub-picture signaling in the SPS. A subpic_split_flag[i] equal to 0 specifies that the i-th sub-picture consists of exactly one rectangular slice. A subpic_split_flag[i] equal to 1 specifies that the i-th sub-picture consists of more than one rectangular slice. num_slices_in_subpic_minus2[i] plus 2 specifies the number of rectangular slices within the i-th sub-picture. The bricks within the i-th sub-picture are indexed in brick scan order. bottom_right_brick_idx_delta[i][j] and sign_bottom_right_brick_idx_delta[i][j] specify signed delta values that are used to derive the brick index of the bottom-right brick of the j-th rectangular slice within the i-th sub-picture relative to the bottom-right brick of the (j - 1)-th rectangular slice of the i-th sub-picture when j is greater than 0, or relative to 0 (i.e., the top-left corner of the i-th sub-picture) when j is equal to 0. When sign_bottom_right_brick_idx_delta[i][j] is equal to 1, the signed delta value may be defined to be equal to bottom_right_brick_idx_delta[i][j], and when sign_bottom_idx_delta[i][j] is equal to 0 but can also be defined using the opposite assignment of the sign_bottom_right_brick_idx_delta[i][j] value, it is defined to be equal to -bottom_right_brick_idx_delta[i][j].
[0429] Figure 15
[0430] In some applications, it may be reasonable to assign the content to be encoded and / or decoded in such a way that only a subset of the rectangular slices are occupied. For example, in viewport-related streams of 360-degree video, only a subset of the rectangular slices may be received. In another example, patch-based coding of volumetric or point-cloud video is applied, and the patches only occupy a subset of the rectangular slices of the picture.
[0431] According to an embodiment for encoding, an uncoded rectangular slice is indicated in the bitstream or along the bitstream (e.g., in the PPS). The uncoded rectangular slice is not encoded into the bitstream as a VCL NAL unit. The uncoded rectangular slice is reconstructed (e.g., reconstructed as a decoded reference picture) using a predefined or indicated method (such as setting the reconstructed sample values to 0 in the sample array).
[0432] According to an embodiment for decoding, an uncoded rectangular slice is decoded from the bitstream or along the bitstream (e.g., from the PPS). The uncoded rectangular slice is not decoded from the bitstream from a VCL NAL unit. Instead, the uncoded rectangular slice is decoded (e.g., decoded as a decoded reference picture) using a predefined or indicated method (such as setting the decoded sample values to 0 in the sample array).
[0433] According to an embodiment applicable to encoding and / or decoding, a flag is indicated and / or decoded for each rectangular slice, which indicates whether the rectangular slice is uncoded.
[0434] According to an embodiment, the number of uncoded rectangular slices is indicated and / or decoded using a variable-length codeword, such as ue(v). The rectangular slices are traversed in a predefined scan order (in reverse raster scan order of the top-left CTU of the rectangular slice). For each traversed rectangular slice, the flag is indicated and / or decoded to determine whether the rectangular slice is uncoded. If the rectangular slice is uncoded, the remaining number of rectangular slices to be assigned as uncoded is decremented by 1. This process is continued until there are no remaining rectangular slices to be assigned as uncoded.
[0435] FIG. 0 is a graphical representation of an example multimedia communication system in which various embodiments may be implemented. The data source 1510 provides a source signal in analog, uncompressed digital, or compressed digital format, or any combination of these formats. The encoder 1520 may include or be connected with preprocessing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into an encoded media bitstream. It should be noted that the bitstream to be decoded may be received directly or indirectly from a remote device that is actually located within any type of network. Additionally, the bitstream may be received from local hardware or software. The encoder 1520 may be capable of encoding more than one media type, such as audio and video, or may require more than one encoder 1520 to encode source signals of different media types. The encoder 1520 may also receive synthetically generated inputs, such as graphics and text, or it may be capable of generating an encoded bitstream of synthetic media. In the following, the processing of only one encoded media bitstream of one media type is considered to simplify the description. However, it should be noted that, generally, a live broadcast service includes multiple streams (usually at least one audio, video, and text caption stream). It should also be noted that the system may include many encoders, but in the figures, only one encoder 1520 is shown to simplify the description without loss of generality. It should also be understood that although the text and examples contained herein may specifically describe the encoding process, those skilled in the art will understand that the same concepts and principles also apply to the corresponding decoding process, and vice versa.
[0436] The encoded media bitstream can be transferred to a storage device 1530. The storage device 1530 can include any type of mass storage to store the encoded media bitstream. The format of the encoded media bitstream in the storage device 1530 can be a basic stand-alone bitstream format, or one or more encoded media bitstreams can be encapsulated into a container file, or the encoded media bitstream can be encapsulated as a segment format suitable for DASH (or a similar streaming system) and stored as a sequence of segments. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figures) can be used to store one or more media bitstreams in the file and create file format metadata, which can also be stored in the file. The encoder 1520 or the storage device 1530 can include the file generator, or the file generator is operatively attached to the encoder 1520 or the storage device 1530. Some systems operate "in real time", i.e., the storage device is omitted and the encoded media bitstream is transferred directly from the encoder 1520 to the transmitter 1540. Then, the encoded media bitstream can be transferred to the transmitter 1540 (also referred to as a server) as needed. The format used in transmission can be a basic stand-alone bitstream format, a packet stream format, a segment format suitable for DASH (or a similar streaming system), or one or more encoded media bitstreams can be encapsulated into a container file. The encoder 1520, the storage device 1530, and the server 1540 can reside in the same physical device, or they can be included in separate devices. The encoder 1520 and the server 1540 can operate with live content, in which case the encoded media bitstream is generally not stored permanently, but buffered in the content encoder 1520 and / or the server 1540 for a short period of time to smooth out processing delays, transfer delays, and variations in the encoded media bitrate.
[0437] The server 1540 sends the encoded media bitstream using a communication protocol stack. The stack can include, but is not limited to, one or more of the Real-Time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the encoded media bitstream as packets. For example, when RTP is used, the server 1540 encapsulates the encoded media bitstream as RTP packets according to the RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be noted again that the system can include more than one server 1540, but for simplicity, the following description only considers one server 1540.
[0438] If the media content is encapsulated in a container file for storage device 1530 or for inputting data to transmitter 1540, transmitter 1540 may include a "transmission file parser" or be operably attached to a "transmission file parser" (not shown in the figures). Specifically, if the container file is not transmitted as such, and at least one of the included encoded media bitstreams is encapsulated for conveyance via a communication protocol, the transmission file parser locates the appropriate portion of the encoded media bitstream to be conveyed via the communication protocol. The transmission file parser may also assist in creating the correct format of the communication protocol, such as packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as a hint track in ISOBMFF, to encapsulate at least one of the included media bitstreams over the communication protocol.
[0439] Server 1540 may or may not be connected to gateway 1550 via a communication network, which may be, for example, a CDN, the Internet, and / or a combination of one or more access networks. The gateway may also or alternatively be referred to as an intermediate box. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It should be noted that the system may generally include any number of gateways, etc., but for simplicity, the following description only considers one gateway 1550. Gateway 1550 may perform different types of functions, such as translating a packet flow from one communication protocol stack to another, merging and branching of data streams, and manipulating the data stream according to the downlink and / or receiver capabilities, such as controlling the bitrate of the forwarded stream according to the prevailing downlink network conditions. In various embodiments, gateway 1550 may be a server entity.
[0440] The system includes one or more receivers 1560, which are generally capable of receiving, demodulating, and de-encapsulating the transmitted signal into an encoded media bitstream. The encoded media bitstream can be transferred to a recording storage device 1570. The recording storage device 1570 can include any type of mass storage to store the encoded media bitstream. The recording storage device 1570 can alternatively or additionally include computing memory, such as random access memory. The format of the encoded media bitstream in the recording storage device 1570 can be a basic stand-alone bitstream format, or one or more encoded media bitstreams can be encapsulated into a container file. If there are multiple encoded media bitstreams associated with each other, such as an audio stream and a video stream, then generally a container file is used, and the receiver 1560 includes or is attached to a container file generator that generates a container file from the input streams. Some systems operate "in real time", i.e., omit the recording storage device 1570 and transfer the encoded media bitstream directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the recorded stream (e.g., an excerpt of the most recent 10 minutes of the recorded stream) is maintained in the recording storage device 1570, and any older recorded data is discarded from the recording storage device 1570.
[0441] The encoded media bitstream can be transferred from the recording storage device 1570 to the decoder 1580. If there are a number of encoded media bitstreams (such as an audio stream and a video stream) associated with each other and encapsulated into a container file or a single media bitstream is encapsulated in a container file, e.g., for easier access, then a file parser (not shown in the figures) is used to de-encapsulate each encoded media bitstream from the container file. The recording storage device 1570 or the decoder 1580 can include the file parser, or the file parser is attached to the recording storage device 1570 or the decoder 1580. It should also be noted that the system can include a number of decoders, but only one decoder 1570 is discussed here to simplify the description without loss of generality.
[0442] The encoded media bitstream can also be processed by the decoder 1570, and its output is one or more uncompressed media streams. Finally, the renderer 1590 can reproduce the uncompressed media stream using, for example, speakers or a display. The receiver 1560, the recording storage device 1570, the decoder 1570, and the renderer 1590 can reside in the same physical device, or they can be included in separate devices.
[0443] Above, some embodiments have been described with reference to and / or using the terminology of VVC / H.266. It should be understood that the embodiments can be similarly implemented using any video encoder and / or video decoder.
[0444] In the foregoing, some example embodiments have been described with reference to specific syntactic structures and / or syntactic elements. It is to be understood that the embodiments can be implemented similarly using any syntactic structure and / or syntactic element. For example, when an embodiment has been described with reference to a syntactic element in the PPS syntax, it is to be understood that the embodiment can be implemented using the same or similar syntactic elements in another syntactic structure, such as SPS.
[0445] In the foregoing, some embodiments have been described with reference to term indication. It is to be understood that a term indication can be understood as encoding or generating one or more syntactic elements in one or more syntactic structures in or along a bitstream.
[0446] In the foregoing, some embodiments have been described with reference to term decoding. It is to be understood that term decoding can be understood as decoding or parsing one or more syntactic elements from one or more syntactic structures in or along a bitstream.
[0447] In the foregoing, in cases where example embodiments have been described with reference to an encoder, it is to be understood that the resulting bitstream and decoder can have corresponding elements therein. Similarly, in cases where example embodiments have been described with reference to a decoder, it is to be understood that the encoder can have a structure and / or computer program for generating a bitstream to be decoded by the decoder. For example, some embodiments related to generating prediction blocks have been described as part of encoding. By generating prediction blocks as part of decoding, embodiments can be implemented similarly, with the difference that encoding parameters such as horizontal offset and vertical offset are decoded from the bitstream as compared to being determined by the encoder.
[0448] The embodiments of the present invention described above describe a codec in terms of separate encoder and decoder devices to assist in understanding the processes involved. However, it is to be understood that the device, structure, and operation can be implemented as a single encoder-decoder device / structure / operation. Additionally, the encoder and decoder may be able to share some or all common elements.
[0449] Although the above examples describe embodiments of the present invention operating within a codec in an electronic device, it is to be understood that the present invention defined in the claims can be implemented as part of any video codec. Thus, for example, embodiments of the present invention can be implemented in a video codec that can perform video encoding over a fixed or wired communication path.
[0450] Accordingly, a user device can include a video codec such as described above in the embodiments of the present invention. It should be understood that the term user device is intended to cover any suitable type of wireless user device, such as a mobile phone, a portable data processing device, or a portable web browser.
[0451] In addition, the elements of a Public Land Mobile Network (PLMN) may also include the above-mentioned video codec.
[0452] In general, the various embodiments of the present invention may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices, although the present invention is not limited thereto. Although the various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be fully understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.
[0453] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any block in the logical flow of the figures may represent a program step, or an interconnected logical circuit, block, and function, or a combination of program steps and logical circuit, block, and function. The software may be stored on such physical media as memory chips, or in memory blocks implemented within the processor, such as magnetic media like hard disks or floppy disks, and optical media like, for example, DVDs and their data variants such as CDs.
[0454] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and, by way of non-limiting example, may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture.
[0455] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of an integrated circuit is generally a highly automated process. Sophisticated and powerful software tools can be used to transform a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0456] Programs (such as those provided by Synopsys of Mountain View, California and Cadence Design of San Jose, California) use well-established design rules and a library of pre-stored design modules to automatically route conductors and position components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (such as Opus, GDSII, etc.) can be sent to a semiconductor fabrication facility or "fab" for fabrication.
[0457] The foregoing description has provided a complete and informative description of exemplary embodiments of the invention by way of example and not limitation. However, in light of the foregoing description, various modifications and adaptations may become apparent to those skilled in the relevant arts when read in conjunction with the accompanying drawings and the appended claims. However, all such modifications and similar modifications of the teachings of the invention will still fall within the scope of the invention.
Claims
1. A video encoding method, comprising: Performing video encoding processing to form a bitstream, including: Determining the number of units to be assigned to partitions in the bitstream, the units being initially unassigned, where the units correspond to portions of a picture in the video; Indicating or inferring in the bitstream the number of partitions of an explicit size of a tile column to be assigned; Indicating in the bitstream the size of the partitions of the explicit size of the tile column, and correspondingly marking the unassigned units as being assigned to partitions in a predefined scan order; Indicating in the bitstream the width of tile columns of equal size; Repeatedly assigning the width of the tile columns of equal size to partitions among the units remaining after assigning the unassigned units to the partitions of the explicit size of the tile column using the predefined scan order, and marking the corresponding unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; and Assigning the unassigned units to a final partition when the number of unassigned units is greater than 0; Wherein the units are coding tree blocks, and wherein the partitions are indicated at least by the tile width in units of the coding tree blocks.
2. A device for video encoding, comprising: Means for performing video encoding processing to form a bitstream, including: Means for determining the number of units to be assigned from the bitstream to partitions, the units being initially unassigned, where the units correspond to portions of a picture of the video; Means for indicating or inferring in the bitstream the number of partitions of an explicit size of a tile column to be assigned; Means for indicating in the bitstream the size of the partitions of the explicit size of the tile column and means for correspondingly marking the unassigned units as being assigned to partitions in a predefined scan order; Means for indicating in the bitstream the width of tile columns of equal size; Means for repeatedly assigning the width of the tile columns of equal size to partitions among the units remaining after assigning the unassigned units to the partitions of the explicit size of the tile column using the predefined scan order and means for marking the corresponding unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; and Means for assigning the unassigned units to a final partition when the number of unassigned units is greater than 0; Wherein the units are coding tree blocks, and wherein the partitions are indicated at least by the tile width in units of the coding tree blocks.
3. A video decoding method, comprising: Using a bitstream to perform video decoding processing to decode a picture, including: Determining the number of units to be assigned from the bitstream or along with the bitstream to partitions, where the units correspond to portions from a picture in the video, the video being embedded in the bitstream; Determining at least using the bitstream from the bitstream or along with the bitstream the number of partitions of an explicit size of a tile column to be assigned; Determine the size of the explicitly sized partitions for the tile column using at least the bitstream, and accordingly mark the unassigned units as assigned to the partitions in a predefined scan order; Determine the width of the tile columns of equal size using at least the bitstream; In the units remaining after assigning the unassigned units to the explicitly sized partitions of the tile column using the predefined scan order, repeatedly assign the width of the tile columns of equal size to partitions, and mark the corresponding unassigned units as assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; and When the number of unassigned units is greater than 0, assign the unassigned units to the last partition; wherein the units are coding tree blocks, and wherein the partitions are indicated at least by the tile width in units of the coding tree blocks.
4. The method according to claim 3, wherein The number of partitions that determine the explicit size of the tile columns to be assigned includes: Decode the number of the explicitly sized partitions of the tile column to be assigned from the syntax structure; Determining the size of the explicitly sized partitions for the tile column includes: decoding the size of the explicitly sized partitions for the tile column from the syntax structure; and Determining the width of the tile columns of equal size includes: decoding the width of the tile columns of equal size from the syntax structure.
5. An apparatus for video decoding to form a bitstream, comprising: Apparatus for performing video decoding processing using a bitstream to decode a picture, comprising: Components for determining the number of units to be assigned from the bitstream or along with the bitstream to partitions, wherein the units correspond to portions from pictures in a video, the video being embedded in the bitstream; Components for determining, using at least the bitstream, the number of explicitly sized partitions of a tile column to be assigned from the bitstream or along with the bitstream; Components for determining the size of the explicitly sized partitions for the tile column using at least the bitstream and for marking the corresponding unassigned units as assigned to the partitions in a predefined scan order; Components for determining the width of the tile columns of equal size using at least the bitstream; Components for repeatedly assigning the width of the tile columns of equal size to partitions in the units remaining after assigning the unassigned units to the explicitly sized partitions of the tile column using the predefined scan order and for correspondingly marking the unassigned units as assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; and Components for assigning the unassigned units to the last partition when the number of unassigned units is greater than 0; wherein the units are coding tree blocks, and wherein the partitions are indicated at least by the tile width in units of the coding tree blocks.
6. A computer program product comprising at least one non-transitory computer-readable storage medium having computer-executable program code instructions stored therein, the computer-executable program code instructions including program code instructions configured to, upon execution: Perform video encoding processing to form a bitstream, including: Determine the number of units to be assigned to partitions in the bitstream, the units being initially unassigned, where the units correspond to portions of a picture in the video; Indicate or infer in the bitstream the number of partitions of an explicit size of a tile column to be assigned; Indicate in the bitstream the size of the partitions of the explicit size for the tile column, and accordingly mark the unassigned units as being assigned to partitions in a predefined scan order; Indicate in the bitstream the width of tile columns of equal size; In the remaining units after assigning the unassigned units to the partitions of the explicit size of the tile column using the predefined scan order, repeatedly assign the width of the tile columns of equal size to partitions, and mark the corresponding unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; And When the number of unassigned units is greater than 0, assign the unassigned units to the last partition; Wherein the units are coding tree blocks, and wherein the partition is indicated at least by a tile width in units of the coding tree blocks.
7. A computer program product comprising at least one non-transitory computer-readable storage medium having computer-executable program code instructions stored therein, the computer-executable program code instructions including program code instructions configured to, upon execution: Perform video decoding processing using the bitstream to decode a picture, including: Determine the number of units to be assigned from the bitstream or as the bitstream is assigned to partitions, where the units correspond to portions of a picture from a video, the video being embedded in the bitstream; Determine at least using the bitstream from the bitstream or as the bitstream the number of partitions of an explicit size of a tile column to be assigned; Determine at least using the bitstream the size of the partitions of the explicit size for the tile column, and accordingly mark the unassigned units as being assigned to partitions in a predefined scan order; Determine at least using the bitstream the width of tile columns of equal size; In the remaining units after assigning the unassigned units to the partitions of the explicit size of the tile column using the predefined scan order, repeatedly assign the width of the tile columns of equal size to partitions, and mark the corresponding unassigned units as being assigned in the predefined scan order until the number of unassigned units is less than the width of the tile columns of equal size; And When the number of unassigned units is greater than 0, assign the unassigned units to the last partition; wherein the unit is a coding tree block, and wherein the partitioning is indicated by at least a tile width in units of the coding tree block.
Citation Information
Patent Citations
Apparatus and method for video coding and decoding, and computer program
CN113940075A
Sub-Pictures for Pixel Rate Balancing on Multi-Core Platforms
US20130202051A1