Video coding with subpicture, slice, and tile support
The method addresses the complexity in signaling picture partitioning in VVC7 by determining slice parameters within sub-pictures, enhancing efficiency and compliance in image encoding and decoding.
Patent Information
- Application Number
- JP2023136525
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2023-08-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-17
AI Technical Summary
The existing video coding standards, such as VVC7, face complexity and inefficiency in signaling picture partitioning, particularly due to the independence of sub-picture partitioning from slice and tile grid signaling, leading to unnecessary resource consumption and potential compliance issues.
A method and apparatus for improving the signaling of picture partitioning by determining parameters associated with slices within sub-pictures using sequence parameters, and using these parameters for decoding and encoding images, thereby simplifying the processing and ensuring compliance with VVC7 constraints.
The proposed solution optimizes the signaling of picture partitioning, reducing unnecessary resource consumption and ensuring compliance with VVC7 standards, thereby improving the efficiency and effectiveness of image encoding and decoding processes.
Smart Images

Figure 0007675768000023 
Figure 0007675768000024 
Figure 0007675768000025
Abstract
Description
[Technical field]
[0001] The present invention relates to partitioning of images and encoding or decoding of an image or a sequence of images comprising images. Embodiments of the invention are of particular, but not exclusive, use when encoding or decoding a sequence of images using a first partitioning of the image into one or more sub-pictures and a second partitioning of the image into one or more slices. [Background technology]
[0002] Video coding includes image coding, where an image corresponds to a single frame of a video or picture. In video coding, before some coding tools, such as motion compensation / prediction (e.g., inter-prediction) or intra-prediction, can be used on the image, the image is first partitioned (e.g., divided) into one or more image parts so that the coding tools can be used and applied on the image parts. The present invention particularly relates to the partitioning of images into two types of image parts, sub-pictures and slices, which are being studied by the Video Coding Experts Group / Moving Picture Experts Group (VCEG / MPEG) standardization group and are being considered for use in the Generic Video Coding (VVC) standard.
[0003] Subpictures are a new concept introduced in VVC to enable bitstream extraction and merging operations of independent spatial regions (or image parts) from different bitstreams. "Independent" here means that those regions (or image parts) are coded / decoded without reference to information obtained from coding / decoding another region or image part. For example, independent spatial regions (i.e., regions or image parts that are coded / decoded without reference to coding / decoding another region / image part from the same image) are used for region of interest (ROI) streaming (e.g., during 3D video streaming) or for streaming of omnidirectional video content (e.g., a sequence of images being streamed using the Omnidirectional MediA Format (OMAF) standard), especially when viewport-dependent streaming techniques are used for streaming. Each image from the omnidirectional video content is divided into independent regions coded in its different versions (e.g., in terms of image quality or resolution). A client terminal (e.g., a device with a display such as a mobile phone) can then select an appropriate version of the independent region to obtain a high-quality version of the independent region in the main line of sight direction, while still being able to use the remaining region of the lower-quality version for the remaining part of the omnidirectional video content to improve encoding efficiency.
[0004] High Efficiency Video Coding (HEVC or H.265) provides motion constrained tile set signaling to indicate independently coded regions (e.g., the bitstream includes data to specify or determine a tile set that has its motion prediction constrained to be "independent" from another region of the image). In HEVC, this signaling is done in a Supplemental Enhancement Information (SEI) message and is only optional. However, in HEVC, slice signaling is done independently from this SEI message, so that the partitioning of an image into one or more slices is defined independently from the partitioning of the same image into one or more tile sets. This means that a slice partitioning does not have the same motion prediction constraints imposed on it.
[0005] Proposals based on older drafts of Generic Video Coding Draft 4 (VVC4) included signaling tile group partitioning dependent on sub-picture signaling. A tile group is an integer number of complete tiles of a picture that are exclusively contained in a single Network Abstraction Layer (NAL) unit. This proposal (JVET-N0107: AHG12: Sub-picture-based coding for VVC, Huawei) was a syntax change proposal to introduce the sub-picture concept for VVC. Sub-picture positions are signaled using luma sample positions in the Sequence Parameter Set (SPS). A flag in the SPS then indicates for each sub-picture whether tile group partitioning (i.e., partitioning of a picture into one or more tile groups) is signaled in the Picture Parameter Set (PPS), with motion prediction being constrained, and each PPS being defined for each sub-picture. Since a PPS is provided per sub-picture, the tile group partitioning is signaled per sub-picture in JVET-N0107.
[0006] However, the latest Versatile Video Coding Draft 7 (VVC7) no longer has this tile group partitioning concept. VVC7 signals the sub-picture layout in the CTU unit in the SPS. A flag in the SPS indicates whether motion prediction is constrained for the sub-picture. These SPS syntax elements are:
[0007] [Table 1]
[0008] In VVC7, slice partitioning is defined in the PPS based on tile partitions as follows:
[0009] [Table 2]
[0010] This means that slice partitioning is defined independently of subpicture partitioning in VVC7. The VVC7 syntax for slice partitioning is done on top of the tile structure without reference to subpictures, since this independence of subpictures avoids any specific processing for subpictures during the encoding / decoding process, resulting in simpler processing for subpictures. Summary of the Invention
[0011] As mentioned above, VVC7 provides several tools for partitioning a picture into regions of pixels (or component samples). Some examples of these tools are sub-pictures, slices and tiles. To accommodate all these tools while maintaining their functionality, VVC7 imposes several constraints on the partitioning of a picture into these regions. For example, tiles must be rectangular and tiles must form a grid. A slice can be either an integer number of tiles or a fraction of a tile (i.e., a slice contains only a portion of a tile, or a "partial tile" or "fraction tile"). A sub-picture is a rectangular region that must contain one or more slices. However, in VVC7, the signaling of sub-picture partitioning is independent of slice and tile grid signaling. Thus, this signaling in VVC7 requires the decoder to check and ensure that the picture partitioning complies with the VVC7 constraints, which is complex and may lead to unnecessary time or resource consumption on the decoder side.
[0012] An object of embodiments of the present invention is to address one or more problems or shortcomings of the partitioning of said images and the encoding or decoding of an image or a sequence of images including said images. For example, one or more embodiments of the present invention aim to improve and optimize the signaling of picture partitioning (e.g., within a VVC7 context) while ensuring that at least some of the constraints that require checking in VVC7 are / are satisfied by design during the signaling or encoding process.
[0013] According to aspects of the invention, there are provided an apparatus / device, a method, a program, a computer readable storage medium and a carrier medium / signal as set forth in the appended claims. Other features of the invention will become apparent from the dependent claims and the description. According to other aspects of the invention, there are provided a system, a method for controlling such a system, an apparatus / device for performing the method as set forth in the appended claims, an apparatus / device for processing, a media storage device storing a signal as set forth in the appended claims, a computer readable storage medium or a non-transitory computer readable storage medium storing a program as set forth in the appended claims, and a bitstream generated using the encoding method as set forth in the appended claims. Other features of the invention will become apparent from the dependent claims and the following description. According to one aspect of the present invention, there is provided a method for decoding image data, the image may include one or more slices that may correspond to an integer number of consecutive complete coding tree unit rows in a tile, the image may include one or more sub-pictures, the method including: obtaining first information indicating a width of a sub-picture and second information indicating a height of the sub-picture from a sequence parameter set; determining parameters related to the slices included in the sub-picture using the first information and the second information; and decoding the image using at least the determined parameters, wherein at least intra prediction is used in decoding the image. According to one aspect of the present invention, there is provided a method for encoding an image, the image may include one or more slices that may correspond to an integer number of consecutive complete coding tree unit rows in a tile, and the image may include one or more sub-pictures, the method including: encoding first information indicating a width of a sub-picture and second information indicating a height of the sub-picture into a sequence parameter set; determining parameters related to the slices included in the sub-picture using the first information and the second information; and encoding the image using at least the determined parameters, wherein at least intra prediction is used in encoding the image. According to one aspect of the present invention, there is provided an apparatus for decoding image data, the image may include one or more slices that may correspond to an integer number of consecutive complete coding tree unit rows in a tile, and the image may include one or more sub-pictures, the apparatus including: an acquisition means for acquiring first information indicating a width of a sub-picture and second information indicating a height of the sub-picture from a sequence parameter set; a determination means for determining parameters related to the slice included in the sub-picture using the first information and the second information; and a decoding means for decoding the image using at least the determined parameters, the decoding means being characterized in that in decoding the image, the decoding means using at least intra prediction. According to one aspect of the present invention, there is provided an apparatus for encoding an image, the image may include one or more slices that may correspond to an integer number of consecutive complete coding tree unit rows in a tile, and the image may include one or more sub-pictures, the apparatus including: a first encoding means for encoding first information indicating a width of the sub-picture and second information indicating a height of the sub-picture into a sequence parameter set; a determination means for determining parameters related to the slice included in the sub-picture using the first information and the second information; and a second encoding means for encoding the image using at least the determined parameters, the second encoding means being characterized in that in encoding the image, the second encoding means using at least intra prediction.
[0014] According to a first aspect of the present invention, there is provided a method of processing image data of one or more images, each image consisting of one or more tiles and divisible into one or more image parts, and an image is divisible into one or more sub-pictures, the method comprising determining one or more image parts comprised in a sub-picture, and processing the one or more images using information obtained from the determination.
[0015] According to a second aspect of the present invention there is provided a method of partitioning one or more images comprising partitioning an image into one or more tiles, partitioning said image into one or more sub-pictures, and partitioning said image into one or more image portions by processing image data of the image according to the first aspect.
[0016] According to a third aspect of the present invention there is provided a method of signalling partitioning of one or more images, the method comprising processing image data of the one or more images according to the first aspect and signalling information for determining the partitioning in a bitstream.
[0017] With respect to the aforementioned aspects of the invention, the following features may be provided in accordance with embodiments of the invention: Preferably, an image portion may include a partial tile; Suitably, an image portion is encoded into or decoded from (e.g., signaled to, communicated to, provided to, or obtained from) a single logical unit (e.g., one network abstraction layer unit or one NAL unit); Suitably, tiles and / or sub-pictures are not encoded into or decoded from (e.g., signaled to, communicated to, provided to, or obtained from) a single logical unit (e.g., one NAL unit).
[0018] According to a fourth aspect of the present invention, there is provided a method of processing image data of one or more images, each image consisting of one or more tiles and divisible into one or more image parts, an image part may comprise a part of a tile (partial tile), an image is divisible into one or more sub-pictures, the method comprising determining one or more image parts comprised in a sub-picture, and processing the one or more images using information obtained from the determination, where a part of a tile (partial tile) is an integer number of consecutive complete coding tree unit (CTU) rows within the tile.
[0019] With respect to the aforementioned aspects of the invention, the following features may be provided in accordance with embodiments of the invention: Suitably, the determining comprises defining the one or more image portions using one or more of: an identifier of the sub-picture; a size, width or height of the sub-picture; whether only a single image portion is included in the sub-picture; and a number of image portions included in the sub-picture.
[0020] Preferably, when a sub-picture contains more than one image portion, each image portion is determined based on the number of tiles contained within it.
[0021] Preferably, if an image portion comprises one or more parts of a tile (partial tiles), said image portion is determined based on the number of rows or columns of coding tree units CTU to be contained therein.
[0022] Suitably, the processing comprises providing or obtaining in or from a picture parameter set PPS information for determining the image portion based on the number of tiles, and, if an image portion comprises one or more parts of a tile (partial tiles), providing or obtaining in or from a header of one or more logical units containing encoded data of said image portion information for identifying said image portion comprising one or more parts of a tile (partial tiles).
[0023] Preferably, an image portion consists of a sequence of tiles in tile raster scan order.
[0024] Suitably, the processing includes providing to or obtaining from the bitstream information for determining one or more of: whether only a single image portion is included in the sub-picture; the number of image portions included in the sub-picture; Suitably, the processing includes providing to or obtaining from the bitstream whether use of an image portion comprising part of a tile (partial tile) is permitted when processing one or more images.
[0025] Suitably, the information provided in or obtained from the bitstream includes information indicating whether or not sub-pictures are used in the video sequence, and one or more image portions of the video sequence are determined to be not permitted to include part of a tile (a partial tile) if the information indicates that sub-pictures are not used in the video sequence.
[0026] Suitably, the information for determining is provided in or obtained from the picture parameter set PPS.
[0027] Suitably, the information for determining is provided in or obtained from the sequence parameter set SPS.
[0028] Preferably, if the information for the determination indicates that the number of image portions included in the sub-picture is one, the sub-picture consists of a single image portion that does not include a portion of a tile (partial tile).
[0029] Preferably, a sub-picture comprises two or more image portions, each image portion comprising one or more portions of a tile (partial tiles).
[0030] Preferably, one or more portions of a tile (partial tiles) are from the same single tile.
[0031] Preferably, the two or more image portions may include one or more portions of tiles (partial tiles) from two or more tiles.
[0032] Preferably, an image portion consists of a number of tiles, said image portion forming a rectangular area within the image.
[0033] Suitably, an image portion is a slice (one or more image portions being one or more slices).
[0034] According to a fifth aspect of the present invention there is provided a method of encoding one or more images, the method comprising either processing image data according to the first or fourth aspect, partitioning according to the second aspect and / or signalling according to the third aspect.
[0035] Suitably, the method further comprises receiving an image, processing image data of the received image in accordance with the first or fourth aspect, encoding the received image, and generating a bitstream.
[0036] Suitably, the method further comprises providing, in the bitstream, one or more of: information for determining an image portion based on a number of tiles in a picture parameter set PPS and, if an image portion comprises one or more portions of a tile (partial tiles), information for identifying the image portion comprising one or more portions of a tile (partial tiles) in a header of one or more logical units comprising coded data of said image portion, a slice segment header, or a slice header; information in the PPS for determining whether only a single image portion is comprised in a sub-picture; information in the PPS for determining the number of image portions comprised in a sub-picture; information in the sequence parameter set SPS for determining whether use of image portions comprising portions of tiles (partial tiles) is allowed when processing one or more images; and information in the SPS indicating whether sub-pictures are used in the video sequence.
[0037] According to a sixth aspect of the present invention there is provided a method of decoding one or more images, the method comprising either processing image data according to the first or fourth aspect, partitioning according to the second aspect and / or signalling according to the third aspect.
[0038] Preferably, the method further comprises receiving a bitstream, decoding information from the received bitstream, processing image data according to either the first or fourth aspect, and obtaining an image using the decoded information and the processed image data.
[0039] Suitably, the method further comprises obtaining from the bitstream one or more of the following: information for determining the image portion based on the number of tiles from a picture parameter set PPS and, if the image portion comprises one or more portions of a tile (partial tiles), information for identifying the image portion comprising one or more portions of a tile (partial tiles) from a header of one or more logical units comprising coded data of said image portion, a slice segment header, or a slice header; information from the PPS for determining whether only a single image portion is comprised in a sub-picture; information from the PPS for determining the number of image portions comprised in a sub-picture; information from a sequence parameter set SPS for determining whether use of an image portion comprising a portion of a tile (partial tile) is allowed when processing one or more images; and information indicating whether sub-pictures are used in the video sequence from the SPS.
[0040] According to a seventh aspect of the present invention there is provided a device for processing image data of one or more images configured to perform a method according to any of the first, fourth, second or third aspects.
[0041] According to an eighth aspect of the present invention there is provided a device for encoding one or more images comprising a processing device according to the seventh aspect. Suitably the device is arranged to carry out a method according to the fifth aspect.
[0042] According to a ninth aspect of the present invention there is provided a device for decoding one or more images comprising a processing device according to the seventh aspect. Suitably the device is arranged to perform a method according to the sixth aspect.
[0043] According to a tenth aspect of the present invention there is provided a program which, when executed on a computer or processor, causes the computer or processor to carry out a method according to the first aspect, the fourth aspect, the second aspect or the third aspect, the fifth aspect or the sixth aspect.
[0044] According to an eleventh aspect of the present invention, there is provided a carrier medium or computer readable storage medium carrying / storing the program of the tenth aspect.
[0045] According to a twelfth aspect of the present invention there is provided a signal carrying an information dataset for an image encoded using the method according to the fifth aspect and represented by a bitstream, the image consisting of one or more tiles and being divisible into one or more image parts, an image part may comprise parts of a tile (partial tiles), the image being divisible into one or more sub-pictures, the information dataset comprising data for determining one or more image parts comprised in a sub-picture.
[0046] Yet another aspect of the invention relates to a program which, when executed by a computer or processor, causes the computer or processor to perform any of the methods of the previous aspects. The program may be provided by itself or may be carried on, by or within a carrier medium. The carrier medium may be non-transitory, for example a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, for example a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet.
[0047] Yet another aspect of the invention relates to a camera comprising a device according to any of the aforementioned device aspects. According to yet another aspect of the invention there is provided a mobile device comprising a device according to any of the aforementioned device aspects and / or a camera embodying the aforementioned camera aspects.
[0048] Any feature in one aspect of the invention may be applied to other aspects of the invention in any suitable combination. In particular, a method aspect may be applied to an apparatus aspect, and vice versa. Furthermore, a feature implemented in hardware may be implemented in software, and vice versa. References herein to software and hardware features should be interpreted accordingly. Any apparatus feature as described herein may be provided as a method feature, and vice versa. As used herein, means-plus-function features may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be understood that a particular combination of the various features described and defined in any aspect of the invention may be implemented and / or provided and / or used independently.
[0049] Further features, aspects, and advantages of the present invention will become apparent from the following description of the embodiments with reference to the accompanying drawings. Each of the embodiments of the present invention described below may be realized alone or in combination with several embodiments. Also, features from various embodiments may be combined where necessary or where a combination of elements or features from the individual embodiments in a single embodiment is beneficial. [Brief description of the drawings]
[0050] Embodiments of the invention will now be described, by way of example only, with reference to the following drawings: [Figure 1] FIG. 1 illustrates partitioning a picture into tiles and slices according to one embodiment of the present invention. [Diagram 2] FIG. 2 illustrates sub-picture partitioning of a picture according to one embodiment of the present invention. [Diagram 3] FIG. 3 illustrates a bitstream according to one embodiment of the present invention. [Figure 4] FIG. 4 is a flow chart illustrating an encoding process according to one embodiment of the present invention. [Diagram 5] FIG. 5 is a flow chart illustrating a decoding process according to one embodiment of the present invention. [Figure 6] FIG. 6 is a flow chart illustrating the decision steps used in signaling slice partitioning according to one embodiment of the present invention. [Figure 7] FIG. 7 illustrates an example of sub-picture and slice partitioning according to one embodiment of the present invention. [Figure 8] FIG. 8 illustrates an example of sub-picture and slice partitioning according to one embodiment of the present invention. [Figure 9a] FIG. 9a is a flow chart illustrating steps of an encoding method according to an embodiment of the invention. [Figure 9b] FIG. 9b is a flow chart illustrating steps of a decoding method according to an embodiment of the invention. [Figure 10] FIG. 10 is a block diagram illustrating steps of an encoding method according to an embodiment of the invention. [Figure 11] FIG. 11 is a block diagram illustrating steps of a decoding method according to an embodiment of the present invention. [Figure 12] FIG. 12 is a block diagram that illustrates generally a data communications system in which one or more embodiments of the present invention may be implemented. [Figure 13] FIG. 13 is a block diagram illustrating components of a processing device capable of implementing one or more embodiments of the present invention. [Figure 14] FIG. 14 is a diagram illustrating a network camera system in which one or more embodiments of the present invention can be implemented. [Figure 15] FIG. 15 illustrates a smartphone in which one or more embodiments of the present invention can be implemented. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0051] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments of the present invention described below relate to improved encoding and decoding of images (or pictures).
[0052] As used herein, "signaling" can refer to inserting (providing / including / encoding) into or extracting / obtaining (decoding) from a bitstream information regarding one or more parameters or syntax elements, such as information for determining any one or more of an identifier for a subpicture, a size / width / height of a subpicture, whether only a single image portion (e.g., a slice) is included in the subpicture, whether the slice is a rectangular slice, and / or the number of slices included in the subpicture. As used herein, "processing" can refer to any type of operation performed on data, such as encoding or decoding image data for one or more images / pictures.
[0053] In this specification, the term "slice" is used as an example of an image portion (another example of such an image portion is an image portion including one or more coding tree units). It is understood that embodiments of the present invention may be implemented based on image portions instead of slices, and appropriately modified parameters / values / syntax, such as headers for image portions (instead of slice headers or slice segment headers). It is also understood that various information described herein as being signaled in a slice header, slice segment header, sequence parameter set (SPS), or picture parameter set (PPS) may be signaled elsewhere, so long as they can provide the same functionality as provided by signaling in those media. It is also understood that any of the following may be referred to as an image portion: slice, tile group, tile, coding tree unit (CTU) / largest coding unit (LCU), coding tree block (CTB), coding unit (CU), prediction unit (PU), transform unit (TU), or block of pixels / samples.
[0054] When a component or tool is described as "active", the component / tool is "enabled" or "available for use" or "in use", when described as "inactive", the component / tool is "disabled" or "unavailable for use" or "not being used", and "inferable" refers to the ability to determine / obtain the associated value or parameter from other information without being explicitly signaled in the bitstream. Additionally, when a flag is described as "active", it is also understood to mean that the flag indicates that the associated component / tool is "active" (i.e., "enabled").
[0055] In this specification, the following terms are used to define the same or functionally equivalent terms as defined in VVC7 unless otherwise specified. The definitions used in VVC7 are shown below.
[0056] Slice: An integer number of contiguous complete CTU rows within a tile of a picture, or an integer number of complete tiles, contained exclusively in one NAL unit.
[0057] Slice Header: The part of an encoded slice that contains data elements pertaining to all tiles or CTU rows within the tiles represented in the slice.
[0058] Tile: A rectangular area of CTUs within a particular tile column and a particular tile row in a picture.
[0059] Subpicture: A rectangular area of one or more slices within a picture.
[0060] Picture (or image): An array of luma samples for monochrome formats, or an array of luma samples and two corresponding arrays of chroma samples for 4:2:0, 4:2:2, and 4:4:4 color formats.
[0061] coded picture: A coded representation of a picture that contains a VCL NAL unit with a particular value of nuh_layer_id in an AU and includes all CTUs of the picture.
[0062] Encoded Representation: A data element that is represented in encoded form.
[0063] Raster Scan: The mapping of a rectangular 2-D pattern onto a 1-D pattern such that the first entry of the 1-D pattern is from the first, top row of the 2-D pattern scanned from left to right, followed similarly by the second, third, etc. rows (going down) of the pattern scanned left to right respectively.
[0064] Block: An MxN (M columns by N rows) array of samples, or an MxN array of transform coefficients.
[0065] Coding block: An MxN block of samples for some values of M and N, such that dividing the CTB into coding blocks is partitioning.
[0066] Coding Tree Block (CTB): An N×N block of samples for some value of N, such that the division of the components into the CTB is a partitioning.
[0067] Coding Tree Unit (CTU): A CTB for luma samples, two corresponding CTBs for chroma samples for a picture with three sample arrays, or a CTB for samples for a monochrome picture, or a picture that is coded using three separate color planes and the syntax structure used to code the samples.
[0068] Coding Unit (CU): A coding block of luma samples, and two corresponding coding blocks of chroma samples for a picture with three sample arrays, or a coding block of samples for a monochrome picture, or a picture that is coded using three separate color planes and syntax structures used to code the samples.
[0069] Component: An array or a single sample from one of the three arrays (luma and two chroma) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color formats, or an array or a single sample of the arrays that make up a picture in monochrome format.
[0070] Picture Parameter Set (PPS): A syntax structure containing syntax elements that apply to zero or more entire coded pictures, as determined by syntax elements in each slice header.
[0071] Sequence Parameter Set (SPS): A syntax structure containing zero or more CVS-wide syntax elements, determined by the contents of syntax elements in the PPS referenced by syntax elements in each slice header.
[0072] As used herein, the following terms are also used to define the same or functionally equivalent as defined below, unless stated otherwise.
[0073] Tile group: An integral number of complete (i.e., entire) tiles of a picture that are contained exclusively in a single NAL unit.
[0074] "Tile fraction", "partial tile", "part of a tile", or "fraction of a tile": an integer number of contiguous complete CTU rows within a tile of a picture that do not form a complete (i.e., entire) tile.
[0075] slice segment: an integer number of contiguous complete CTU rows within a tile of a picture, or an integer number of complete tiles, that are exclusively contained in one NAL unit.
[0076] Slice Segment Header: The part of a coded slice segment that contains data elements pertaining to all tiles or CTU rows within the tiles represented in the slice segment.
[0077] Slice, when slice segments are present: A set of one or more slice segments that collectively represent an integer number of complete tiles or tiles of a picture. EMBODIMENTS OF THE PRESENT DISCLOSURE Picture / Image and Bitstream Partitioning 3.1 Partitioning a Picture into Tiles and Slices Video compression relies on block-based video coding in most coding systems such as HEVC or the emerging VVC standard. In these coding systems, video is composed of a sequence of frames or pictures or images or samples that may be displayed at different times (e.g., at different temporal positions within the video). In the case of multi-layered video (e.g., scalable, stereo, or 3D video), several pictures may need to be decoded to be able to form the final / result image that is displayed at a particular time. A picture may also be composed of two or more image components (i.e., the image data of a picture includes two or more image components). An example of such an image component would be a component for encoding luma, chroma, or depth information.
[0078] Compression of video sequences uses several different partitioning techniques (i.e., different schemes / frameworks / arrangements / mechanisms for partitioning / diving pictures) for each picture and how these partitioning techniques are implemented during the compression process.
[0079] Figure 1 illustrates a partitioning of a picture into tiles and slices according to an embodiment of the present invention that is compatible with VVC7. Pictures 101, 102 are divided into coding tree units (CTUs), shown by dotted lines. A CTU is the basic unit of encoding and decoding in VVC7. For example, in VVC7, a CTU can encode an area of 128x128 pixels.
[0080] A coding tree unit (CTU) may also be called a block (of pixels or component samples (values)), a macroblock, or a coding block. It may be used to simultaneously code / decode different image components of a picture, or it may be limited to only one image component, so that the different image components of a picture can be coded / decoded separately / individually. If the data of an image contains separate data for each component, the CTU groups multiple coding tree blocks (CTBs), one CTB for each component.
[0081] As shown in Fig. 1, a picture can also be partitioned according to a grid of tiles (i.e., into one or more grids of tiles) represented by thin solid lines. A tile is a picture portion (part / portion of a picture) that is a rectangular area (of pixels / component samples) that can be defined independently of the CTU partitioning. A tile can also correspond to a sequence of CTUs, e.g. in VVC7, as in the example shown in Fig. 1, and a partitioning technique can constrain the tile boundaries to coincide / align with the boundaries of the CTUs.
[0082] Tiles are defined such that tile boundaries break spatial dependencies of the encoding / decoding process (i.e., in a given picture, tiles are defined / specified such that they can be encoded / decoded independently from other spatially "adjacent" tiles of the same picture. This means that the encoding / decoding of CTUs within a tile is not based on pixels / samples or reference data from other tiles in the same picture.
[0083] Some encoding / decoding systems, for example the systems for embodiments of the present invention or VVC7, provide the concept of slices (i.e. also use a partitioning technique based on one or more slices). This mechanism allows partitioning a picture into one or several groups of tiles, collectively called slices. Each slice is composed of one or more tiles or sub-tiles. Two different types of slices are provided as shown by pictures 101 and 102. The first type of slices is restricted to slices forming a rectangular area / region in the picture, as represented by the thick solid lines in picture 101. Picture 101 illustrates the partitioning of the picture into six different rectangular slices (0)-(5). The second type of slices is restricted to consecutive tiles in raster scan order (so that slices form a sequence of tiles), as represented by the thick solid lines in picture 102. Picture 102 illustrates the partitioning of the picture into three different slices (0)-(2) composed of consecutive tiles in raster scan order. Rectangular slices are often a structure / arrangement / configuration of choice for processing a region of interest (RoI) in a video. A slice can be coded into (or decoded from) a bitstream as one or more Network Abstraction Layer (NAL) units. A NAL unit is a logical unit of data for encapsulation of data in the coding / decoding bitstream (e.g., a packet containing an integer number of bytes, where multiple packets together form the coded video data). In the VVC7 coding / decoding system, a slice is typically coded as a single NAL unit. If a slice is coded as several NAL units in the bitstream, each NAL unit for the slice is called a slice segment. A slice segment includes a slice segment header that includes the coding parameters of that slice segment. According to a variant, the header of the first slice segment NAL unit of a slice includes all the coding parameters for the slice.The slice segment headers of subsequent NAL units of a slice may contain fewer parameters than the first NAL unit, in which case the first slice segment is an independent slice segment and the subsequent segments are dependent slice segments (because they depend on the coding parameters from the NAL unit of the first slice segment).
[0084] 3.2 Partitioning into Subpictures 2 illustrates a sub-picture partitioning of a picture, i.e. into one or more sub-pictures, according to an embodiment of the present invention. A sub-picture represents a picture portion (a part or a portion of a picture) covering a rectangular area of the picture. Each sub-picture can have different size and coding parameters than another sub-picture. The sub-picture layout, i.e. the geometry of a sub-picture within a picture (e.g. as defined using the position and dimensions / width / height of the sub-picture), allows for grouping a set of slices of a picture and can constrain (i.e. impose constraints on) the temporal motion prediction between two pictures.
[0085] In FIG. 2, the tile partitioning of picture 201 is a 4×5 tile grid. The slice partitioning defines 24 slices, one slice per tile, except for the last tile column on the right side, where each tile is divided into two slices (i.e., the slice contains a partial tile). Picture 201 is also divided into two sub-pictures 202 and 203. A sub-picture is defined as one or more slices forming a rectangular area. Sub-picture 202 (shown as a dotted area) includes slices from the first three tile columns (starting from the left) and sub-picture 203 (shown as a hatched area with diagonal lines through them) for the remaining slices (the last two tile columns on the right side). As shown in FIG. 2, VVC7 and embodiments of the present invention provide a partitioning scheme that allows to define single slice partitioning and single tile partitioning at the picture level (e.g., per picture using syntax elements provided in the PPS). Sub-picture partitioning is applied on top of tile and slice partitioning. Another aspect of subpictures is that each subpicture is associated with a set of flags. This allows for indicating (using one or more of the set of flags) that temporal prediction is constrained to use data from reference frames that are part of the same subpicture (e.g., temporal prediction is constrained such that a predictor of a subpicture cannot use reference data from another subpicture). For example, referring to FIG. 2, CTB 204 belongs to subpicture 202. When temporal prediction is indicated as constrained for subpicture 202, temporal prediction for subpicture 202 cannot use reference blocks (or reference data) coming from subpicture 203. As a result, slices of subpicture 202 are codeable / decodable independently of slices of subpicture 203. This property / attribute / characteristic / capability is useful in viewport-dependent streaming, which involves segmenting an omnidirectional video sequence into spatial parts, each spatial part representing a particular viewing direction of the 360 (degree) content.The viewer can then select the segment that corresponds to the desired / relevant viewing direction, and this property of the subpictures can be used to encode / decode the segment without having to access data from the rest of the 360 content.
[0086] Another use of sub-pictures is to generate streams with regions of interest. Sub-pictures provide a spatial representation for these regions of interest that can be coded / decoded independently. Sub-pictures are designed to allow / enable easy access to the coded data corresponding to these regions. As a result, it is possible to extract the coded data corresponding to a sub-picture and generate a new bitstream that includes data of only a single sub-picture or a combination / composite of the sub-picture with one or more other sub-pictures, i.e., sub-picture-based bitstream generation can be used to improve flexibility and scalability.
[0087] 3.3 Bitstream FIG. 3 illustrates a bitstream configuration (i.e., structure, organization, or arrangement) according to one embodiment of the present invention that conforms to the requirements of a VVC7 coding system. The bitstream 300 is composed of data representing / indicating an ordered sequence of syntax elements and coded (image) data. The syntax elements and coded (image) data are arranged (i.e., packaged / grouped) into NAL units 301-308. There are different NAL unit types. The Network Abstraction Layer (NAL) provides the functionality / capability to encapsulate the bitstream into packets of various protocols such as Real Time Protocol / Internet Protocol (RTP / IP), ISO Base Media File Format, etc. The Network Abstraction Layer also provides a framework for packet loss resilience.
[0088] The NAL units are split into VCL NAL units and non-VCL NAL units, where VCL stands for video coding layer. The VCL NAL units contain the actual coded video data. The non-VCL NAL units contain additional information, which may be parameters required for decoding the coded video data or supplemental data that may improve the usability of the decoded video data. The NAL units 306 in Figure 3 correspond to slices (i.e., contain the actual coded video data of the slice) and constitute the VCL NAL units of the bitstream.
[0089] Different NAL units 301-305 correspond to different parameter sets, and these NAL units are non-VCL NAL units. DPS NAL unit 301 stands for decoding parameter set NAL unit and contains parameters that are constant for a given decoding process. VPS NAL unit 302, VPS stands for video parameter set NAL unit and contains parameters defined for an entire video (e.g., an entire video includes one or more sequences of pictures / images) and is therefore applicable when decoding the encoded video data of the entire bitstream. DPS NAL units may define parameters that are more static (in the sense that they are stable and do not change much during the decoding process) than parameters in VPS NAL units. In other words, parameters of DPS NAL units change less frequently than parameters of VPS NAL units. SPS NAL unit 303, SPS stands for sequence parameter set and contains parameters defined for a video sequence (i.e., a sequence of pictures or images). In particular, SPS NAL units may define the relevant parameters of a video sequence and the sub-picture layout. The parameters associated with each subpicture specify the coding constraints that apply to the subpicture. According to a variant, it includes a flag indicating that temporal prediction between subpictures is restricted, so that data coming from the same subpicture is available for use during the temporal prediction process. Another flag can enable or disable loop filters (i.e. post-filtering) that cross subpicture boundaries.
[0090] The PPS NAL unit 304, PPS stands for Picture Parameter Set, and contains parameters defined for a picture or a group of pictures. The APS NAL unit 305, APS stands for Adaptive Parameter Set, and contains parameters for a loop filter, typically an Adaptive Loop Filter (ALF) or a reshaping model (or luma mapping with chroma scaling model) or a scaling matrix used at slice level. The bitstream may also contain SEI NAL units (not shown in FIG. 3), which represent Supplemental Enhancement Information NAL units. The periodicity of occurrence (or frequency of inclusion) of these parameter sets (or NAL units) in the bitstream is variable. A VPS defined for the entire bitstream may occur only once in the bitstream. In contrast, an APS defined for a slice may occur once for each slice in each picture. In practice, different slices may depend on (e.g., reference) the same APS, and thus there are typically fewer APS NAL units than slices in a bitstream for a picture.
[0091] The AUD NAL unit 307 is an access unit delimiter NAL unit that separates two access units. An access unit is a set of NAL units that may comprise one or more coded pictures with the same decoding timestamp (i.e., a group of NAL units associated with one or more coded pictures with the same timestamp).
[0092] The PH NAL unit 308 is a picture header NAL unit that groups parameters common to a set of slices of a single coded picture. A picture may reference one or more APSs to indicate the AFL parameters, reconstruction models, and scaling matrices used by the slices of the picture.
[0093] Each VCL NAL unit 306 contains the video / image data for a slice. A slice can correspond to an entire picture or a subpicture, a single tile, or multiple tiles, or a fraction of a tile (a partial tile). For example, the slice of Figure 3 contains several tiles 320. A slice consists of a slice header 310 and a raw byte sequence payload (RBSP) 311 that contains coded pixel / component sample data coded as coding blocks 340.
[0094] The syntax of a PPS such as VVC7 includes syntax elements that specify the size of a picture in luma samples, and also includes syntax elements that specify the partitioning of each picture into tiles and slices.
[0095] The PPS contains syntax elements that allow (i.e. can determine) slice locations within a picture / frame. Since a subpicture forms a rectangular region in a picture / frame, it is possible to determine the set of slices, portions of tiles, or tiles that belong to a subpicture from a parameter set NAL unit (i.e., one or more of the DPS, VPS, SPS, PPS, and APS NAL units).
[0096] Encoding and Decoding Processes 3.4 Encoding process FIG. 4 illustrates an encoding method for encoding pictures of a video into a bitstream according to an embodiment of the present invention.
[0097] In a first step 401, the picture is divided into sub-pictures. For each sub-picture, the size of the sub-picture is determined as a function of the spatial access granularity required by the application (e.g., the sub-picture size can be expressed as a function of the size / scale / granularity level of the region / spatial part / area in the picture that the application / usage scenario requires, and the sub-picture size can be small enough to contain a single region / spatial part / area). Typically, in viewport-dependent streaming techniques, the sub-picture size is set to cover a given range of fields of view (e.g., the range of a 60° horizontal field of view). For adaptive streaming of regions of interest, the width and height of each sub-picture is made to depend on the regions of interest present in the input video sequence. Typically, the size of each sub-picture is made to contain one region of interest. The size of the sub-picture is determined in luma sample units or in multiples of the size of the CTB. Furthermore, in step 401, the position of each sub-picture in the coded picture is determined. The positions and sizes of the sub-pictures form sub-picture layout information that is typically signaled in a non-VCL NAL unit, such as a parameter set NAL unit, for example, the sub-picture layout information is encoded in the SPS NAL unit in step 402.
[0098] The SPS syntax for such an SPS NAL unit typically includes the following syntax elements:
[0099] [Table 3]
[0100] The descriptor string gives the encoding scheme used to encode the syntax element, e.g., u(n), where n is an integer value, means that the syntax element is encoded using n bits, ue(v) means that the syntax element is encoded using an unsigned integer zeroth order Exp-Golomb encoded syntax element that is a variable length encoding left bit first, se(v) is equivalent to ue(v) but for a signed integer, and u(v) means that the syntax element is encoded using a fixed length encoding with a particular length in bits determined from other parameters.
[0101] The presence of subpictures signaled in the bitstream for a picture depends on the value of the flag subpics_present_flag. If this flag is equal to 0, it indicates that the bitstream does not contain any information related to partitioning the picture into one or more subpictures. In such a case, it is presumed that there is a single subpicture covering the whole picture. When this flag is equal to 1, a set of syntax elements specifies the layout of subpictures within a frame (i.e., a picture): the signaling involves determining / defining / specifying the subpictures of a picture (the number of subpictures in a picture is coded in the sps_num_subpics_minus1 syntax element) using a loop (i.e., a programming structure for repeating a sequence of instructions until a certain condition is met), which involves defining the position and size of each subpicture. The index of this "for loop" is the subpicture index. The syntax elements subpic_ctu_top_left_x[i] and subpic_ctu_top_left_y[i] correspond to the column index and row index, respectively, of the first CTU of the ith subpicture. The subpic_width_minus1[i] and subpic_height_minus[i] syntax elements signal the width and height of the ith subpicture in CTUs.
[0102] In addition to the subpicture layout, the SPS specifies constraints on subpicture boundaries: for example, subpic_treated_as_pic_flag[i] equal to 1 indicates that the boundaries of the ith subpicture are treated as picture boundaries for temporal prediction. This ensures that the coded blocks of the ith subpicture are predicted from data of reference pictures belonging to the same subpicture. When equal to 0, this flag indicates that the temporal prediction may be constrained or unconstrained. A second flag (loop_filter_across_subpic_enabled_flag[i]) specifies whether the loop filtering process can use data (usually pixel values) from another subpicture. These two flags make it possible to indicate whether a subpicture is coded independently of other subpictures or not. This information is useful in deciding whether a subpicture can be extracted, derived from, or merged with other subpictures.
[0103] In step 403, the encoder determines partitions in tiles and slices in pictures of the video sequence and describes these partitionings in one non-VCL NAL unit, such as a PPS. This step is further described below with reference to Figure 6. The signaling of slice and tile partitioning is constrained by sub-pictures, such that each sub-picture contains at least a portion of one slice and tile (i.e., a partial tile) or one or more tiles.
[0104] In step 404, at least one slice forming a sub-picture is encoded into a bitstream.
[0105] 3.5 Decryption process Figure 5 shows a general decoding process for a slice according to an embodiment of the present invention. For each VCL NAL unit, the decoder determines the PPS and SPS that apply to the current slice. Typically, it determines the identifiers of the PPS and SPS used for the current picture. For example, the picture header of the slice signals the identifier of the PPS in use. The PPS associated with this PPS identifier also references the SPS using another identifier (the SPS identifier).
[0106] In step 501, the decoder determines the sub-picture partition, determining the size of the sub-pictures of the picture / frame, typically their width and height, for example by parsing a parameter set that describes / indicates the sub-picture layout. In an embodiment compliant with VVC7 and this part of VVC7, the parameter set that contains the information for determining this sub-picture partition is the SPS. In a second step 502, the decoder parses the syntax elements of one parameter set NAL unit (or non-VCL NAL unit) that relate to the partitioning of the picture into tiles. For example, for a VVC7 compliant stream, the tile partitioning signaling is in the PPS NAL unit. During this determination step, the decoder initializes a set of variables that describe / define the characteristics of the tiles present in each sub-picture. For example, the following information can be determined for the i-th sub-picture (see step 601 in FIG. 6): A flag indicating whether the subpicture contains a fraction of a tile, i.e. a partial tile (see step 603 in FIG. 6 ). An integer value indicating the number of tiles in the subpicture (step 602 in FIG. 6) An integer value that specifies the width of the subpicture in tiles (step 604 in Figure 6) An integer value that specifies the height of the subpicture in tiles (step 604 in Figure 6) A list of tile indices present in the subpicture in raster scan order (step 605 in Figure 6) This FIG. 6 illustrates slice partitioning signaling according to one embodiment of the present invention, which includes decision steps that can be used in both the encoding and decoding processes.
[0107] In step 503, the decoder relies on slice partition signaling (within one non-VCL NAL unit, e.g., typically in the PPS for VVC7) and previously determined information to infer (i.e., derive or determine) the slice partitioning for each subpicture. Specifically, the decoder can infer (i.e., derive or determine) the number of slices, the height and width of one or more of the slices. The decoder can also obtain information present in the slice header to determine the decoding position of the CTB present in the slice data.
[0108] In a final step 504 , the decoder decodes the slices of the sub-pictures that form the picture at the positions determined in step 503 .
[0109] Partitioning Signaling 3.6 Signaling Slice Partitioning According to one embodiment of the present invention, a slice can consist of an integer number of complete and contiguous CTU rows or columns within a tile of a picture, or an integer number of complete tiles (the latter possibility means that a slice can contain partial tiles as long as they contain contiguous CTU rows or columns).
[0110] Two modes of slices may also be supported / provided for use: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of complete tiles in tile raster scan order of the picture. In rectangular slice mode, a slice contains either multiple complete tiles that collectively form a rectangular area of the picture, or multiple contiguous complete CTU rows (or columns) of one tile that collectively form a rectangular area of the picture. The tiles in a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice.
[0111] VVC7's syntax for specifying slice structure (layout and / or partitioning) is unrelated to VVC7's syntax for subpictures. For example, slice partitioning (i.e., partitioning a picture into slices) is done on top of (i.e., based on or with reference to) the tile structure / partitioning, without reference to a subpicture (i.e., without reference to syntax elements used to form a subpicture). On the other hand, VVC7 imposes some constraints (i.e., restrictions) on subpictures, e.g., a subpicture must contain one or more slices, and the slice header contains a slice_address syntax element that is an index of the slice with respect to that subpicture (i.e., an index defined for the associated subpicture, such as an index of the slice among slices in the subpicture). VVC7 also only allows slices containing partial tiles in rectangular slice modes, and raster scan slice modes do not prescribe such slices containing partial tiles. The current syntax system used in VVC7 does not enforce all of these constraints by design, and therefore implementation of this syntax system leads to a system that is prone to producing bitstreams that do not comply with the VVC7 specification / requirements.
[0112] Therefore, embodiments of the present invention use information from the sub-picture layout definition to specify slice partitions and seek to provide better coding efficiency for sub-picture, tile, and slice signaling.
[0113] The requirement to have the ability to define multiple slices within a tile (also called "tile fraction" slices or slices containing partial tiles) comes from the omni-directional streaming requirements. It was identified that slices need to be defined within tiles for the BEAMER (Bitstream Extraction And MERging) operation of OMAF streams. This means that "tile fraction" slices can be in different subpictures to enable the BEAMER operation, which means that it makes little sense to have a subpicture containing one complete tile with multiple slices.
[0114] Following the first three embodiments of the present invention (embodiment 1, embodiment 2 and embodiment 3), we define slice partition determination and signaling based on sub-pictures. The first embodiment 1 includes a syntax system that prohibits / prohibits / prohibits a sub-picture from including more than one "tile fraction" slice, while the second embodiment 2 includes a syntax system that allows / allows only if the sub-picture includes at most one tile (i.e., if the sub-picture includes more than one tile, all its slices include an integer number of complete tiles). The embodiment 3 provides explicit signaling of the use of "tile fraction" slices. When "tile fraction" slices are not used, it avoids signaling syntax elements related to such slices, which can improve the coding efficiency of slice partition signaling. The embodiments 1 to 3 allow / allow the use of tile fraction slices only in rectangular slice mode.
[0115] A fourth embodiment 4 is an alternative to the first three embodiments, in which a unified syntax system is used to allow / permit the use of tile fraction slices in both raster scan and rectangular slice modes. A fifth embodiment 5 is an alternative to the other embodiments, in which the sub-picture layout and slice partitioning are not signaled in the bitstream, but are inferred (i.e., determined or derived) from the tile signaling.
[0116] EMBODIMENT 1 In a first embodiment, the syntax for slice partitions in VVC7 is modified to avoid specifying a large number of constraints that are difficult to implement correctly, and to rely on the subpicture layout to infer / derive parameters for slice partitioning, such as the size of a slice. Since a subpicture is represented by (i.e., composed of) a set / group of complete slices, it is possible to infer the size of a slice when the subpicture contains a single slice. Similarly, the size of the last slice can be inferred / derived / determined from the number of slices in the subpicture and the size of previously processed / encountered slices in the subpicture.
[0117] In embodiment 1, a slice is allowed to contain a fraction / portion of a tile (i.e., a partial tile) only if the slice covers the entire / whole subpicture (in other words, there is a single slice in the subpicture). Slice size signaling is required when there are more than one slice in the subpicture. On the other hand, the slice size is the same as the subpicture size when there is one slice in the subpicture. As a result, when a slice contains a fraction of a tile, the slice size is not signaled in the parameter set NAL unit (because it is the same as the subpicture size). Therefore, it is possible to constrain slice width and height to be in tiles only for slice size signaling scenarios.
[0118] The syntax of the PPS for this embodiment includes specifying the number of slices contained in each subpicture. If the number of slices is greater than one, the size (width and height) of the slices is expressed in tiles. The size of the last slice is not signaled or inferred / derived / determined from the subpicture layout as described above.
[0119] According to a variation of this embodiment, the PPS syntax includes the following syntax elements with the following semantics (ie, definitions or functions):
[0120] PPS Syntax
[0121] [Table 4]
[0122] PPS Semantics Slices are defined per subpicture. A "for loop" is used with the num_slices_in_subpic_minus1 syntax element to form / process the correct number of slices in that particular subpicture. The syntax element num_slices_in_subpic_minus1[i] indicates the number of slices (in the subpicture with subpicture index equal to i) minus 1, i.e. the syntax element indicates one less than the number of slices in the subpicture. When equal to 0, it indicates that the subpicture contains a single slice of size equal to the subpicture size. If the number of slices is greater than 1, the size of the slice is expressed in units of an integer number of tiles. The size of the last slice is inferred from the subpicture size (and the size of the other tiles in the subpicture). With this approach, it is possible to define "tile fraction" slices if they cover the complete subpicture with the following semantics for the syntax element:
[0123] pps_num_subpics_minus1 plus 1 specifies the number of subpictures in the coded picture that refer to the PPS. It is a bitstream conformance requirement that the value of pps_num_subpic_minus1 be equal to sps_num_subpics_minus1 (the number of subpictures defined at the SPS level).
[0124] single_slice_per_subpic_flag equal to 1 specifies that each subpicture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each subpicture consists of one or more rectangular slices. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag is equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1 (the number of subpictures defined at the SPS level).
[0125] num_slices_in_subpic_minus1[i] plus 1 specifies the number of rectangular slices in the ith subpicture. The value of num_slices_in_subpic_minus1 must be in the range from 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Annex A. If no_pic_partition_flag is equal to 1, the value of num_slices_in_subpic_minus1[0] is inferred to be equal to 0.
[0126] This syntax element determines the length of slice_address in the slice header, which is Ceil(log2(num_slices_in_subpic_minus1[SubPicIdx]+1)) bits, where SubPicIdx is the index of the subpicture of the slice. The value of slice_address is in the range from 0 to num_slices_in_subpic_minus1[SubPicIdx], inclusive.
[0127] A tile_idx_delta_present_flag equal to 0 specifies that no tile_idx_delta values are present in the PPS and all rectangular slices in all subpictures of the picture referencing the PPS are specified in raster order. A tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values may be present in the PPS and all rectangular slices in all subpictures of the picture referencing the PPS are specified in the order indicated by the values of tile_idx_delta.
[0128] slice_width_in_tiles_minus1[i][j] plus 1 specifies the width of the jth rectangular slice in tile columns in the ith subpicture. The value of slice_width_in_tiles_minus1[i][j] must be in the range from 0 to NumTileColumns-1, inclusive, where NumTileColumns is the number of tile columns in the tile grid. If not present, the value of slice_width_in_tiles_minus1[i][j] is inferred as a function of the subpicture size.
[0129] slice_height_in_tiles_minus1[i][j] plus 1 specifies the height of the jth rectangular slice in units of tile rows in the ith subpicture. The value of slice_height_in_tiles_minus1[i][j] ranges from 0 to NumTileRows-1, inclusive, where NumTileRows is the number of tile rows in the tile grid. If not present, the value of slice_height_in_tiles_minus1[i][j] is inferred as a function of the ith subpicture size.
[0130] tile_idx_delta[i][j] specifies the tile index difference between the jth and the (j+1)th rectangular slices of the ith subpicture. The value of tile_idx_delta[i][j] must be in the range -NumTilesInPic[i]+1 to NumTilesInPic[i]-1, inclusive, where NumTilesInPic[i] is the number of tiles in the picture. If not present, the value of tile_idx_delta[i][j] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i][j] is not equal to 0.
[0131] Therefore, according to this variant, A subpicture is a rectangular region of one or more slices within a picture. The addresses of the slices are defined relative to the subpicture. This association / relationship between a subpicture and its slices is reflected in a syntax system that defines the slices in a "for loop" that applies to each subpicture. · By design, it avoids undesirable partitioning that can lead to having two or more tile fraction slices from two different tiles in the same subpicture. It is possible to infer slice partitioning from both tile and sub-picture partitioning, which improves the coding efficiency of the signaling. According to a further variant of this variant, this inference / derivation of slice partitioning is performed using the following process: For rectangular slices, a list NumCtuInSlice[i] for i ranging from 0 to num_slices_in_pic_minus1, inclusive, specifies the number of CTUs in the ith slice, and a matrix CtbAddrInSlice[i][j] for i ranging from 0 to num_slices_in_pic_minus1, inclusive, and j ranging from 0 to NumCtuInSlice[i]-1, inclusive, specifies the picture raster scan address of the jth CTB in the ith slice, and is derived as follows if single_slice_per_subpic_flag is equal to 0:
[0132] [Table 5]
[0133] Here, the function AddCtbsToSlice(sliceIdx, startX, stopX, startY, stopY) fills the CtbAddrInSlice array for the slice with indices equal to SliceIdx. It fills the array with CTB addresses in CTB raster scan order, with the vertical addresses of CTB rows being between startY and stopY, and the horizontal addresses of CTB columns being between startX and stopX.
[0134] The process includes applying a processing loop to each subpicture. For each subpicture, the tile index of the first tile in the subpicture is determined from the horizontal and vertical addresses of the first tile of the subpicture (i.e., the subpicTileTopLeftX[i] and subpicTileTopLeftX[i] variables) and the number of tile columns specified by the tile partition information. This value infers / indicates / represents the index of the first tile in the first slice of the subpicture. For each subpicture, a second processing loop is applied to each slice of the subpicture. The number of slices is equal to 1 plus the num_slices_in_subpics_minus1[i][j] variable and is either coded in the PPS or inferred / derived / determined from other information included in the bitstream. If the slice is the last one in the subpicture, the width of the slice in the tile is inferred / inferred / derived / determined to be equal to the subpicture width in the tile minus the horizontal address of the column of the first tile in the slice plus the horizontal address of the column of the first tile in the subpicture. Similarly, the height of the slice in the tile is inferred / inferred / derived / determined to be equal to the subpicture height in the tile minus the vertical address of the row of the first tile in the slice plus the vertical address of the row of the first tile in the subpicture. The index of the first tile in the previous slice is either coded in the slice partitioning information (e.g., as a difference from the tile index of the first tile in the previous slice) or inferred / inferred / derived / determined to be equal to the next tile in the raster scan order of tiles in the subpicture.
[0135] If the subpicture contains a fraction of a tile (i.e., a partial tile), the width and height in CTUs of the slice are presumed / estimated / derived / determined to be equal to the width and height of the subpicture in CTUs. The CtbAddrInSlice[sliceIdx] array is filled with the CTUs of the subpicture in raster scan order. Otherwise, the subpicture contains one or more tiles and the CtbAddrInSlice[sliceIdx] array is filled with the CTUs of the tiles contained in the slice. A tile contained in a subpicture slice is a tile having a vertical address of a tile column, a vertical address defined in the range, [tileX, tileX + slice_width_in_tiles_minus1[i][j]], and a horizontal address of a tile column, a horizontal address defined in the range, [tileY, tileY + slice_height_in_tiles_minus1[i][j]], where tileX is the vertical address of the tile column of the first tile in the slice, tileY is the horizontal address of the tile row of the first tile in the slice, the j index of the slice in the subpicture, and the i subpicture index.
[0136] Finally, the processing loop for a slice includes a determination step of determining the first tile in the next slice of the subpicture. If the tile index offset (tile_idx_delat[i][j]) is coded (tile_idx_delta_present_flag[i] is equal to 1), the tile index of the next slice is set equal to the index of the first tile of the current slice plus the value of the tile index offset value. Otherwise (i.e., the tile index offset is not coded), tileIdx is set equal to the tile index of the first tile of the subpicture plus the product of the number of tile columns in the picture and the height in tile units minus 1 for the current slice.
[0137] According to a further variant, if a sub-picture contains a partial tile, instead of signaling the number of slices contained in the sub-picture in the PPS, it is inferred / inferred / derived / determined that the sub-picture consists of this partial tile. For example, to do this, the following alternative PPS syntax could be used instead:
[0138] Alternative PPS Syntax In a further variant, if a subpicture represents a fraction of a tile (i.e., the subpicture contains a partial tile), the number of slices in the subpicture is inferred to be equal to 1. The coded data size of the syntax elements is further reduced as the number of subpictures representing tile fractions (i.e., the number of subpictures containing partial tiles) increases, improving the compression of the stream. The syntax of the PPS is, for example, as follows:
[0139] [Table 6]
[0140] The new semantic / definition of num_slices_in_subpic_minus1[i] is:
[0141] num_slices_in_subpic_minus1[i] plus 1 specifies the number of rectangular slices in the i-th subpicture. The value of num_slices_in_subpic_minus1 ranges from 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Annex A. If not present, the value of num_slices_in_subpic_minus1[i] is inferred to be equal to 0 for i in the range from 0 to pps_num_subpic_minus1, inclusive.
[0142] The tileFractionSubpicture[i] variable specifies whether the i-th subpicture covers a fraction of a tile (a partial tile), i.e., whether the size of the subpicture is strictly smaller (i.e., smaller) than the tile to which the subpicture's first CTU belongs.
[0143] The decoder determines this variable from the subpicture layout and tile grid information as follows: if both the top and bottom horizontal boundaries of the subpicture are tile boundaries, then tileFractionSubpicture[i] is set equal to 0. In contrast, if at least one of the top or bottom horizontal subpicture boundaries is not a tile boundary, then tileFractionSubpicture[i] is set equal to 1.
[0144] In yet another variation, the presence or absence of slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] is inferred from the width and height of the subpicture in tiles. For example, the variables subPictureWidthInTiles[i] and subPictureHeightInTiles[i] define the width and height of the ith subpicture in tiles, respectively. If the subpicture is a fraction of a tile, the width of the subpicture is set to 1 (since a fraction of a tile has a width equal to the width of the tile) and the height is set by convention to 0 to indicate that the height of the subpicture is less than one full tile. It should be understood that any other preset / predetermined values can be used, the main constraint being that the two values are set / determined such that they do not represent possible subpicture sizes within a tile. For example, the values may be set equal to the maximum number of tiles in a picture plus 1. In that case, it can be inferred that the subpicture is a fraction of a tile, since no subpicture width or height larger than the maximum number of tiles is possible.
[0145] These variables are initialized once the tile partitioning is determined (typically based on the num_exp_tile_columns_minus1, num_exp_tile_rows_minus1, tile_column_width_minus1[i], and tile_row_height_minus1[i] syntax elements). The processing may include performing a processing loop for each subpicture. That is, for each subpicture, if the subpicture covers a fraction of a tile, the subpicture's width and height in tile units are set to 1 and 0, respectively. Otherwise, the subpicture's width in tiles is determined as follows: the horizontal address of the CTU column of the first CTU of the subpicture (determined from the subpicture layout syntax element subpic_ctu_top_left_x[i] for the i-th subpicture) is used to determine the horizontal address of the tile column of the first tile in the subpicture.
[0146] Then, for each tile column of the tile grid, the rightmost and leftmost horizontal addresses of the tile column are determined. If the horizontal address of the CTU column of the first CTU of the subpicture is between these two addresses, the horizontal address indicates the horizontal address of the tile column of the first tile in the subpicture. The same process is applied to determine the horizontal address of the tile column containing the rightmost CTU column of the subpicture. The rightmost CTU column has a horizontal address of the CTU column equal to the sum of the horizontal address of the CTU column of the first CTU of the subpicture and the width of the subpicture in CTU units (subpic_width_minus1[i]+1). The width of the subpicture in the tile is equal to the difference between the horizontal address of the tile column of the last CTU column and the horizontal address of the first CTU of the subpicture. The same principle is applied when determining the height of the subpicture in tiles. The process determines the subpicture height in the tile as the difference between the vertical addresses of the tile row of the first CTU of the subpicture and the last CTU row of the subpicture.
[0147] In one further variant, if the number of tiles in a subpicture is equal to the number of slices in the subpicture, slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] are inferred to be absent and equal to 0. Indeed, in the equivalent case, the subpicture and slice constraints impose that each slice contains exactly one tile. The number of tiles in a subpicture is equal to 1 if the subpicture is a tile fraction (tileFractionSubpicture[i] is equal to 1); otherwise, it is equal to the product of subPictureHeightInTiles[i] and subPictureWidthInTiles[i].
[0148] In another further variant, if the subpicture width in tiles is equal to 1, slice_width_in_tiles_minus1[i][j] is not present and is inferred to be equal to 0.
[0149] In another further variation, if the subpicture height in tiles is equal to 1, slice_height_in_tiles_minus1[i][j] is not present and is inferred to be equal to 0.
[0150] In yet another variation, any combination of the three preceding further variations is used.
[0151] In some of the previous variants, the presence of the syntax element is inferred from the sub-picture partition information. If the sub-picture partitioning is defined in a different parameter set NAL unit, the parsing of the slice partitioning depends on information from the other parameter set NAL unit. This dependency may limit the use of the variants in certain applications, since the parsing of a parameter set that includes the slice partitioning cannot be performed without storing information from another parameter set. With respect to the decoding of the parameters, i.e., the determination of the value encoded by the syntax element, this dependency is not a limitation, since the decoder needs all parameter sets used to decode the pixel sample in any way (but there may be some latency, since it has to wait for all relevant parameter sets to be decoded). As a result, in a further variant, the inference of the presence of the syntax element is enabled only when the sub-picture, tile, and slice partitioning are signaled in the same parameter set NAL unit. For example, the variables subPictureWidthInTiles[i], which specifies the width of the ith subpicture in tiles, subPictureHeightInTiles[i], which specifies the height of the ith subpicture in tiles, subpicTileTopLeftX[i], which specifies the horizontal address of the column of the first tile of the ith subpicture, and subpicTileTopLeftY[i], which specifies the vertical address of the row of the first tile of the ith subpicture, for i in the range 0 to pps_num_subpicture_minus1, inclusive, are determined as follows:
[0152] [Table 7]
[0153] The tileFractionSubpicture[i] variable, which specifies whether a subpicture contains a fraction of a tile, is derived as follows:
[0154] [Table 8]
[0155] A list SliceSubpicToPicIdx[i][k] specifying the number of rectangular slices in the i-th subpicture and the picture-level slice index of the k-th slice in the i-th subpicture is derived as follows.
[0156] [Table 9]
[0157] Where: CtbToTileRowBd[ctbAddrY] converts the vertical CTB address (ctbAddrY) to the first tile row boundary in CTB units. CtbToTileColBd[ctbAddrX] converts a horizontal CTB address (ctbAddrX) to the left tile column boundary in CTB units. ColWidth[i] is the width of the i-th tile column in the CTB. RowHeight[i] is the height of the i-th tile row in the CTB. tileColBd[i] is the position of the i-th tile column boundary in the CTB. tileRowBd[i] is the position of the i-th tile row boundary in the CTB. NumTileColumns is the number of tile columns NumTileRows is the number of tile rows 7 shows an example of sub-picture and slice partitioning using signaling of the above-mentioned embodiment / variant / further variant. In this example, a picture 700 is partitioned into nine sub-pictures, labelled (1)-(9), and a 4x5 tile grid (tile boundaries shown in thick solid lines). The slice partitioning (areas included in each slice are shown in thin solid lines just inside the slice boundaries) is as follows for each sub-picture: Subpicture (1): 3 slices containing 1 tile, 2 tiles and 3 tiles respectively. The height of the slices is equal to 1 tile and their widths in tiles are 1, 2 and 3 respectively (i.e. the 3 slices consist of horizontally arranged rows of tiles). Subpicture (2): Two slices of equal size, one tile wide by one tile high (i.e. each of the two slices consists of a single tile) Subpictures (3)-(6): A "tile fraction" slice, i.e. a slice consisting of a single partial tile Subpicture (7): Two slices with a size of two tile rows (i.e. each of the two slices consists of two vertically arranged rows of tiles). Subpicture(8): 1 slice of a row of 3 tiles Subpicture (9): Two slices with sizes of 1 tile row and 2 tile rows. For sub-picture (1), the width and height of the two first slices are coded and the size of the last slice is estimated.
[0158] For subpicture(2), the width and height of the first two slices are inferred because there are two slices for two tiles in the subpicture.
[0159] For subpictures (3)-(6), the number of slices in each subpicture is equal to 1, and since each subpicture is a fraction of a tile, the width and height of the slice are inferred to be equal to the subpicture size.
[0160] For sub-pictures (7), the width and height of the first slice and the width and height of the last slice are inferred from the sub-picture size.
[0161] For subpictures (8), the width and height of the slice are inferred to be equal to the subpicture size, since there is a single slice within the subpicture.
[0162] For subpictures (9), the slice height is estimated to be equal to 1 (because the height of a subpicture within a tile is equal to 1) and the width of the first slice is coded, while the width of the last slice is equal to the subpicture width minus the width of the first slice.
[0163] EMBODIMENT 2 In a second embodiment, embodiment 2, the constraint / restriction that "tile fraction" slices (i.e. slices that are partial tiles) cover the entire subpicture is relaxed / removed. As a result, a subpicture can contain one or more slices, each of which contains one or more tiles, but can also contain one or more "tile fraction" slices.
[0164] In this embodiment, sub-picture partitioning allows / enables prediction / derivation / determination of slice positions and sizes.
[0165] According to a variation of this embodiment, this can be done using the following PPS syntax:
[0166] PPS Syntax For example, the PPS syntax looks like this:
[0167] [Table 10]
[0168] The semantics of the syntax elements single_slice_per_subpic_flag, tile_idx_delta_present_flag, num_slices_in_subpic_minus1[i] and tile_idx_delta[i][j] are the same as in the previous embodiment.
[0169] The slice_width_minus1 and slice_height_minus1 syntax elements (parameters) specify the slice size in tiles or CTUs, depending on the subpicture partitioning.
[0170] The variable newTileIdxDeltaRequired is set equal to 1 if the last CTU of the slice is the last CTU of a tile. If the slice is not a "tile fraction" slice, newTileIdxDeltaRequired is set equal to 1. If the slice is a "tile fraction" slice, newTileIdxDeltaRequired is set to 0 if the slice is not the last of the tiles in the slice, otherwise it is the last of the tile and newTileIdxDeltaRequired is set to 1.
[0171] In a first further variant, a sub-picture is constrained / restricted to contain fractional tile slices of a single tile. In this case, if the sub-picture contains more than one tile, the size is in tiles. Otherwise, the sub-picture contains a single tile or a part of a tile (fractional tile_ and the slice height is defined in CTU units. The width of the slice is necessarily equal to the sub-picture width and therefore can be inferred and does not need to be coded in PPS.
[0172] slice_width_minus1[i][j] plus 1 specifies the width of the jth rectangular slice. The value of slice_width_in_tiles_minus1[i][j] must be in the range from 0 to NumTileColumns-1, inclusive, where NumTileColumns is the number of tile columns in the tile grid. If not present (i.e., if subPictureWidthInTiles[i]*subPictureHeightInTiles[i] == 1 or subPictureWidthInTiles[i] is equal to 1), the value of slice_width_in_tiles_minus1[i][j] is inferred to be equal to 0.
[0173] slice_height_in_tiles_minus1[i][j] plus 1 specifies the height of the jth rectangular slice in the ith subpicture. The value of slice_height_in_tiles_minus1[i][j] ranges from 0 to NumTileRows-1, inclusive, where NumTileRows is the number of tile rows in the tile grid. If not present (i.e., if subPictureHeightInTiles[i] is equal to 1 and num_slices_in_subpic_minus1[i] == 0), the value of slice_height_in_tiles_minus1[i][j] is inferred to be equal to 0.
[0174] For i in the range of 0 to pps_num_subpic_minus1 and j in the range of 0 to num_slices_in_subpic_minus1[i], the variables SliceWidthInTiles[i][j], which specifies the width in tiles of the jth rectangular slice of the ith subpicture, SliceHeightInTiles[i][j], which specifies the height in tiles of the jth rectangular slice of the ith subpicture, and SliceHeightInCTU[i][j], which specifies the height in CTB units of the jth rectangular slice of the ith subpicture, are derived as follows:
[0175] [Table 11]
[0176] In this algorithm, SliceHeightInCTU[i][j] is valid only if sliceHeightInTiles[i][j] is equal to 0.
[0177] In a further alternative variant, a subpicture may contain tile fraction slices from several tiles (i.e. slices containing partial tiles from more than one tile), with the restriction / constraint / condition that all subpicture slices must be tile fraction slices. Thus, it is not possible to define a subpicture containing a first tile fraction slice containing a partial tile of a first tile and a second tile fraction slice containing a partial tile of a second tile different from the first tile, and another slice covering a third (different) tile entirely (i.e. the third tile is a complete / whole tile). In this case, if the subpicture contains more slices than the subpicture tiles, it is inferred that the slice_height_minus1[i][j] syntax element is in CTU units and slice_width_minus1[i][j] is equal to the subpicture width in CTU units. Otherwise, the size of the slice is in tiles.
[0178] In a further variation, the slice PPS syntax is as follows:
[0179] [Table 12]
[0180] In this further variant, the same principles apply, except that separate syntax elements define the slice width and height in tiles or slice height in CTUs, depending on whether the slice is defined in CTUs or not. The slice width and height are defined by slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] if the slice is signaled in tiles, and slice_width_in_ctu_minus1[i][j] is used if the width is in CTUs. The variable sliceInCtuFlag[i] equal to 1 indicates that the i-th subpicture contains only tile fraction slices (i.e. there are no whole / complete tile slices in the subpicture). Equal to 0 indicates that the i-th subpicture contains a slice with one or more tiles.
[0181] The variable sliceInCtuFlag[i], for i in the range 0 to pps_num_subpic_minus1, is derived as follows:
[0182] [Table 13]
[0183] The determination of the sliceInCtuFlag[i] variable introduces a parsing dependency between slice and sub-picture partitioning information, so that in a variant, sliceInCtuFlag[i] is signaled and not inferred when slice, tile, and sub-picture partitioning are signaled in different parameter set NAL units.
[0184] In a further variant, a subpicture may include tile fraction slices from several tiles, without any particular constraint / limitation. Thus, it is possible to define a subpicture that includes a first tile fraction slice that includes a partial tile of a first tile, a second tile fraction slice that includes a partial tile of a second tile different from the first tile, and another slice that covers a third (different) tile in its entirety (i.e., the third tile is a complete / whole tile). In this case, if the number of tiles in the subpicture is greater than 1, a flag indicates whether the slice size is specified in CTU units or tiles. For example, the following syntax of the PPS signals a slice_in_ctu_flag[i] syntax element that indicates whether the slice size of the i-th subpicture is expressed in CTU units or tiles. slice_in_ctu_flag[i] equal to 0 indicates that the slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] syntax elements are present and slice_height_in_ctu_minus1[i][j] is not present, i.e. the slice size is expressed in tiles. slice_in_ctu_flag[i] equal to 1 indicates that slice_height_in_ctu_minus1[i][j] is present and slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] syntax elements are not present, i.e. the slice size is expressed in CTUs.
[0185] [Table 14]
[0186] 8 shows an example of sub-picture and slice partitioning using signaling of the above-mentioned embodiments / variations / further. In this example, a picture 800 is partitioned into six sub-pictures, labeled (1)-(6), and a 4x5 tile grid (tile boundaries shown by thick solid lines). The slice partitioning (areas included in each slice are shown by thin solid lines just inside the slice boundaries) is as follows for each sub-picture: Subpicture (1): Three slices with sizes of 1 tile, 2 tiles, and 3 tile rows (i.e., three slices consisting of horizontally arranged rows of tiles). Subpicture (2): Two slices of equal size, each one tile in size (i.e. each of the two slices consists of a single tile) Subpicture (3): 4 "tile fraction" slices, i.e., 4 slices each consisting of a single partial tile Subpicture (4): Two slices with a size of two tile rows (i.e., each of the two slices consists of two vertically arranged rows of tiles). Subpicture(5): 1 slice of a row of 3 tiles Subpicture (6): Two slices with sizes of 1 tile row and 2 tile rows. For subpicture (3), the number of slices in the subpicture is equal to 4 and the subpicture contains only two tiles. The slice width is inferred to be equal to the subpicture width and the slice height is specified in CTU units.
[0187] For subpictures (1), (2), (4), and (5), the number of slices is less than the tiles in the subpicture, and therefore width and height are specified in tiles only when necessary, i.e., when they cannot be inferred / derived / determined from other information.
[0188] For sub-picture (1), the width and height of the two first slices are coded and the size of the last slice is estimated.
[0189] For subpicture(2), the width and height of the first two slices are inferred because there are two slices for two tiles in the subpicture.
[0190] For sub-pictures (4), the width and height of the first slice and the size of the last slice are inferred from the sub-picture size.
[0191] For subpictures (5), the width and height are inferred to be equal to the subpicture size since there is a single slice within the subpicture.
[0192] For subpictures (5), the slice height is estimated to be equal to 1 (since the subpicture height in tiles is equal to 1) and the width of the first slice is encoded, while the width of the last slice is estimated to be equal to the subpicture width minus the size of the first slice.
[0193] EMBODIMENT 3 In this third embodiment, embodiment 3, we specify in the bitstream whether tile fraction slices are enabled or disabled. The principle is to include in one of the parameter set NAL units (or non-VCL NAL units) a syntax element that indicates whether the use of "tile fraction" slices is allowed or not.
[0194] According to a variant, the SPS includes a flag indicating whether "tile fraction" slices are allowed or not, and thus the "tile fraction" signaling can be skipped if this flag indicates that "tile fraction" slices are not allowed. For example, if the flag is equal to 0, it indicates that "tile fraction" slices are not allowed. If the flag is equal to 1, "tile fraction" slices are allowed. The NAL unit can include a syntax element indicating the location of the "tile fraction" slices. This can be done, for example, using the following SPS syntax element and its semantics:
[0195] SPS syntax: enable / disable tile fraction slicing
[0196] [Table 15]
[0197] SPS Semantics sps_tile_fraction_slices_enabled_flag specifies whether "tile fraction" slices are enabled in the coded video sequence. sps_tile_fraction_slices_enabled_flag equal to 0 indicates that the slice contains an integer number of tiles. sps_tile_fraction_slices_enabled_flag equal to 1 indicates that the slice can contain an integer number of tiles or an integer number of CTU rows from one tile.
[0198] In a further variant, sps_tile_fraction_slices_enabled_flag is specified at the PPS level, providing more granularity to adaptively apply / define the presence of "tile fraction" slices. In yet another variant, a flag can be placed in the picture header NAL unit to enable / enable the adaptation of the presence of "tile fraction" slices on a picture basis. The flag can be present in multiple NAL units with different values to allow overriding the configuration defined in a higher level parameter set. For example, the value of the flag in the picture header overrides the value in the SPS, which overrides the value in the PPS.
[0199] In alternative variations, the value of sps_tile_fraction_slices_enabled_flag may be constrained or inferred from other syntax elements, for example, sps_tile_fraction_slices_enabled_flag is inferred to be equal to 0 if subpictures are not used in the video sequence (i.e., subpics_present_flag is equal to 0).
[0200] The variants of the first and second embodiments can take into account the value of sps_tile_fraction_slices_enabled_flag to infer the presence or absence of tile fraction slices signaling in a similar manner. For example, the above PPS can be modified as follows:
[0201] [Table 16]
[0202] Slice height signaling is inferred on a per-tile basis when sps_tile_fraction_slices_enabled_flag is equal to 0.
[0203] EMBODIMENT 4 In VVC7, tile fraction slices are only valid in rectangular slice mode. The fourth embodiment described below has the advantage that it allows the use of tile fraction slices also in raster scan slice mode. This offers the possibility to adjust the bit length of the coded slices more precisely, since slice boundaries are not constrained to align with tile boundaries as in VVC7.
[0204] The principle involves defining slice partitioning (or providing related information) in two places: The parameter set defines slices in terms of tiles. Tile fraction slices are signaled in the slice header. In a variant, sps_tile_fraction_slices_enabled_flag is pre-determined to be equal to 1, and there is always a "tile fraction" signaling in the slice header.
[0205] To achieve this, in effect, the semantics of a slice is modified from that of VVC7 and the previous embodiments / variants / further variants: A slice is a set of one or more slice segments that collectively represent a tile of a picture or an integer number of complete tiles. A slice segment represents an integer number of complete tiles or an integer number of contiguous complete CTU rows (i.e., a "tile fraction") within a tile of a picture that are exclusively contained in a single NAL unit, i.e., one or more tiles or "tile fractions". A "tile fraction" slice is a set of contiguous CTU rows of a tile. It is possible for a slice segment to include all CTU rows of a tile. In such a case, a slice segment includes a single slice segment.
[0206] According to a variant, the PPS syntax of any previous embodiment is modified to include tile fraction specific signaling. An example of such a PPS syntax change is given below:
[0207] PPS Syntax The PPS syntax of any of the previous embodiments is modified to remove the tile fraction specific signaling. For example, the PPS syntax becomes:
[0208] [Table 17]
[0209] Uses the same semantics as the syntax element.
[0210] Slice Segment Syntax
[0211] [Table 18]
[0212] A slice segment NAL unit consists of a slice segment header and slice segment data, which is similar to the VVC7 NAL unit structure of a slice. The slice header from the previous embodiment becomes a slice segment header with the same syntax elements as the slice header, but as a slice segment header, includes additional syntax elements to locate / identify the slice segment within a slice (e.g., as described / defined in the PPS).
[0213] [Table 19]
[0214] The slice segment header contains signaling to specify in which CTU row the slice segment starts within the slice. slice_ctu_row_offset specifies the CTU line offset of the first CTU in the slice.
[0215] If rect_slice_flag is equal to 0 (i.e., the slice mode is raster scan slice mode), the CTU line offset is relative to the first row of the tile with index equal to slice_address. If rect_slice_flag is equal to 1 (i.e., rectangular slice mode), the CTU line offset is relative to the first CTU of the slice with index equal to slice_address of the subpicture identified by slice_subpic_id. The CTU line offset is coded using variable or fixed length coding. In the fixed length case, the number of CTU rows in the slice is determined from the PPS, and the bit length of the syntax element is equal to the log2 of the number of CTU rows minus 1.
[0216] There are two ways to indicate the end of a slice segment.
[0217] In the first method, the slice segment indicates the number of CTU rows in the slice segment (-1). The number of CTU rows is coded using variable length coding or fixed length coding. In the case of fixed length coding, the number of CTU rows in the slice is determined from the PPS. The bit length of the syntax element is equal to the log2 of the difference of the number of CTU rows in the slice minus the CTU row offset, minus 1.
[0218] An example of the syntax of a slice header is as follows:
[0219] [Table 20]
[0220] num_ctu_rows_in_slice_minus1 plus 1 specifies the number of CTU rows in a slice segment NAL unit. The range of num_ctu_rows_in_slice_minus1 is from 0 to the number of CTU rows in tiles contained in the slice minus 2.
[0221] If sps_tile_fraction_slices_enabled_flag is equal to 1 and num_tiles_in_slice_segment_minus1 is equal to 0, the variable NumCtuInCurrSlice, which specifies the number of CTUs in the current slice, is equal to the number of CTU rows multiplied by the width of the tiles present in the slice in CTUs.
[0222] In the second method, the slice segment data includes signaling at the end of each CTU row to specify whether the slice segment ends. The advantage of this second method is that the encoder does not need to pre-determine the number of CTUs in a given slice segment. This reduces the latency for encoders that can output slice headers in real time, whereas the first method requires buffering the slice header to indicate the number of CTU rows in the slice segment at the end of encoding the slice segment.
[0223] EMBODIMENT 5 The fifth embodiment is a modification to the signaling of the sub-picture layout, which may lead to improvements in certain situations. In fact, increasing the number of sub-pictures or slices or tiles in a video sequence limits / restricts the effectiveness / efficiency of temporal and intra prediction mechanisms. As a result, the compression efficiency of the video sequence may be reduced. For this reason, there is a high probability that the sub-picture layout can be pre-determined / determined / predicted / estimated as a function of the application requirements (e.g., the size of the ROI). In the encoding process, an optimal tile partitioning for the sub-picture layout is generated. In the best case scenario, each sub-picture contains exactly one tile. To limit the impact on compression efficiency, the encoder tries to minimize the number of slices per sub-picture by using a single slice per sub-picture. As a result, the best option for the encoder is to define one slice and one tile per sub-picture.
[0224] In such a case, the sub-picture layout and the slice layout are the same. This embodiment adds a flag to the SPS to indicate such a particular case / scenario / situation. If this flag is equal to 1, then the sub-picture layout does not exist and it can be inferred / derived / determined to be the same as the slice partition. Otherwise, if the flag is equal to 0, the sub-picture layout is explicitly signaled in the bitstream according to what has been described above with respect to the previous embodiment / variant / further variant.
[0225] According to a variant, the SPS includes the sps_single_slice_per_subpicture flag for this purpose.
[0226] [Table 21]
[0227] sps_single_slice_per_subpicture equal to 1 indicates that each subpicture contains a single slice and there are no subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i] and subpic_height_minus1[i] for i in the range from 0 to sps_num_subpics_minus1, inclusive. sps_single_slice_per_subpicture equal to 0 indicates that the subpicture may or may not contain a single slice and there are subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i] and subpic_height_minus1[i] for i in the range from 0 to sps_num_subpics_minus1, inclusive.
[0228] According to yet another variant, the PPS syntax includes the following syntax element to indicate that the sub-picture layout can be inferred from the slice layout:
[0229] [Table 22]
[0230] If either pps_single_slice_per_subpic_flag or sps_single_slice_per_subpic_flag is equal to 1, there is a single slice per subpicture. If sps_single_slice_per_subpic_flag is equal to 1, there is no slice layout in the SPS and pps_single_slice_per_subpic_flag must be equal to 0. The PPS then specifies the slice partitions: the i-th subpicture has a size and position corresponding to the i-th slice (i.e., the i-th subpicture and the i-th slice have the same size and position).
[0231] If sps_single_slice_per_subpic_flag is equal to 0, the slice layout is present in the SPS, and pps_single_slice_per_subpic_flag may be equal to 1 or 0. If sps_single_slice_per_subpic_flag is equal to 1, the SPS specifies the subpicture partitions. The i-th slice has a size and position corresponding to the i-th subpicture (i.e., the i-th slice and the i-th subpicture have the same size and position).
[0232] In order to maintain the same sub-picture layout for a coded video sequence, an encoder may constrain all PPSs that reference an SPS with sps_single_slice_per_subpic_flag equal to 1 to describe / define / impose the same slice partitioning.
[0233] In a variant, another flag (pps_single_slice_per_tile) is provided to indicate that the PPS has one slice per tile. If this flag is equal to 1, the slice partitioning is inferred to be equal to (i.e., the same as) the tile partitioning. In such a case, if sps_single_slice_per_subpic_flag is equal to 1, the subpicture and slice partitioning is inferred to be the same as the tile partitioning.
[0234] Implementation of the embodiments of the present invention One or more of the above embodiments / variations may be implemented in the form of an encoder or decoder that performs the method steps of one or more of the above embodiments / variations. The following embodiments illustrate such implementations.
[0235] FIG. 9a is a flow chart illustrating steps of an encoding method according to an embodiment / variant of the invention, and FIG. 9b is a flow chart illustrating steps of a decoding method according to an embodiment / variant of the invention.
[0236] According to the encoding method of Figure 9a, sub-picture partition information is obtained at 9911 and slice partition information is obtained at 9912. Using this obtained information, information for determining one or more of the number of slices in a sub-picture, whether only a single slice is included in the sub-picture, and / or whether a slice can include a tile fraction is determined at 9915. Data for obtaining this determined information is then encoded at 9919, for example by providing the data in a bitstream.
[0237] According to the decoding method of Figure 9b, at 9961, data is decoded (e.g., from a bitstream) to obtain information for determining the number of slices in a subpicture, whether only a single slice is included in the subpicture, and / or whether a slice can include a tile fraction. At 9964, this obtained information is used to determine one or more of: the number of slices in a subpicture, whether only a single slice is included in the subpicture, and / or whether a slice can include a tile fraction. Then, at 9967, subpicture partition information and / or slice partition information is determined based on this determination and its results.
[0238] It will be understood that any of the above embodiments / variations may be used by the encoder of FIG. 10 (e.g., when performing the division into blocks 9402, entropy encoding 9409, and / or bitstream generation 9410) or the decoder of FIG. 11 (e.g., when performing the bitstream processing 9561, entropy decoding 9562, and / or video signal generation 9569).
[0239] Fig. 10 shows a block diagram of an encoder according to an embodiment of the invention. The encoder is represented by connected modules, each module adapted to carry out at least one corresponding step of a method implementing at least one embodiment of encoding images of a sequence of images according to one or more embodiments / variants of the invention, for example in the form of program instructions to be executed by a central processing unit (CPU) of the device.
[0240] An original sequence of digital images i0 to in9401 is received as input by the encoder 9400. Each digital image is represented by a set of samples, sometimes also called picture elements (hereafter referred to as pixels). A bitstream 9410 is output by the encoder 9400 after the implementation of the encoding process. The bitstream 9410 contains data of a number of image parts or coding units, such as slices, each slice including a slice header for transmitting coded values of coding parameters used to code the slice, and a slice body containing the coded video data. The input digital images i0 to in9401 are divided into blocks of pixels by the module 9402. The blocks correspond to image parts (hereafter an image part represents any type of part of an image, such as a tile, a slice, a slice segment or a subpicture) and may be of variable size (for example 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes can also be considered). An encoding mode is selected for each input block.
[0241] Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal prediction (e.g. inter coding, MERGE, SKIP). Possible coding modes are tested. Module 9403 performs an intra prediction process in which a given block to be coded is predicted by a predictor calculated from pixels in the neighborhood of said block to be coded. An indication of the selected intra predictor and the difference between the given block and its predictor are coded to provide a residual in case intra coding is selected. Temporal prediction is implemented by a motion estimation module 9404 and a motion compensation module 9405. First, a reference image is selected from a set of reference images 9416, and a portion of the reference image, also called a reference region or image portion, that is the closest region (closest in terms of pixel value similarity) to the given block to be coded is selected by the motion estimation module 9404. Then, the motion compensation module 9405 uses the selected region to predict the block to be coded. The difference between the selected reference area and the given block, also called residual block / data, is calculated by the motion compensation module 9405. The selected reference area is indicated with motion information (e.g., motion vector). Thus, in both cases (spatial prediction and temporal prediction), the residual is calculated by subtracting a predictor from the original block if the original block is not in skip mode. In intra prediction implemented by the module 9403, the prediction direction is coded. In inter prediction implemented by the modules 9404, 9405, 9416, 9418, 9417, at least one motion vector or information (data) for identifying such a motion vector is coded for temporal prediction. If inter prediction is selected, the motion vector and information related to the residual block are coded. To further reduce the bit rate, assuming that the motion is uniform, the motion vector is coded by the difference to the motion vector predictor. A motion vector predictor from a set of motion information predictor candidates is obtained from the motion vector field 9418 by the motion vector predictive coding module 9417.The encoder 9400 further includes a selection module 9406 for selecting a coding mode by applying a coding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (such as a DCT) is applied to the residual block by a transform module 9407, and the resulting transformed data is then quantized by a quantization module 9408 and entropy coded by an entropy coding module 9409. Finally, the coded residual block of the current block being coded is inserted into the bitstream 9410 if it is not in skip mode and the selected coding mode requires coding of the residual block.
[0242] The encoder 9400 also performs decoding of the encoded image to generate reference images (e.g., those in the reference image / picture 9416) for motion estimation of subsequent images. This allows the encoder and decoder receiving the bitstream to have the same reference frame (e.g., a reconstructed image or a reconstructed image portion is used). The inverse quantization ("inverse quantization") module 9411 performs inverse quantization ("inverse quantization") of the quantized data, followed by an inverse transformation performed by the inverse transformation module 9412. The intra prediction module 9413 uses the prediction information to decide which predictor should be used for a given block, and the motion compensation module 9414 actually adds the residual obtained by the module 9412 to a reference area obtained from the set of reference images 9416. Then, post-filtering is applied by the module 9415 to filter the reconstructed frame of pixels (image or image portion) to obtain another reference image for the set of reference images 9416.
[0243] 11 shows a block diagram of a decoder 9560 that may be used to receive data from an encoder according to one embodiment of the present invention. The decoder is represented by connected modules, each module adapted to carry out a corresponding step of a method implemented by the decoder 9560, for example in the form of program instructions executed by the device's CPU.
[0244] The decoder 9560 receives a bitstream 9561 containing coding units (e.g. data corresponding to image portions, blocks or coding units), each coding unit consisting of a header containing information about coding parameters and a body containing the coded video data. As explained with respect to FIG. 10, the coded video data is entropy coded, and the motion information (e.g. an index of a motion vector predictor) is coded with a predefined number of bits for a given image portion (e.g. a block or CU). The received coded video data is entropy decoded by module 9562. The residual data is then inverse quantized by module 9563 and then an inverse transform is applied by module 9564 to obtain pixel values.
[0245] Mode data indicating the coding mode is also entropy decoded, and based on this mode, intra-type or inter-type decoding is performed on the coded block (unit / set / group) of image data. In the case of intra-mode, the intra-predictor is determined by the intra-prediction module 9565 based on the intra-prediction mode specified in the bitstream (e.g., the intra-prediction mode is determinable using data provided in the bitstream). If the mode is inter-mode, motion prediction information is extracted / obtained from the bitstream to find (identify) the reference region used by the encoder. The motion prediction information includes, for example, a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 9570 to obtain a motion vector. The motion vector decoding module 9570 applies motion vector decoding to each image portion (e.g., current block or CU) coded by motion prediction. Once the index of the motion vector predictor of the current block is obtained, the actual value of the motion vector associated with the image portion (e.g., current block or CU) can be decoded and used to apply motion compensation by the module 9566. The reference image portion indicated by the decoded motion vector is extracted / obtained from a set of reference images 9568 so that module 9566 can perform motion compensation. The motion vector field data 9571 is updated with the decoded motion vector to be used for prediction of the motion vector to be decoded later. Finally, a decoded block is obtained. If appropriate, post-filtering is applied by a post-filtering module 9567. A decoded video signal 9569 is finally obtained and provided by the decoder 9560.
[0246] FIG. 12 illustrates a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system includes a transmitting device, in this case a server 9201, operable to transmit data packets of a data stream 9204 to a receiving device, in this case a client terminal 9202, via a data communication network 9200. The data communication network 9200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network made up of several different networks. In certain embodiments of the present invention, the data communication system may be a digital television broadcasting system in which a server 9201 transmits the same data content to multiple clients. The data stream 9204 provided by the server 9201 may be composed of multimedia data representing video and audio data. The audio and video data streams may be captured by the server 9201 using a microphone and a camera, respectively, in some embodiments of the present invention. In some embodiments, the data stream may be stored on the server 9201, or may be received by the server 9201 from another data provider, or may be generated at the server 9201. The server 9201 in particular comprises an encoder for encoding the video and audio streams to provide a compressed bitstream for transmission, which is a more compact representation of the data presented as input to the encoder. In order to obtain a better ratio of the quality of the transmitted data to the amount of the transmitted data, the compression of the video data may for example be according to the High Efficiency Video Coding (HEVC) format, or the H.264 / AVC (Advanced video Coding) format, or the VVC (Versatile video Coding) format.The client 9202 receives the transmitted bitstream and decodes the reconstructed bitstream to play the video image on a display device and the audio data by a speaker. Although a streaming scenario is considered in this embodiment, it will be understood that in some embodiments of the invention, data communication between the encoder and the decoder may be performed using a media storage device, such as an optical disk, for example. In one or more embodiments of the invention, the video image may be transmitted along with data representing a compensation offset to be applied to the reconstructed pixels of the image to provide filtered pixels in the final image.
[0247] Fig. 13 shows a schematic representation of a processing device 9300 adapted to implement at least one embodiment / variant of the invention. The processing device 9300 can be a device such as a microcomputer, a workstation, a user terminal or a light portable device. The device / apparatus 9300 comprises: a central processing unit 9311, such as a microprocessor, indicated by CPU; a read-only memory 9307, indicated by ROM, for storing computer programs / instructions for operating the device 9300 and / or for implementing the invention; a random access memory 9312, indicated by RAM, for storing executable codes of the methods of the embodiments / variants of the invention as well as registers adapted for recording variables and parameters necessary for implementing the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the embodiments / variants of the invention; and a communication interface 9302, connected to a communication network 9303, over which the digital data to be processed are transmitted and received, a communication bus 9313 connected to the communication interface 9302.
[0248] Optionally, the device 9300 may also include the following components: a data storage means 9304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments / variations of the invention, and data used or generated during the implementation of one or more embodiments / variations of the invention; a disk drive 9305 for a disk 9306 (e.g., a storage medium), adapted to read data from the disk 9306 or to write data to said disk 9306; or a screen 9309 for displaying data and / or serving as a graphical interface with a user, by means of a keyboard 9310, a touch screen, or any other indication / input means. The device 9300 may be connected to various peripherals, such as for example a digital camera 9320 or a microphone 9308, each of which is connected to an input / output card (not shown) to provide multimedia data to the device 9300. A communication bus 9313 provides communication and interoperability between the various elements included in or connected to the device 9300. The representation of the bus is not limiting, in particular the central processing unit 9311 is operable to communicate instructions to any element of the device 9300 directly or by another element of the device 9300. The disk 9306 can be replaced by any information carrier, such as for example a compact disk (CD-ROM), a ZIP disk or a memory card, rewritable or not, and generally speaking is an information storage means readable by a microcomputer or processor, integrated or not integrated in the device, possibly removable, and configured to store one or more programs whose execution enables the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the invention to be performed. The executable code can be stored either in the read-only memory 9307, the hard disk 9304, or in a removable digital medium, such as for example the disk 9306, as previously described.According to a variant, the executable code of the program can be received by the communication network 9303, via the interface 9302, for example to be stored in one of the storage means of the device 9300 before being executed in the hard disk 9304. The central processing unit 9311 is configured to control and direct the execution of the instructions or parts of the program or software code of the program according to the invention, with instructions stored in one of the aforementioned storage means. On power-up, the program or programs stored in a non-volatile memory, for example on the hard disk 9304, the disk 9306 or the read-only memory 9307, are transferred to the random access memory 9312, which then contains the executable code of the program or programs, as well as registers for storing variables and parameters necessary for implementing the invention. In this embodiment, the device is a programmable device that uses software to implement the invention. However, alternatively, the invention may be implemented in hardware (for example in the form of an application specific integrated circuit or ASIC).
[0249] Implementation of the embodiments of the present invention It is also understood that according to other embodiments of the invention, a decoder according to the aforementioned embodiment / variant is provided in a user terminal such as a computer, a mobile phone (cell phone), a tablet or any other type of device (e.g. a display device) capable of providing / displaying content to a user. According to yet another embodiment, an encoder according to the aforementioned embodiment / variant is provided in an image capture device, also comprising a camera, a video camera or a network camera (e.g. a closed circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such embodiments are provided below with reference to figures 14 and 15.
[0250] FIG. 14 is a diagram illustrating a network camera system 9450 including a network camera 9452 and a client device 9454. The network camera 9452 includes an imaging unit 9456, an encoding unit 9458, a communication unit 9460, and a control unit 9462. The network camera 9452 and the client device 9454 are communicatively connected to each other via a network 9200. The imaging unit 9456 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)), captures an image of an object, and generates image data based on the image. The image may be a still image or a video image. The imaging unit may also include zoom means and / or pan means, which are adapted to zoom or pan (optically or digitally), respectively. The encoding unit 9458 encodes the image data by using the encoding method described in one or more of the above embodiments / variations. The encoding unit 9458 uses at least one of the encoding methods described in the above embodiments / variations. As another example, the encoding unit 9458 can use a combination of the encoding methods described in the previous embodiments / variations. The communication unit 9460 of the network camera 9452 transmits the encoded image data encoded by the encoding unit 9458 to the client device 9454. Furthermore, the communication unit 9460 may receive commands from the client device 9454. The commands include commands for setting parameters for encoding by the encoding unit 9458. The control unit 9462 controls other units in the network camera 9452 according to the commands received by the communication unit 9460 and user input. The client device 9454 includes a communication unit 9464, a decoding unit 9466, and a control unit 9468. The communication unit 9464 of the client device 9454 may transmit commands to the network camera 9452. Furthermore, the communication unit 9464 of the client device 9454 receives the encoded image data from the network camera 9452.The decoding unit 9466 decodes the encoded image data by using the decoding method described in one or more of the above-mentioned embodiments / variations. As another example, the decoding unit 9466 can use a combination of the decoding methods described in the above-mentioned embodiments / variations. The control unit 9468 of the client device 9454 controls other units in the client device 9454 according to a user operation or command received by the communication unit 9464. The control unit 9468 of the client device 9454 may also control the display device 9470 to display an image decoded by the decoding unit 9466. The control unit 9468 of the client device 9454 may also control the display device 9470 to display a GUI (Graphical User Interface) for specifying the values of parameters of the network camera 9452, for example, the values of parameters for encoding by the encoding unit 9458. The control unit 9468 of the client device 9454 may also control other units in the client device 9454 according to a user's operation input to a GUI displayed by the display device 9470. In addition, the control unit 9468 of the client device 9454 may control the communication unit 9464 of the client device 9454 to send a command to the network camera 9452 specifying the value of a parameter of the network camera 9452 in response to a user operation input to a GUI displayed by the display device 9470.
[0251] FIG. 15 is a diagram showing a smartphone 9500. The smartphone 9500 includes a communication unit 9502, a decoding / encoding unit 9504, a control unit 9506, and a display unit 9508. The communication unit 9502 receives encoded image data via the network 9200. The decoding / encoding unit 9504 decodes the encoded image data received by the communication unit 9502. The decoding / encoding unit 9504 decodes the encoded image data by using the decoding method described in one or more of the above embodiments / variations. The decoding / encoding unit 9504 may also use at least one of the encoding method or the decoding method described in the above embodiments / variations. In another example, the decoding / encoding unit 9504 may use a combination of the decoding method or the encoding method described in the above embodiments / variations. The control unit 9506 controls other units in the smartphone 9500 according to a user operation or a command received by the communication unit 9502. For example, the control unit 9506 controls the display unit 9508 to display images decoded by the decoding / encoding unit 9504. The smartphone may further include an image recording device 9510 (e.g., a digital camera and associated circuitry) for recording images or videos. Such recorded images or videos may be encoded by the decoding / encoding unit 9504 under the direction of the control unit 9506. The smartphone may further include a sensor 9512 configured to sense an orientation of the mobile device. Such sensors may include an accelerometer, gyroscope, compass, global positioning (GPS) unit, or similar position sensor. Such a sensor 9512 may determine whether the smartphone changes orientation, and such information may be used when encoding the video stream.
[0252] Although the present invention has been described with reference to embodiments and variations thereof, it should be understood that the present invention is not limited to the disclosed embodiments / variations. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the present invention, as defined in the appended claims. All of the features disclosed in this specification (including any accompanying claims, abstract, and drawings), and / or all of the steps of any method or process so disclosed, can be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings) can be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless otherwise specified. Thus, unless otherwise specified, each feature disclosed is merely one example of a generic series of equivalent or similar features.
[0253] It is also understood that any result of the above-mentioned comparison, determination, inference, evaluation, selection, execution, performance, or consideration, e.g., a selection made during an encoding, processing, or partitioning process, may be indicated in or determinable / inferable from data in the bitstream, e.g., a flag or information indicating the result, such that the indicated or determined / inferred result may be used in processing instead of actually performing the comparison, determination, evaluation, selection, execution, performance, or consideration, e.g., during a decoding or partitioning process. It is understood that where a "table" or "lookup table" is used, other data types, such as arrays, may also be used to perform the same function, so long as the data type is capable of performing the same function (e.g., representing a relationship / mapping between different elements).
[0254] In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage. Reference numerals appearing in the claims are for illustration purposes only and shall have no limiting effect on the scope of the claims.
[0255] In the above embodiments / variations, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit.
[0256] A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium, which includes any medium that facilitates transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0257] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and disks reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer readable media.
[0258] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.
Claims
1. 1. A method for decoding data for an image, the image may include one or more slices, the slices may correspond to an integer number of consecutive complete coding tree unit rows in a tile, the image may include one or more sub-pictures, The method comprises: obtaining first information indicating a width of a sub-picture and second information indicating a height of the sub-picture from a sequence parameter set; determining parameters associated with the slice included in the sub-picture using the first information and the second information; Decoding the image using at least the determined parameters. Including, determining a parameter associated with the slice using a number of slices included in the sub-picture; In decoding the image, at least intra prediction is used. A method comprising:
2. 2. The method of claim 1, wherein said determining further comprises using a subpicture identifier to determine parameters associated with the slice.
3. 3. The method of claim 1, further comprising determining a parameter associated with a slice based on whether only a single slice is included in the subpicture.
4. 4. The method according to claim 1, wherein the sub-picture comprises two or more slices.
5. 5. A method according to any one of claims 1 to 4, wherein the slice may consist of one or more tiles, the slice forming a rectangular area in the image.
6. 1. A method for encoding an image, comprising the steps of: The image may include one or more slices, which may correspond to an integer number of consecutive complete coding tree unit rows in a tile; The image may include one or more sub-pictures; The method comprises: encoding first information indicating a width of a sub-picture and second information indicating a height of the sub-picture into a sequence parameter set; determining parameters associated with the slice included in the sub-picture using the first information and the second information; encoding said image using at least said determined parameters; Including, determining a parameter associated with the slice using a number of slices included in the sub-picture; In encoding the image, at least intra prediction is used. A method comprising:
7. 7. The method of claim 6, wherein said determining further comprises using a subpicture identifier to determine parameters associated with the slice.
8. 8. The method of claim 6 or 7, further comprising determining a parameter associated with the slice based on whether only a single slice is included in the sub-picture.
9. 9. A method according to any one of claims 6 to 8, wherein the sub-picture comprises two or more slices.
10. 10. A method according to any one of claims 6 to 9, wherein the slice may consist of one or more tiles, the slice forming a rectangular area within the image.
11. An apparatus for decoding image data, comprising: The image may include one or more slices, which may correspond to an integer number of consecutive complete coding tree unit rows in a tile; The image may include one or more sub-pictures; The apparatus comprises: an acquisition means for acquiring first information indicating a width of a sub-picture and second information indicating a height of the sub-picture from a sequence parameter set; a determining means for determining a parameter associated with the slice included in the sub-picture using the first information and the second information; a decoding means for decoding the image using at least the determined parameters; Including, The determining means determines a parameter associated with the slice further using the number of slices included in the sub-picture; The decoding means uses at least intra prediction in decoding the image. An apparatus comprising:
12. 1. An apparatus for encoding an image, comprising: The image may include one or more slices, which may correspond to an integer number of consecutive complete coding tree unit rows in a tile; The image may include one or more sub-pictures; The apparatus comprises: a first encoding means for encoding first information indicating a width of a sub-picture and second information indicating a height of the sub-picture into a sequence parameter set; a determining means for determining a parameter associated with the slice included in the sub-picture using the first information and the second information; a second encoding means for encoding the image using at least the determined parameters; Including, The determining means determines a parameter associated with the slice further using the number of slices included in the sub-picture; The second encoding means uses at least intra prediction in encoding the image. An apparatus comprising:
13. A program for causing a computer to execute the method according to any one of claims 1 to 5.
14. A program for causing a computer to execute the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Sub-picture identifier signaling in video coding
WO2020146662A1
Indication of one slice per subpicture in subpicture-based video coding
WO2021061443A1