Supports video encoding of sub-pictures, slices and blocks
By determining the image part contained in the sub-picture and processing the image data, the complexity problem caused by signal notification independence in the sub-picture partition in VVC7 is solved, and more efficient signal notification and lower resource consumption are achieved.
Patent Information
- Application Number
- CN202080088832.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-17
AI Technical Summary
In VVC7, signaling of sub-picture partitions is independent of signaling of stripes and block grids, resulting in the decoder needing to check and ensure that the image partitions comply with the constraints of VVC7, which can be complex and time-consuming.
By determining the image portion contained in the sub-picture and using the information to process the image data, a method is provided to optimize signaling of the image partition to ensure that the constraints of VVC7 are satisfied during the encoding or decoding process.
This method simplifies the encoding and decoding process of signaling, improves efficiency, reduces resource consumption, and ensures that partitions comply with VVC7 specifications.
Smart Images

Figure CN114846791B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to partitioning of images, and encoding or decoding of images or image sequences comprising images. Embodiments of the invention are particularly, but not exclusively, used in the case of encoding or decoding of image sequences using a first partitioning of the image into one or more sub-pictures and a second partitioning of the image into one or more slices. Background Art
[0002] Video coding includes image coding (an image is equivalent to a single frame, or picture, of a video). In video coding, before some coding tools such as motion compensation / prediction (e.g., inter-frame prediction) or intra-frame prediction can be used on an image, the image is first partitioned (e.g., split) into one or more image parts so that the coding tools can be used on the image parts. The present invention is particularly related to partitioning an image into two types of image parts, namely sub-pictures and slices, which have been studied by the Video Coding Experts Group / Moving Picture Experts Group (VCEG / MPEG) standardization group and are considered for the Versatile Video Coding (VVC) standard.
[0003] Sub-pictures are a new concept introduced in VVC to enable bitstream extraction and merging operations of independent spatial regions (or image portions) from different bitstreams. "Independent" here means that these regions (or image portions) are encoded / decoded without reference to information obtained from encoding / decoding other regions or image portions. For example, independent spatial regions (i.e., regions or image portions that are encoded / decoded without reference to encoding / decoding of other regions / image portions from the same image) are used for region of interest (ROI) streaming (e.g., during 3D video streaming) or for streaming of omnidirectional video content (e.g., image sequences streamed using the Omnidirectional Media Format (OMAF) standard), in particular when viewport-dependent streaming methods are used for streaming. Individual images from omnidirectional video content are segmented into independent regions that are encoded with different versions of them (e.g., in terms of image quality or resolution). The client terminal (e.g., a device with a display, such as a mobile phone, etc.) is then able to select an appropriate version of the independent region to obtain a high-quality version of the independent region in the primary viewing direction, while still being able to use a lower-quality version of the remaining region for the remainder of the omnidirectional video content to improve coding efficiency.
[0004] High Efficiency Video Coding (HEVC or H.265) provides for signaling of motion restricted block sets (e.g., the bitstream includes data for specifying or determining a set of blocks whose motion prediction is restricted to make them "independent" of other areas of the picture) to indicate independently coded areas. In HEVC, this signaling is done in the Supplemental Enhancement Information (SEI) message and is only optional. However, in HEVC, the signaling of slices is made independent of the SEI message, so that the partitioning of an image into one or more slices is defined independently of the partitioning of the same image into one or more sets of blocks. This means that slice partitions do not have the same motion prediction restrictions imposed on them.
[0005] Proposals based on old drafts of the Universal Video Coding Draft 4 (VVC4) include signaling block group partitions that rely on the signaling of sub-pictures. A block group is an integer number of complete blocks of a picture that are exclusively contained in a single network abstraction layer (NAL) unit. This proposal (JVET-N0107: AHG 12: Sub-picture-based coding for VVC, Huawei) is a syntax change proposal for introducing the sub-picture concept of VVC. The sub-picture position is signaled using the luma sample position in the sequence parameter set (SPS). Next, a flag in the SPS indicates whether motion prediction is constrained for each sub-picture, but the block group partition is signaled in the picture parameter set (PPS) (i.e., the picture is partitioned into one or more block groups), where each PPS is defined for each sub-picture. Since a PPS is provided for each sub-picture, a block group partition is signaled for each sub-picture in JVET-N0107.
[0006] However, the latest Universal Video Coding Draft 7 (VVC7) no longer has this block group partition concept. VVC7 signals the sub-picture layout in CTU units in the SPS. The flag in the SPS indicates whether motion prediction is constrained for sub-pictures. The SPS syntax elements for these are as follows:
[0007]
[0008]
[0009] In VVC7, stripe partitions are defined in PPS based on block partitions as follows:
[0010]
[0011]
[0012] This means that in VVC7, slice partitioning is defined independently of sub-picture partitioning. The VVC7 syntax for slice partitioning is performed on the block structure without reference to sub-pictures, because this independence of sub-pictures avoids any specific processing of sub-pictures during the encoding / decoding process, making the processing of sub-pictures simpler. Summary of the invention
[0013] As mentioned above, VVC7 provides several tools to partition pictures into pixel regions (or component samples). Some examples of these tools are sub-pictures, strips, and blocks. In order to accommodate all these tools while maintaining their functionality, VVC7 imposes some constraints on partitioning pictures into these regions. For example, blocks must have a rectangular shape, and blocks must form a grid. A strip can be an integer number of blocks or a fragment of a block (i.e., a strip includes only a portion of a block or a "partial block" or a "fragment block"). A sub-picture is a rectangular area that must contain one or more strips. However, in VVC7, the signaling of sub-picture partitioning is independent of the signaling of strips and block grids. Therefore, this signaling in VVC7 requires the decoder to check and ensure that the picture partitioning complies with the constraints of VVC7, which can be complex and lead to unnecessary time or resource consumption at the decoder end.
[0014] Embodiments of the present invention aim to address one or more problems or disadvantages of the aforementioned partitioning of images and the encoding or decoding of images or image sequences comprising the images. For example, one or more embodiments of the present invention aim to improve and optimize the signaling of picture partitioning (e.g., within the VVC7 context) while ensuring that at least some of the constraints that need to be checked in VVC7 are achieved / satisfied by design during the signaling or encoding process.
[0015] According to various aspects of the present invention, there are provided devices / apparatus, methods, programs, computer-readable storage media and carrier media / signals as set forth in the appended claims. Other features of the present invention will be apparent from the dependent claims and the specification. According to other aspects of the present invention, there are provided systems as set forth in the appended claims, methods for controlling such systems, devices / apparatus for performing methods, devices / apparatus for processing, medium storage devices storing signals as set forth in the appended claims, computer-readable storage media or non-transitory computer-readable storage media storing programs as set forth in the appended claims, and bitstreams generated using the encoding methods as set forth in the appended claims. Other features of the present invention will be apparent from the dependent claims and the subsequent specification.
[0016] According to a first aspect of the present invention, a method for processing image data of one or more images is provided, each image consisting of one or more blocks and capable of being divided into one or more image parts, wherein the image can be divided into one or more sub-pictures, and the method comprises: determining one or more image parts included in the sub-picture; and using information obtained from the determination to process the one or more images.
[0017] According to a second aspect of the present invention, a method for partitioning one or more images is provided, the method comprising: partitioning the image into one or more blocks; partitioning the image into one or more sub-pictures; and partitioning the image into one or more image parts by processing image data of the image according to the first aspect.
[0018] According to a third aspect of the invention, there is provided a method of signalling partitioning of one or more images, the method comprising: processing image data of one or more images according to the first aspect; and signalling information for determining the partitioning in a bitstream.
[0019] For the aforementioned aspects of the present invention, the following features may be provided according to embodiments of the present invention. Suitably, the image portion may include a portion of a block. Suitably, the image portion is encoded in a single logical unit (e.g., in a network abstraction layer unit or a NAL unit) or decoded from a single logical unit (e.g., signaled, transmitted, provided, or obtained from a single logical unit). Suitably, blocks and / or sub-pictures are not encoded in a single logical unit (e.g., a NAL unit) or decoded from a single logical unit (e.g., signaled, transmitted, provided, or obtained from a single logical unit).
[0020] According to a fourth aspect of the present invention, there is provided a method for processing image data of one or more images, each image consisting of one or more blocks and being divisible into one or more image parts, wherein the image part can include a part of the block (partial block), and the image can be divisible into one or more sub-pictures, and the method comprises: determining one or more image parts included in the sub-picture; and using information obtained from the determination to process the one or more images. The part of the block (partial block) may be an integer number of consecutive complete coding tree unit (CTU) rows within the block.
[0021] For the aforementioned aspects of the invention, the following features may be provided according to embodiments of the invention. Suitably, the determining comprises defining the one or more image parts using one or more of the following: an identifier of a sub-picture; a size, width or height of the sub-picture; whether only a single image part is included in the sub-picture; and the number of image parts included in the sub-picture.
[0022] Suitably, when the number of image parts included in the sub-picture is greater than 1, each image part is determined based on the number of blocks included therein.
[0023] Suitably, when the image portion comprises one or more than one parts of a block (partial blocks), the image portion is determined based on the number of rows or columns of coding tree units (CTUs) to be comprised therein.
[0024] Suitably, the processing comprises providing in a picture parameter set (PPS) or obtaining from the PPS information for determining an image portion based on the number of blocks; and when the image portion comprises one or more parts (partial blocks) of blocks, providing in a header of one or more logical units of encoded data comprising the image portion or obtaining from the header information for identifying the image portion comprising one or more parts (partial blocks) of blocks.
[0025] Suitably, the image portion consists of a sequence of tiles in a tile raster scan order.
[0026] Suitably, the processing comprises providing in the bitstream or obtaining from the bitstream information for determining one or more of the following: whether only a single image portion is included in the sub-picture; the number of image portions included in the sub-picture. Suitably, the processing comprises providing in the bitstream or obtaining from the bitstream whether use of an image portion comprising a portion of a tile (partial tile) is permitted when processing one or more images.
[0027] Suitably, the information provided in or obtained from the bitstream comprises: information indicating whether a sub-picture is used in the video sequence; and when the information indicates that a sub-picture is not used in the video sequence, one or more image portions of the video sequence are determined as not being allowed to include a portion of a block (partial block).
[0028] Suitably, the information used for the determination is provided in a picture parameter set (PPS) or obtained from the PPS.
[0029] Suitably, the information used for the determination is provided in a sequence parameter set (SPS) or obtained from the SPS.
[0030] Suitably, when the information used for determining indicates that the number of image parts included in the sub-picture is 1, the sub-picture is composed of a single image part that does not include a part of the tile (partial tile).
[0031] Suitably, the sub-picture comprises two or more image parts, each image part comprising one or more parts of a block (part-blocks).
[0032] Suitably, one or more parts of a block (part-blocks) are from the same single block.
[0033] Suitably, the two or more image parts can comprise one or more parts of a block (partial blocks) from more than one block.
[0034] Suitably, the image portion consists of a plurality of blocks, said image portion forming a rectangular area in the image.
[0035] Suitably, the image parts are stripes (one or more than one image parts are one or more than one stripe).
[0036] According to a fifth aspect of the present invention, a method for encoding one or more images is provided, the method comprising any one of the following items: processing image data according to the first aspect or the fourth aspect; partitioning according to the second aspect; and / or signaling according to the third aspect.
[0037] Suitably, the method further comprises: receiving an image; processing image data of the received image according to the first aspect or the fourth aspect; and encoding the received image and generating a bitstream.
[0038] Suitably, the method further comprises providing one or more of the following items in a bitstream: information in a picture parameter set (PPS) for determining an image portion based on the number of blocks, and when the image portion comprises one or more parts (partial blocks) of a block, information in a header of one or more logical units of encoded data comprising the image portion, a slice segment header or a slice header for identifying the image portion comprising one or more parts (partial blocks) of a block; in the PPS, information for determining whether only a single image portion is included in a sub-picture; in the PPS, information for determining the number of image portions included in a sub-picture; in a sequence parameter set (SPS), information for determining whether use of an image portion comprising a part of a block (partial block) is allowed when processing one or more images; and in the SPS, information indicating whether a sub-picture is used in a video sequence.
[0039] According to a sixth aspect of the present invention, a method for decoding one or more images is provided, the method comprising any one of the following items: processing image data according to the first aspect or the fourth aspect; partitioning according to the second aspect; and / or signaling according to the third aspect.
[0040] Suitably, the method further comprises: receiving a bitstream; decoding information from the received bitstream and processing image data according to any one of the first aspect or the fourth aspect; and obtaining an image using the decoded information and the processed image data.
[0041] Suitably, the method further comprises obtaining one or more of the following items from a bitstream: information from a picture parameter set (PPS) for determining an image portion based on the number of blocks; and when the image portion comprises one or more parts (partial blocks) of a block, information from a header of one or more logical units of encoded data comprising the image portion, a slice segment header or a slice header for identifying the image portion comprising one or more parts (partial blocks) of a block; information from the PPS for determining whether only a single image portion is included in a sub-picture; information from the PPS for determining the number of image portions included in a sub-picture; information from a sequence parameter set (SPS) for determining whether use of an image portion comprising a part of a block (partial block) is allowed when processing one or more images; and information from the SPS indicating whether a sub-picture is used for a video sequence.
[0042] According to a seventh aspect of the present invention, there is provided an apparatus for processing image data of one or more images, the apparatus being configured to perform a method according to any one of the first, fourth, second and third aspects.
[0043] According to an eighth aspect of the present invention, there is provided an apparatus for encoding one or more images, the apparatus comprising processing apparatus according to the seventh aspect. Suitably, the apparatus is configured to perform the method according to the fifth aspect.
[0044] According to a ninth aspect of the present invention, there is provided an apparatus for decoding one or more images, the apparatus comprising the processing apparatus according to the seventh aspect. Suitably, the apparatus is configured to perform the method according to the sixth aspect.
[0045] According to a tenth aspect of the present invention, there is provided a program which, when executed on a computer or a processor, causes the computer or the processor to execute the method according to the first, fourth, second or third, fifth or sixth aspect.
[0046] According to an eleventh aspect of the present invention, there is provided a carrier medium or a computer-readable storage medium carrying / storing the program of the tenth aspect.
[0047] According to the twelfth aspect of the present invention, there is provided a signal carrying an information data set for an image, the image being encoded using the method according to the fifth aspect and represented by a bit stream, the image being composed of one or more blocks and being capable of being divided into one or more image parts, wherein the image part may include a part of a block (partial block), and the image being capable of being divided into one or more sub-pictures, wherein the information data set includes data for determining one or more image parts included in a sub-picture.
[0048] Further aspects of the invention relate to a program which, when executed by a computer or processor, causes the computer or processor to perform any of the methods of the foregoing aspects. The program may be provided separately, or may be carried on, by or in a carrier medium. The carrier medium may be non-transitory, such as a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transient, such as a signal or other transmission medium. The signal may be transmitted via any suitable network, including the Internet.
[0049] Further aspects of the invention relate to a camera comprising an apparatus according to any one of the aforementioned apparatus aspects. According to yet another aspect of the invention, there is provided a mobile device comprising an apparatus according to any one of the aforementioned apparatus aspects and / or a camera embodying the above-mentioned camera aspects.
[0050] Any feature in one aspect of the present invention may be applied to other aspects of the present invention in any appropriate combination. In particular, method aspects may be applied to device aspects, and vice versa. In addition, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be interpreted accordingly. Any device features as described herein may also be provided as method features, and vice versa. As used herein, component plus function features may be expressed alternatively according to their corresponding structures (such as appropriately programmed processors and associated memories). It should also be understood that specific combinations of various features described and defined in any aspect of the present invention may be independently implemented and / or provided and / or used.
[0051] By referring to the following description of the embodiments of the present invention with reference to the accompanying drawings, further features, aspects and advantages of the present invention will become apparent. The various embodiments of the present invention described below can be implemented individually or as a combination of multiple embodiments. In addition, features from different embodiments can be combined when necessary or when a combination of elements or features from various embodiments is beneficial in a single embodiment. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which:
[0053] Figure 1 Shows partitioning of a picture into blocks and slices according to an embodiment of the present invention;
[0054] Figure 2 shows sub-picture partitioning of an image according to an embodiment of the present invention;
[0055] Figure 3 shows a bit stream according to an embodiment of the present invention;
[0056] Figure 4 is a flow chart illustrating an encoding process according to an embodiment of the present invention;
[0057] Figure 5 is a flowchart illustrating a decoding process according to an embodiment of the present invention;
[0058] Figure 6 is a flow chart illustrating determination steps used when signaling stripe partitions according to an embodiment of the present invention;
[0059] Figure 7 An example of sub-picture and slice partitioning according to an embodiment of the present invention is shown;
[0060] Figure 8 An example of sub-picture and slice partitioning according to an embodiment of the present invention is shown;
[0061] Figure 9a is a flow chart showing the steps of an encoding method according to an embodiment of the present invention;
[0062] Figure 9b is a flow chart showing the steps of a decoding method according to an embodiment of the present invention;
[0063] Fig.10 is a block diagram showing steps of an encoding method according to an embodiment of the present invention;
[0064] Fig.11 is a block diagram showing steps of a decoding method according to an embodiment of the present invention;
[0065] Fig.12 is a block diagram schematically illustrating a data communication system in which one or more embodiments of the present invention may be implemented;
[0066] Fig.13 is a block diagram illustrating components of a processing device that may implement one or more embodiments of the present invention;
[0067] Fig.14is a diagram illustrating a network camera system in which one or more embodiments of the present invention may be implemented; and
[0068] Fig.15 is a diagram illustrating a smart phone in which one or more embodiments of the present invention may be implemented. DETAILED DESCRIPTION
[0069] The embodiments of the invention described below are directed to improving the encoding and decoding of images (or pictures).
[0070] In the present specification, "signaling" may refer to inserting (providing / including / encoding in) information about one or more parameters or syntax elements into a bitstream or extracting / obtaining (decoding) the information from a bitstream, the information being, for example, any one or more of information for determining an identifier of a sub-picture, a size / width / height of a sub-picture, whether only a single image portion (e.g., a slice) is included in a sub-picture, whether a slice is a rectangular slice, and / or the number of slices included in a sub-picture. In the present specification, "processing" may refer to any type of operation performed on data, for example, encoding or decoding image data of one or more images / pictures.
[0071] In this specification, the term "slice" is used as an example of an image portion (other examples of such an image portion would be an image portion comprising one or more coding tree units). It should be understood that embodiments of the present invention may also be implemented based on image portions instead of slices and appropriately modified parameters / values / syntax, such as an image portion header (instead of a slice header or a slice segment header). It should also be understood that various information described herein as being signaled in a slice header, a slice segment header, a sequence parameter set (SPS), or a picture parameter set (PPS) may be signaled elsewhere as long as it is able to provide the same functionality provided by signaling the information in these mediums. It should also be understood that any of a slice, a block group, a block, a coding tree unit (CTU) / largest coding unit (LCU), a coding tree block (CTB), a coding unit (CU), a prediction unit (PU), a transform unit (TU), or a pixel / sample block may be referred to as an image portion.
[0072] It should also be understood that when a component or tool is described as "active", the component / tool is "enabled" or "usable" or "used"; when described as "inactive", the component / tool is "disabled" or "unavailable" or "not used"; and "can be inferred" means that the relevant value or parameter can be determined / obtained from other information without explicit signaling in the bitstream. In addition, it should also be understood that when a flag is described as "active", it means that the flag indicates that the relevant component / tool is "active" (i.e., "valid").
[0073] In this specification, unless otherwise specified, the following terms are used with the same or functionally equivalent definitions as they are defined in VVC7. The definitions used in VVC7 are as follows.
[0074] Slice: An integer number of complete blocks of a picture, or an integer number of consecutive complete CTU rows within a block, contained exclusively in a single NAL unit.
[0075] Slice header: A portion of a coded slice that contains data elements related to all blocks or CTU rows within blocks represented in the slice.
[0076] Block: A rectangular area of a CTU within a specific block column and a specific block row in a picture.
[0077] Sub-image: A rectangular area of one or more strips within an image.
[0078] Picture (or image): An array of luma samples in monochrome format or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0079] Coded picture: A coded representation of a picture that includes VCL NAL units with a specific value of nuh_layer_id within an AU and includes all CTUs of the picture.
[0080] Encoded representation: A data element represented in an encoded form.
[0081] Raster Scan: The mapping of a rectangular two-dimensional pattern to a one-dimensional pattern such that the first entry in the one-dimensional pattern is from the top row of the two-dimensional pattern scanned from left to right, followed similarly by the second, third, etc. rows of the pattern (downwards), each scanned from left to right.
[0082] Block: an M×N (M columns×N rows) array of samples, or an M×N array of transform coefficients.
[0083] Coding block: A block of M×N samples for some values of M and N such that the partitioning of CTBs into coding blocks is a partition.
[0084] Coding Tree Block (CTB): An NxN block of samples for some value of N such that the partitioning of components into CTBs is a partition.
[0085] Coding Tree Unit (CTU): A CTB of luma samples, two corresponding CTBs of chroma samples of a picture with three sample arrays, or a CTB of samples of a monochrome picture or a picture coded using three separate color planes and a syntax structure for coding the samples.
[0086] Coding Unit (CU): A coding block of luma samples of a picture with three sample arrays, two corresponding coding blocks of chroma samples, or a coding block of samples of a monochrome picture or a picture coded using three separate color planes and a syntax structure for coding the samples.
[0087] Component: An array or a single sample from one of the three arrays (luminance and two chrominance) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample from an array that makes up a picture in monochrome format.
[0088] Picture Parameter Set (PPS): A syntax structure containing syntax elements that apply to zero or more entire coded pictures as determined by the syntax elements found in each slice header.
[0089] Sequence Parameter Set (SPS): A syntax structure containing syntax elements that apply to zero or more complete CVSs as determined by the contents of the syntax elements found in the PPS referenced by the syntax elements found in each slice header.
[0090] In this specification, unless otherwise specified, the following terms are also used in the same or functionally equivalent definitions as they are defined below.
[0091] Block group: An integer number of complete (ie, whole) blocks of a picture that are exclusively contained in a single NAL unit.
[0092] “Block slice”, “partial block”, “portion of a block” or “slice of a block”: an integer number of consecutive complete CTU rows within a block of a picture that do not form a complete (ie, entire) block.
[0093] Slice segment: A picture that contains exclusively an integer number of complete blocks in a single NAL unit or an integer number of consecutive complete CTU rows within a block.
[0094] Slice segment header: A portion of an encoded slice segment that contains data elements related to all blocks or CTU rows within blocks represented in the slice segment.
[0095] Slice when slice segments are present: A collection of one or more slice segments that together represent a block or an integer number of complete blocks of a picture
[0096] Embodiments of the present invention
[0097] Partitioning and bitstreaming of pictures / images
[0098] 3.1 Partition the image into blocks and strips
[0099] In most coding systems (such as HEVC or the emerging VVC standard), the compression of video relies on block-based video coding. In these coding systems, the video consists of a sequence of frames or pictures or images or samples that can be displayed at different time points (e.g., at different time positions within the video). In the case of multi-layer video (e.g., scalable, stereoscopic or 3D video), it may be necessary to decode several pictures to be able to form the final / resulting image to be displayed at a specific time point. A picture may also consist of more than one image component (i.e., the image data of the picture includes more than one image component). Examples of such image components would be components for encoding brightness, chrominance or depth information.
[0100] Compression of video sequences uses several different partitioning techniques (ie, different schemes / frameworks / arrangements / mechanisms for partitioning / splitting pictures) for individual pictures and how these partitioning techniques are implemented during the compression process.
[0101] Figure 1 The partitioning of a picture into blocks and slices according to an embodiment of the present invention is shown, which is compatible with VVC7. Pictures 101 and 102 are divided into coding tree units (CTUs) represented by dotted lines. CTU is the basic unit of encoding and decoding of VVC7. For example, in VVC7, a CTU can encode an area of 128×128 pixels.
[0102] A coding tree unit (CTU) may also be referred to as a block (of pixels or component samples (values)), a macroblock, or even a coding block. A coding tree unit may be used to encode / decode different image components of a picture simultaneously, or may be limited to only one image component so that different image components of a picture may be encoded / decoded separately / independently. When the data of an image includes separate data for each component, a CTU groups multiple coding tree blocks (CTBs), one CTB for each component.
[0103] like Figure 1 As shown, the picture can also be partitioned according to a grid of blocks (i.e., divided into one or more grids of blocks) represented by thin solid lines. A block is a picture portion (a part / portion of a picture) that is a rectangular area (of pixels / component samples) that can be defined independently of the CTU partition. For example, in VVC7, a block can also correspond to a sequence of CTUs to be Figure 1 In the example represented in , the partitioning technique may constrain the boundaries of blocks to be consistent / aligned with the boundaries of CTUs.
[0104] Blocks are defined so that block boundaries break spatial dependencies of the encoding / decoding process (i.e., in a given picture, a block is defined / specified so that it can be encoded / decoded independently of other spatially "neighboring" blocks of the same picture). This means that encoding / decoding of CTUs in a block is not based on pixels / samples or reference data from other blocks in the same picture.
[0105] Some encoding / decoding systems (e.g., embodiments of the present invention or embodiments for VVC7) provide the concept of a stripe (i.e., also use a partitioning technique based on one or more stripes). This mechanism enables a picture to be partitioned into one or several block groups, which are collectively referred to as stripes. Each stripe consists of one or several blocks or partial blocks. As shown in pictures 101 and 102, two different types of stripes are provided. The first type of stripe is limited to stripes that form a rectangular area / region in the picture as indicated by the thick solid line in picture 101. Picture 101 shows that the picture is partitioned into six different rectangular stripes (0) to (5). The second type of stripe is limited to continuous blocks in the raster scan order as indicated by the thick solid line in picture 102 (so that they form a sequence of blocks). Picture 102 shows that the picture is partitioned into three different stripes (0) to (2) consisting of continuous blocks in raster scan order. Generally, rectangular stripes are structures / arrangements / configurations for dealing with the selection of regions of interest (ROIs) in videos. A slice can be encoded in a bitstream (or decoded from a bitstream) as one or several network abstraction layer (NAL) units. A NAL unit is a logical unit of data used to encapsulate data in an encoded / decoded bitstream (e.g., a packet containing an integer number of bytes, where multiple packets together form the encoded video data). In the encoding / decoding system of VVC7, a slice is typically encoded as a single NAL unit. When a slice is encoded as several NAL units in a bitstream, each NAL unit of a slice is referred to as a slice segment. A slice segment includes a slice segment header containing coding parameters for the slice segment. According to a variation, the header of the first slice segment NAL unit of a slice contains all coding parameters for the slice. The slice segment header of the subsequent NAL unit of the slice may contain fewer parameters than the first NAL unit. In this case, the first slice segment is an independent slice segment, and the subsequent segments are dependent slice segments (because they depend on the coding parameters of the NAL unit from the first slice segment).
[0106] 3.2 Partitioning into sub-images
[0107] Figure 2Sub-picture partitioning of a picture according to an embodiment of the present invention is shown, i.e., partitioning a picture into one or more sub-pictures. A sub-picture represents a portion of a picture (a part or portion of a picture) that covers a rectangular area of the picture. Each sub-picture may have a different size and encoding parameters than another sub-picture. Sub-picture layout, i.e., the geometry of sub-pictures in a picture (e.g., as defined using the position and size / width / height of the sub-pictures), allows grouping of sets of slices of a picture and may constrain (i.e., impose restrictions on) temporal motion prediction between two pictures.
[0108] exist Figure 2 , the block partition of picture 201 is a 4×5 block grid. Strip partitioning defines 24 strips, which include one strip for each block (except for the last column of blocks on the right hand side, where for the last column of blocks, each block is partitioned into two strips (i.e., a strip includes partial blocks)). Picture 201 is also partitioned into two sub-pictures 202 and 203. A sub-picture is defined as one or more strips forming a rectangular area. Sub-picture 202 (shown as a dotted area) includes the strips in the first three block columns (starting from the left), and sub-picture 203 (shown as a shaded area with diagonal lines thereon) includes the remaining strips (in the last two block columns on the right hand side). As shown Figure 2 As shown in, VVC7 and embodiments of the present invention provide for allowing single slice partitions and single block partitions to be defined at the picture level (e.g., for each picture using syntax elements provided in the PPS). Sub-picture partitioning is applied on top of block and slice partitions. Another aspect of sub-pictures is that each sub-picture is associated with a set of flags. This makes it possible to indicate (using one or more of the set of flags) that the temporal prediction is constrained to use data from a reference frame that is part of the same sub-picture (e.g., the temporal prediction is restricted so that the predictor of a sub-picture cannot use reference data from another sub-picture). For example, reference Figure 2 , CTB 204 belongs to sub-picture 202. When it is indicated that temporal prediction is constrained for sub-picture 202, temporal prediction of sub-picture 202 cannot use reference blocks (or reference data) from sub-picture 203. As a result, slices of sub-picture 202 can be encoded / decoded independently of slices of sub-picture 203. This feature / attribute / property / capability is useful in viewport-dependent streaming that includes segmenting an omni-directional video sequence into spatial portions, where each spatial portion represents a specific viewing direction of 360 (degree) content. The viewer can then select the segment corresponding to the desired / relevant viewing direction, and using this property of the sub-picture, the segment can be encoded / decoded without accessing data from the remaining portion of the 360 content.
[0109] Another use of sub-pictures is to generate streams with regions of interest. Sub-pictures provide spatial representations for these regions of interest that can be independently encoded / decoded. Sub-pictures are designed to enable / allow easy access to the coded data corresponding to these regions. As a result, the coded data corresponding to the sub-picture can be extracted and a new bitstream can be generated that includes data for only a single sub-picture or a combination / synthesis of a sub-picture with one or more other sub-pictures, i.e., sub-picture-based bitstream generation can be used to increase flexibility and scalability.
[0110] 3.3 Bit Stream
[0111] Figure 3 The organization (i.e., structure, configuration, or arrangement) of a bitstream according to an embodiment of the present invention that complies with the requirements of a coding system of VVC7 is shown. The bitstream 300 consists of data representing / indicating an ordered sequence of syntax elements and encoded (image) data. The syntax elements and the encoded (image) data are placed (i.e., packaged / packed) into NAL units 301 to 308. There are different NAL unit types. The network abstraction layer (NAL) provides the ability / performance to encapsulate the bitstream into packets for different protocols, such as Real Time Protocol / Internet Protocol (RTP / IP), ISO base media file format, etc. The network abstraction layer also provides a framework for packet loss resilience.
[0112] NAL units are divided into VCL NAL units and non-VCL NAL units, VCL stands for video coding layer. VCL NAL units contain actual coded video data. Non-VCL NAL units contain additional information. The additional information can be parameters required to decode the coded video data, or supplementary data that can enhance the usability of the decoded video data. Figure 3 The NAL units 306 in VCL_VCL_NAL_UNIT_S306 correspond to slices (ie, they include the actual coded video data for the slices) and constitute the VCL NAL units of the bitstream.
[0113] Different NAL units 301 to 305 correspond to different parameter sets, which are non-VCL NAL units. DPS NAL unit 301 represents a decoding parameter set NAL unit, which contains constant parameters for a given decoding process. VPS NAL unit 302 (VPS stands for video parameter set NAL unit) contains parameters defined for the entire video (for example, the entire video includes one or more sequences of pictures / images), so it is applicable when decoding the encoded video data of the entire bit stream. DPS NAL units can define parameters that are more static than those in VPS NAL units (in the sense that the parameters are stable and do not change so much during the decoding process). In other words, the parameters of DPS NAL units change less frequently than those of VPS NAL units. SPS NAL unit 303 (SPS stands for sequence parameter set) contains parameters defined for a video sequence (i.e., a sequence of pictures or images). Specifically, SPS NAL units can define sub-picture layouts and associated parameters of a video sequence. The parameters associated with each sub-picture specify the coding constraints applied to the sub-picture. According to a variant, the parameters include a flag indicating that temporal prediction between sub-pictures is restricted, so that data from the same sub-picture can be used during the temporal prediction process. Another flag may enable or disable loop filters (ie post-filtering) across sub-picture boundaries.
[0114] The PPS NAL unit 304 (PPS stands for Picture Parameter Set) contains the parameters defined for a picture or group of pictures. The APSNAL unit 305 (APS stands for Adaptive Parameter Set) contains the parameters of the loop filter, which is typically an adaptive loop filter (ALF) or a shaper model (or luma mapping with chroma scaling model) or a scaling matrix used at slice level. The bitstream may also contain SEI NAL units ( Figure 3 ), which represents the supplemental enhancement information NAL unit. The periodicity (or frequency included) of these parameter sets (or NAL units) in the bitstream is variable. A VPS defined for the entire bitstream may appear only once in the bitstream. In contrast, an APS defined for a slice may appear once for each slice in each picture. In practice, different slices may rely on (e.g., refer to) the same APS, so in the bitstream of a picture, there are usually fewer APS NAL units than slices.
[0115] The AUD NAL unit 307 is an access unit delimiter NAL unit that separates two access units. An access unit is a set of NAL units that may include one or more coded pictures with the same decoding timestamp (ie, a group of NAL units associated with one or more coded pictures with the same timestamp).
[0116] The PH NAL unit 308 is a picture header NAL unit that groups parameters common to a set of slices of a single coded picture.A picture may refer to one or more APSs to indicate the AFL parameters, shaper models, and scaling matrices used by slices of the picture.
[0117] Each VCL NAL unit 306 contains video / image data of a slice. A slice may correspond to an entire picture or a sub-picture, a single block or multiple blocks or a fragment (partial block) of a block. For example, Figure 3 A slice includes a number of blocks 320. A slice consists of a slice header 310 and a raw byte sequence payload (RBSP) 311, which contains coded pixel / component sample data encoded as coded blocks 340.
[0118] The syntax of the PPS as in VVC7 includes a syntax element specifying the size of a picture in units of luma samples, and also includes a syntax element for specifying partitioning of each picture in units of blocks and slices.
[0119] The PPS contains syntax elements that allow the determination (i.e., enable determination) of the slice positions in a picture / frame. Since a sub-picture forms a rectangular area in a picture / frame, the set of slices, part of a block, or a block belonging to a sub-picture can be determined from parameter set NAL units (i.e., one or more of DPS, VPS, SPS, PPS, and APSNAL units).
[0120] Encoding and decoding process
[0121] 3.4 Encoding Processing
[0122] Figure 4 A coding method for encoding a picture of a video in a bit stream according to an embodiment of the present invention is shown.
[0123] In a first step 401, an image is segmented into sub-pictures. For each sub-picture, the size of the sub-picture is determined according to the spatial access granularity required by the application (e.g., the sub-picture size can be expressed according to the size / scale / granularity level of the region / spatial portion / area in the picture required by the application / usage scenario, and the sub-picture size can be small enough to contain a single region / spatial portion / area). Typically, for a viewport-dependent streaming method, the sub-picture size is set to cover a predetermined range of the field of view (e.g., a range of a 60° horizontal field of view). For adaptive streaming of regions of interest, the width and height of each sub-picture are made dependent on the region of interest present in the input video sequence. Typically, the size of each sub-picture is made to contain a region of interest. The size of the sub-picture is determined in units of luma samples or in multiples of the size of a CTB. In addition, the position of each sub-picture in the coded picture is determined in step 401. The position and size of the sub-picture form sub-picture layout information, which is typically signaled in a non-VCL NAL unit (such as a parameter set NAL unit, etc.). For example, in step 402, the sub-picture layout information is encoded in an SPS NAL unit.
[0124] The SPS syntax for such an SPS NAL unit typically contains the following syntax elements:
[0125]
[0126] The descriptor column gives the encoding method used to encode the syntax element, for example, u(n) (where n is an integer value) means that the syntax element is encoded using n bits, and ue(v) means that the syntax element is encoded using unsigned integer 0-order exponential Golomb coding (where the left-leading bit is a variable length encoding). Except for signed integers, se(v) is equivalent to ue(v). u(v) means that the syntax element is encoded using a fixed-length encoding (with a specific length in bits determined from other parameters).
[0127] The presence of sub-pictures signaled in the bitstream of a picture depends on the value of the flag subpics_present_flag. When the flag is equal to 0, it indicates that the bitstream does not contain information related to the partitioning of the picture into one or more sub-pictures. In this case, it is inferred that there is a single sub-picture covering the entire picture. When the flag is equal to 1, a set of syntax elements specifies the layout of sub-pictures in a frame (i.e., picture): signaling includes the use of a loop (i.e., a programming structure for repeating a sequence of instructions until a certain condition is met) to determine / define / specify the sub-pictures of the picture (the number of sub-pictures in a picture is encoded with the sps_num_subpics_minus1 syntax element), which includes defining the position and size of each sub-picture. The index of this "for loop" is the sub-picture index. The syntax elements subpic_ctu_top_left_x[i] and subpic_ctu_top_left_y[i] correspond to the column index and row index of the first CTU of the i-th sub-picture, respectively. The subpic_width_minus1[i] and subpic_height_minus1[i] syntax elements signal the width and height of the i-th sub-picture in units of CTU.
[0128] In addition to the sub-picture layout, the SPS also specifies constraints on sub-picture boundaries: for example, subpic_treated_as_pic_flag[i] equal to 1 indicates that the boundary of the i-th sub-picture is treated as a picture boundary for temporal prediction. This ensures that the coding block of the i-th sub-picture is predicted from the data of the reference picture belonging to the same sub-picture. When equal to 0, this flag indicates that temporal prediction may or may not be constrained. The second flag (loop_filter_across_subpic_enabled_flag[i]) specifies whether the loop filtering process is allowed to use data (usually pixel values) from another sub-picture. These two flags make it possible to indicate whether a sub-picture is encoded independently of other sub-pictures. This information is useful in determining whether a sub-picture can be extracted, derived, or merged with other sub-pictures.
[0129] In step 403, the encoder determines the partitions in the blocks and slices of the video sequence and describes these partitions in a non-VCL NAL unit such as PPS. Figure 6 This step is further described.The signaling of slice and block partitioning is subject to sub-picture constraints, such that each sub-picture includes at least one slice, and a portion of a block (ie, a partial block) or one or more than one block.
[0130] In step 404, at least one slice forming a sub-picture is encoded in a bitstream.
[0131] 3.5 Decoding Process
[0132] Figure 5 A general decoding process for a slice according to an embodiment of the present invention is shown. For each VCL NAL unit, the decoder determines the PPS and SPS that apply to the current slice. Typically, identifiers for the PPS and SPS for the current picture are determined. For example, the picture header of the slice signals the identifier of the PPS in use. The PPS associated with this PPS identifier then also references an SPS that uses another identifier (the identifier of the SPS).
[0133] In step 501, the decoder determines the sub-picture partitioning, for example by parsing a parameter set describing / indicating the sub-picture layout to determine the size of the sub-picture of the picture / frame, typically its width and height. For VVC7 and embodiments compliant with this part of VVC7, the parameter set including information for determining the sub-picture partitioning is an SPS. In a second step 502, the decoder parses the syntax elements of a parameter set NAL unit (or non-VCL NAL unit) related to the partitioning of the picture into blocks. For example, for a VVC7 compliant stream, the signaling of the block partitioning is in the PPS NAL unit. During this determination step, the decoder initializes a set of variables describing / defining the characteristics of the blocks present in each sub-picture. For example, the following information for the i-th sub-picture can be determined (see Figure 6 Step 601):
[0134] A flag indicating whether the sub-picture contains a fragment of a block (i.e., a partial block) (see Figure 6 Step 603)
[0135] An integer value indicating the number of blocks in the sub-picture ( Figure 6 Step 602)
[0136] An integer value that specifies the width of the sub-image in blocks ( Figure 6 Step 604)
[0137] An integer value that specifies the height of the sub-image in blocks ( Figure 6 Step 604)
[0138] A list of block indices that exist in the sub-picture in raster scan order ( Figure 6 Step 605)
[0139] Figure 6 Signaling of stripe partitioning according to an embodiment of the present invention is shown, which involves determination steps that can be used in both encoding and decoding processes.
[0140] In step 503, the decoder relies on the signaling of the slice partition (in a non-VCL NAL unit, for example, typically in the PPS of VVC7) and previously determined information to infer (i.e., derive or determine) the slice partition of each sub-picture. In particular, the decoder can infer (i.e., derive or determine) the number of slices, the width and height of one or more slices. The decoder can also obtain information present in the slice header to determine the decoding position of the CTB present in the slice data.
[0141] In a final step 504 , the decoder decodes the slices forming the sub-picture at the positions determined in step 503 of the picture.
[0142] Signaling a partition
[0143] 3.6 Signaling Stripe Partitions
[0144] According to an embodiment of the present invention, a slice may be composed of an integer number of complete blocks of a picture or an integer number of complete and continuous CTU rows or columns within a block (the latter possibility means that the slice can include partial blocks as long as the partial blocks include continuous CTU rows or columns).
[0145] Two stripe modes may also be supported / provided for use, namely raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a sequence of complete blocks of a picture in a block raster scan order. In rectangular stripe mode, a stripe contains multiple complete blocks that together form a rectangular area of a picture, or multiple consecutive complete CTU rows (or columns) that together form a block of a rectangular area of a picture. Blocks within the stripe are scanned in a block raster scan order within a rectangular area corresponding to the rectangular stripe.
[0146] The syntax of VVC7 for specifying a slice structure (layout and / or partitioning) is independent of the syntax of VVC7 for sub-pictures. For example, slice partitioning (i.e., partitioning a picture into slices) is performed on top of (i.e., based on or with reference to) a block structure / partitioning without reference to a sub-picture (i.e., without reference to syntax elements used to form a sub-picture). On the other hand, VVC7 imposes some restrictions (i.e., constraints) on sub-pictures, for example, a sub-picture must contain one or more than one slice, and a slice header includes a slice_address syntax element, which is an index of the slice relative to its sub-picture (i.e., an index defined for the associated sub-picture, such as an index of a strip within a strip in a sub-picture). VVC7 also only allows strips that include partial blocks in a rectangular strip mode, and a raster scan strip mode does not define such strips that include partial blocks. The current syntax system used in VVC7 is not designed to enforce all of these constraints, so implementations of the syntax system result in a system that is prone to generating bitstreams that do not conform to the VVC7 specifications / requirements.
[0147] Therefore, embodiments of the present invention use information from sub-picture layout definitions to specify slice partitions in an attempt to provide better coding efficiency for the signaling of sub-pictures, blocks, and slices.
[0148] The requirement to have the ability to define multiple slices within a tile (also called "tile fragment" slices or slices that include part of a tile) comes from omni-directional streaming requirements. It has been identified that a stripe needs to be defined within a tile for BEAMER (Bitstream Extraction And MERging) operations for OMAF streams. This means that "tile fragment" slices can exist in different sub-pictures to allow BEAMER operations, which then means that it makes little sense to have a sub-picture containing one complete tile with multiple slices.
[0149] The following first three embodiments of the present invention (Embodiment 1, Embodiment 2 and Embodiment 3) define the determination and signaling of slice partitioning based on sub-pictures. The first embodiment 1 includes a syntax system that does not allow / prevent / prohibit a sub-picture from containing multiple "block fragment" slices, while the second embodiment 2 includes a syntax system that allows / permits only when the sub-picture contains at most one block (i.e., if the sub-picture contains more than one block, all of its slices contain an integer number of complete blocks). Embodiment 3 provides explicit signaling of the use of "block fragment" slices. When "block fragment" slices are not used, signaling syntax elements related to such slices is avoided, which can improve the coding efficiency of signaling slice partitioning. Embodiments 1 to 3 allow / permit the use of block fragment slices only in rectangular slice mode.
[0150] A fourth embodiment 4 is an alternative to the first three embodiments, in which a unified syntax system is used to allow / permit the use of tile fragment stripes in both raster scan and rectangular strip modes. A fifth embodiment 5 is an alternative to the other embodiments, in which the sub-picture layout and slice partitioning are not signaled in the bitstream, but are inferred (i.e., determined or derived) from the signaling of the tiles.
[0151] Example 1
[0152] In the first embodiment (Embodiment 1), the syntax of the stripe partition of VVC7 is modified to avoid specifying many constraints that are difficult to implement correctly, and the parameters for the stripe partition, such as the size of the stripe, are also inferred / derived depending on the sub-picture layout. Since the sub-picture is represented by a set / group of complete stripes (i.e., composed of a set / group of complete stripes), the size of the stripe can be inferred when the sub-picture contains a single stripe. Similarly, based on the number of stripes in the sub-picture and the size of the stripes processed / encountered earlier in the sub-picture, the size of the last stripe can be inferred / derived / determined.
[0153] In Embodiment 1, a slice is allowed to include a fragment / part of a block (i.e., a partial block) only when the slice covers the entire / whole sub-picture (in other words, a single slice exists in the sub-picture). When more than one slice exists in a sub-picture, it is necessary to signal the slice size. On the other hand, when a single slice exists in a sub-picture, the slice size is the same as the sub-picture size. Therefore, when a slice includes a fragment of a block, the slice size is not signaled in the parameter set NAL unit (because it is the same as the sub-picture size). Therefore, the slice width and height can be constrained to be in units of blocks only for the scenario where the slice size is signaled.
[0154] The syntax of the PPS of this embodiment includes specifying the number of slices included in each sub-picture. When the number of slices is greater than 1, the size (width and height) of the slice is expressed in units of blocks. As described above, the size of the final slice is not signaled and is inferred / derived / determined from the sub-picture layout.
[0155] According to a variation of this embodiment, the PPS syntax contains the following syntax elements with the following semantics (ie, definitions or functions):
[0156] PPS Syntax
[0157]
[0158] PPS semantics
[0159] Slices are defined for each sub-picture. A "for loop" is used with the num_slices_in_subpic_minus1 syntax element to form / process the correct number of slices in that particular sub-picture. The syntax element num_slices_in_subpic_minus1[i] indicates the number of slices (in the sub-picture with sub-picture index equal to i) minus 1, that is, the syntax element indicates a value one less than the number of slices in the sub-picture. When equal to 0, it indicates that the sub-picture contains a single slice of size equal to the sub-picture size. When the number of slices is greater than 1, the size of the slice is expressed in units of an integer number of blocks. The size of the final slice is inferred from the sub-picture size (and the sizes of other blocks in the sub-picture). Using this approach, a "block fragment" slice can be defined when it covers a complete sub-picture through the following semantics of the syntax element:
[0160] pps_num_subpics_minus1 plus 1 specifies the number of sub-pictures in the coded picture that references the PPS. A requirement for bitstream conformance is that the value of pps_num_subpic_minus1 shall be equal to sps_num_subpics_minus1 (which is the number of sub-pictures defined at the SPS level).
[0161] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may consist of one or more than one rectangular slice. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1 (which is the number of sub-pictures defined at the SPS level).
[0162] num_slices_in_subpic_minus1[i] plus 1 specifies the number of rectangular slices in the i-th sub-picture. The value of num_slices_in_subpic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1 (inclusive), where MaxSlicesPerPicture is specified in Appendix A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_subpic_minus1[0] is inferred to be equal to 0.
[0163] This syntax element determines the length of slice_address of the slice header, which is Ceil(log2(num_slices_in_subpic_minus1[SubPicIdx]+1)) bits (where SubPicIdx is the index of the sub-picture of the slice). The value of slice_address can be in the range of 0 to num_slices_in_subpic_minus1[SubPicIdx] (inclusive).
[0164] tile_idx_delta_present_flag equal to 0 specifies that tile_idx_delta values are not present in the PPS and all rectangular slices in all sub-pictures of the pictures referencing the PPS are specified in raster order. tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values may be present in the PPS and all rectangular slices in all sub-pictures of the pictures referencing the PPS are specified in the order indicated by the values of tile_idx_delta.
[0165] slice_width_in_tiles_minus1[i][j] plus 1 specifies the width of the j-th rectangular strip in tile columns in the ith sub-picture. The value of slice_width_in_tiles_minus1[i][j] shall be in the range of 0 to NumTileColumns-1, inclusive (where NumTileColumns is the number of tile columns in the tile grid). When not present, the value of slice_width_in_tiles_minus1[i][j] is inferred from the sub-picture size.
[0166] slice_height_in_tiles_minus1[i][j] plus 1 specifies the height of the j-th rectangular strip in tile rows in the i-th sub-picture. The value of slice_height_in_tiles_minus1[i][j] shall be in the range of 0 to NumTileRows-1, inclusive (where NumTileRows is the number of tile rows in the tile grid). When not present, the value of slice_height_in_tiles_minus1[i][j] is inferred from the size of the i-th sub-picture.
[0167] tile_idx_delta[i][j] specifies the difference in tile index between the j-th rectangular slice and the (j+1)-th rectangular slice of the i-th sub-picture. The value of tile_idx_delta[i][j] shall be in the range of –NumTilesInPic[i]+1 to NumTilesInPic[i]-1 (inclusive), where NumTilesInPic[i] is the number of tiles in the picture. When not present, the value of tile_idx_delta[i][j] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i][j] shall not be equal to 0.
[0168] Thus, according to this variation:
[0169] A sub-picture is a rectangular area of one or more stripes within a picture. The address of a stripe is defined relative to a sub-picture. This association / relationship between a sub-picture and its stripes is reflected in the syntax system, which defines the stripes in a "for loop" applied to each sub-picture.
[0170] • By design, unwanted partitioning that could result in having more than one block fragment slice from two different blocks in the same sub-picture is avoided.
[0171] • Slice partitioning can be derived from both block and sub-picture partitioning, which improves the signaling coding efficiency.
[0172] According to another variation of this variation, the following process is used to perform such derivation / derivation of slice partitions. For rectangular slices, when single_slice_per_subpic_flag is equal to 0, the list NumCtuInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) (specifying the number of CTUs in the i-th slice) and the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtuInSlice[i]-1 (inclusive)) (specifying the picture raster scan address of the j-th CTB in the i-th slice) are derived as follows:
[0173]
[0174]
[0175] Wherein the function AddCtbsToSlice(sliceIdx, startX, stopX, startY, stopY) fills the CtbAddrInSlice array of the strip with an index equal to SliceIdx. The array is filled with CTB addresses for CTBs in raster scan order, where the vertical address of a CTB row is between startY and stopY; and the horizontal address of a CTB column is between startX and stopX.
[0176] The processing includes applying a processing loop to each sub-picture. For each sub-picture, the tile index of the first tile in the sub-picture is determined based on the horizontal and vertical addresses of the first tile of the sub-picture (i.e., the subpicTileTopLeftX[i] and SubpicTileTopLeftX[i] variables) and the number of tile columns specified by the tile partition information. This value infers / indicates / represents the index of the first tile in the first slice of the sub-picture. For each sub-picture, a second processing loop is applied to each slice of the sub-picture. The number of slices is equal to one plus a variable num_slices_in_subpics_minus1[i][j], which is encoded in the PPS or inferred / derived / determined from other information included in the bitstream. When the slice is the last of the sub-picture, the width of the slice in units of tiles is inferred / derived / derived / determined to be equal to the width of the sub-picture in units of tiles minus the horizontal address of the column of the first tile of the slice plus the horizontal address of the column of the first tile of the sub-picture. Similarly, the height of a slice in a block is inferred / derived / derived / determined to be equal to the height of the sub-picture in units of blocks minus the vertical address of the row of the first block of the slice plus the vertical address of the row of the first block of the sub-picture. The index of the first block of the previous slice is encoded in the block partition information (e.g., as a difference in the block index of the first block of the previous slice), or it is inferred / derived / derived / determined to be equal to the next block in raster scan order of the blocks in the sub-picture.
[0177] When a sub-picture contains a fragment of a tile (i.e., a partial tile), the width and height of the slice in CTU units are inferred / derived / derived / determined to be equal to the width and height in CTU units of the sub-picture. The CtbAddrInSlice[sliceIdx] array is filled with the CTUs of the sub-picture in raster scan order. Otherwise, the sub-picture contains one or more tiles, and the CtbAddrInSlice[sliceIdx] array is filled with the CTUs of the tiles included in the slice. The tiles included in the stripe of the sub-picture are tiles having a vertical address of a tile column and a horizontal address of the tile column, the vertical address is defined from the range of [tileX, tileX+slice_width_in_tiles_minus1[i][j]], and the horizontal address is defined from the range of [tileY, tileY+slice_height_in_tiles_minus1[i][j]], where tileX is the vertical address of the tile column of the first tile of the stripe; tiley is the horizontal address of the tile row of the first tile of the stripe; i is the sub-picture index, and j is the index of the stripe in the sub-picture.
[0178] Finally, the processing loop for the slice includes a determination step for determining the first tile in the next slice of the sub-picture. When the tile index offset (tile_idx_delat[i][j]) is encoded (tile_idx_delta_present_flag[i] is equal to 1), the tile index of the next slice is set to a value equal to the index of the first tile of the current slice plus the tile index offset value. Otherwise (i.e., when the tile index offset is not encoded), tileIdx is set to a value equal to the tile index of the first tile of the sub-picture plus the product of the number of tile columns in the picture and the height of the current slice in tiles minus 1.
[0179] According to another variant, when a sub-picture includes a partial block, instead of signaling the number of slices included in the sub-picture in the PPS, it is inferred / derived / derived / determined that the sub-picture consists of the partial block. For example, the following alternative PPS syntax may be used to do this.
[0180] Alternative PPS syntax
[0181] In another variant, when a sub-picture represents a fragment of a block (i.e., the sub-picture includes a partial block), the number of slices in the sub-picture is inferred to be equal to one. As the number of sub-pictures representing fragments of blocks (i.e., the number of sub-pictures including partial blocks) increases, the coded data size of the syntax element is further reduced, which improves the compression of the stream. The syntax of the PPS is, for example, as follows:
[0182]
[0183]
[0184] The new semantics / definition of num_slices_in_subpic_minus1[i] is as follows:
[0185] num_slices_in_subpic_minus1[i] plus 1 specifies the number of rectangular slices in the i-th sub-picture. The value of num_slices_in_subpic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1 (inclusive), where MaxSlicesPerPicture is specified in Appendix A. When not present, the value of num_slices_in_subpic_minus1[i] is inferred to be equal to 0 (i in the range of 0 to pps_num_subpic_minus1 (inclusive).
[0186] The tileFractionSubPicture[i] variable specifies whether the i-th sub-picture covers a fragment (partial block) of a tile, ie the size of the sub-picture is strictly lower (ie smaller) than the tile to which the first CTU of the sub-picture belongs.
[0187] The decoder determines this variable based on the sub-picture layout and tile grid information as follows: When both the top and bottom horizontal boundaries of the sub-picture are tile boundaries, tileFractionSubpicture[i] is set equal to 0. In contrast, when at least one of the top or bottom horizontal sub-picture boundaries is not a tile boundary, tileFractionSubpicture[i] is set equal to 1.
[0188] In yet another variation, the presence or absence of slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] is inferred from the width and height of the sub-picture in units of tiles. For example, the variables subPictureWidthInTiles[i] and subPictureHeightInTiles[i] define the width and height of the i-th sub-picture in units of tiles, respectively. When the sub-picture is a fragment of a tile, the width of the sub-picture is set to 1 (because a fragment of a tile has a width equal to the width of the tile), and the height is set to 0 by convention to indicate that the height of the sub-picture is less than a full tile. It should be understood that any other preset / predetermined values may be used, the main constraint being that the two values are set / determined so that they do not represent possible sub-picture sizes in units of tiles. For example, the values may be set equal to the maximum number of tiles in a picture plus 1. In this case, it may be inferred that the sub-picture is a fragment of a tile, since a width or height of a sub-picture greater than the maximum number of tiles is not possible.
[0189] These variables are initialized once the tile partitioning is determined (typically based on the num_exp_tile_columns_minus1, num_exp_tile_rows_minus1, tile_column_width_minus1[i], and tile_row_height_minus1[i] syntax elements). The processing may include a processing loop for each sub-picture. That is, for each sub-picture, if the sub-picture covers a fragment of a tile, the width and height of the sub-picture in tiles are set to 1 and 0, respectively. Otherwise, the determination of the width of the sub-picture in tiles is determined as follows: The horizontal address of the CTU column of the first CTU of the sub-picture (determined from the sub-picture layout syntax element subpic_ctu_top_left_x[i] of the i-th sub-picture) is used to determine the horizontal address of the tile column of the first tile in the sub-picture.
[0190] Then, for each block column of the block grid, the horizontal addresses of the leftmost and rightmost CTU columns of the block column are determined. When the horizontal address of the CTU column of the first CTU of the sub-picture is between these two addresses, the horizontal address indicates the horizontal address of the block column of the first block in the sub-picture. The same process applies to determining the horizontal address of the block column containing the rightmost CTU column of the sub-picture. The horizontal address of the CTU column of the rightmost CTU column is equal to the sum of the horizontal address of the CTU column of the first CTU of the sub-picture and the width of the sub-picture in CTU units (subpic_width_minus1[i] plus 1). The width of the sub-picture in blocks is equal to the difference between the horizontal address of the block column of the last CTU column and the horizontal address of the first CTU of the sub-picture. The same principle applies when determining the height of the sub-picture in blocks. The process determines the height of the sub-picture in blocks as the difference between the vertical addresses of the block rows of the first CTU of the sub-picture and the last CTU row of the sub-picture.
[0191] In another variation, when the number of tiles in the sub-picture is equal to the number of slices in the sub-picture, slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] do not exist and are inferred to be equal to 0. In fact, the sub-picture and slice constraints enforce that each slice contains exactly one tile, when equal. When the sub-picture is a tile fragment (tileFractionSubpicture[i] is equal to 1), the number of tiles in the sub-picture is equal to 1. Otherwise, it is equal to the product of subPictureHeightInTiles[i] and SubPictureWidthInTiles[i].
[0192] In another variation, when the width of the sub-picture in tiles is equal to 1, slice_width_in_tiles_minus1[i][j] does not exist and is inferred to be equal to 0.
[0193] In another variation, when the height of the sub-picture in tiles is equal to 1, slice_height_in_tiles_minus1[i][j] does not exist and is inferred to be equal to 0.
[0194] In yet another variation, any combination of three of the foregoing another variations is used.
[0195] In some previous variants, the presence of the syntax element is inferred from the sub-picture partition information. When sub-picture partitions are defined in different parameter set NAL units, the parsing of slice partitions depends on information from other parameter set NAL units. This dependency may limit the use of the variant in certain applications, since the parsing of the parameter set including the slice partition cannot be done without storing information from the other parameter set. For the decoding of the parameter, i.e. determining the value encoded by the syntax element, this dependency is not a limitation, since the decoder needs to use all parameter sets to decode the pixel samples in any way (there may be some delay caused by having to wait for all relevant parameter sets to be decoded). Therefore, in another variant, the inference of the presence of the syntax element is enabled only when sub-picture, block and slice partitions are signaled in the same parameter set NAL unit. For example, a variable subpictureWidthInTiles[i] specifying the width of the i-th sub-picture in units of tiles, a variable subpictureHeightInTiles[i] specifying the height of the i-th sub-picture in units of tiles, a variable subpictileTopLeftX[i] specifying the horizontal address of the column of the first tile in the i-th sub-picture, and subpicTileTopLefty[i] specifying the vertical address of the row of the first tile in the i-th sub-picture (where i is in the range of 0 to pps_num_subpicture_minus1 (inclusive)) are determined as follows:
[0196]
[0197] The tileFractionSubpicture[i] variable that specifies whether a sub-picture contains a tile's fragment is exported as follows:
[0198]
[0199] The list SliceSubpicToPicIdx[i][k] specifying the number of rectangular slices in the ith sub-picture and the picture-level slice index of the kth slice of the ith sub-picture is derived as follows:
[0200]
[0201] in:
[0202] CtbToTileRowBd[ctbAddrY] converts the vertical CTB address (ctbAddrY) to the upper tile row boundary in units of CTB.
[0203] CtbTotileColBd[ctbAddrX] converts the horizontal CTB address (ctbAddrX) to the left tile column boundary in units of CTB.
[0204] ColWidth[i] is the width of the i-th block column in CTB units.
[0205] RowHeight[i] is the height of the i-th block row in CTB units
[0206] tileColBd[i] is the position of the i-th tile column boundary in CTB units.
[0207] tileRowBd[i] is the position of the i-th tile row boundary in CTB units.
[0208] NumTileColumns is the number of tile columns
[0209] NumTileRows is the number of tile rows
[0210] Figure 7 An example of sub-picture and slice partitioning using the above embodiment / variant / further variant is shown. In this example, the picture 700 is partitioned into 9 sub-pictures labeled (1) to (9) and a 4×5 block grid (block boundaries are shown in thick solid lines). For each sub-picture, the slice partitioning (the area included in each slice is shown in thin solid lines, which is just within the slice boundary) is as follows:
[0211] Sub-picture (1): 3 stripes, each containing 1 block, 2 blocks, and 3 blocks. The stripe height is equal to 1 block, and the stripe width is 1, 2, and 3 in units of blocks (i.e., 3 stripes consisting of a row of blocks arranged in the horizontal direction).
[0212] Sub-image (2): 2 equal-sized stripes, 1 block wide and 1 block high (i.e., 2 stripes each consisting of a single block)
[0213] Sub-pictures (3) to (6): 1 "block fragment" slice, i.e. a slice consisting of a single partial block
[0214] Sub-picture (7): 2 stripes of column size of 2 blocks (i.e., 2 stripes each consisting of columns of 2 blocks arranged in the vertical direction)
[0215] Sub-image (8): 1 strip of 3 blocks
[0216] Sub-image (9): 2 stripes with a row size of 1 block and a row size of 2 blocks
[0217] For sub-picture (1), the width and height of the first two slices are encoded, while the size of the last slice is inferred.
[0218] For sub-picture (2), the width and height of the first two slices are inferred since there are two slices for two blocks in the sub-picture.
[0219] For sub-pictures (3) to (6), the number of slices in each sub-picture is equal to 1, and the width and height of the slices are inferred to be equal to the sub-picture size, because each sub-picture is a fragment of a block.
[0220] For sub-pictures (7), the width and height of the first slice and the width and height of the last slice are inferred from the sub-picture size.
[0221] For a sub-picture (8), the width and height of the stripe are inferred to be equal to the sub-picture size, since there is a single stripe in the sub-picture.
[0222] For a sub-picture (9), the height of the slice is inferred to be equal to 1 (because the sub-picture height in blocks is equal to 1), and the width of the first slice is encoded, while the width of the last slice is equal to the width of the sub-picture minus the width of the first slice.
[0223] Example 2
[0224] In a second embodiment (Embodiment 2), the constraint / restriction that a "block fragment" slice (i.e., a slice that is a part of a block) should cover the entire sub-picture is relaxed / removed. As a result, a sub-picture may contain one or more slices, each of which includes one or more blocks, but may also contain one or more "block fragment" slices.
[0225] In this embodiment, sub-picture partitioning allows / enables prediction / derivation / determination of slice positions and sizes.
[0226] According to a variation of this embodiment, the following PPS syntax may be used to perform this operation.
[0227] PPS Syntax
[0228] For example, the PPS syntax is as follows:
[0229]
[0230]
[0231] The semantics of the syntax elements single_slice_per_subpic_flag, tile_idx_delta_present_flag, num_slices_in_subpic_minus1[i] and tile_idx_delta[i][j] are the same as in the previous embodiment.
[0232] The slice_width_minus1 and slice_height_minus1 syntax elements (parameters) specify the slice size in units of blocks or CTUs, depending on sub-picture partitioning.
[0233] When the last CTU of the slice is the last CTU of the block, the variable newTileIdxDeltaRequired is set equal to one. When the slice is not a "block fragment" slice, newTileIdxDeltaRequired is equal to 1. When the slice is a "block fragment" slice, if the slice is not the last in the block in the slice, newTileIdxDeltaRequired is set to 0. Otherwise, it is the last in the block and newTileIdxDeltaRequired is set to 1.
[0234] In a first further variant, a sub-picture is constrained / restricted to contain a slice of fragment blocks of a single block. In this case, if the sub-picture contains more than one block, the size is in units of blocks. Otherwise, the sub-picture contains a single block or a portion of a block (partial block), and the slice height is defined in units of CTUs. The width of the slice must be equal to the sub-picture width, so it can be inferred and does not need to be encoded in the PPS.
[0235] slice_width_minus1[i][j] plus 1 specifies the width of the j-th rectangular strip. The value of slice_width_in_tiles_minus1[i][j] should be in the range of 0 to NumTileColumns-1 (inclusive) (where NumTileColumns is the number of tile columns in the tile grid). When not present (i.e., when SubPictureWidthInTiles[i]*SubPictureHeightInTiles[i] == 1 or SubPictureWidthinTiles[i] is equal to 1), the value of slice_width_in_tiles_minus1[i][j] is inferred to be equal to 0.
[0236] slice_height_in_tiles_minus1[i][j] plus 1 specifies the height of the j-th rectangular slice in the i-th sub-picture. The value of slice_height_in_tiles_minus1[i][j] shall be in the range of 0 to NumTileRows-1 (inclusive) (where NumTileRows is the number of tile rows in the tile grid). When not present (i.e., when subPictureHeightInTiles[i] is equal to 1 and num_slices_in_subpic_minus1[i] == 0), the value of slice_height_in_tiles_minus1[i][j] is inferred to be equal to 0.
[0237] The variable SliceWidthInTiles[i][j] that specifies the width of the j-th rectangular strip of the i-th sub-picture in units of blocks, the variable SliceHeightInTiles[i][j] that specifies the height of the j-th rectangular strip of the i-th sub-picture in units of blocks, and the variable SliceHeightInCTUs[i][j] that specifies the height of the j-th rectangular strip of the i-th sub-picture in units of CTB are derived as follows (i is in the range of 0 to pps_num_subpic_minus1, and j is in the range of 0 to num_slices_in_subpic_minus1[i]).
[0238]
[0239] Using this algorithm, SliceHeightInCTUs[i][j] is only valid when sliceHeightInTiles[i][j] is equal to 0.
[0240] In an alternative further variant, a sub-picture is allowed to contain block segment slices from several blocks (i.e., slices including partial blocks from more than one block), with the limitation / restriction / condition that all slices of the sub-picture must be block segment slices. Thus, a sub-picture may not be defined that contains a first block segment slice including partial blocks of a first block and a second block segment slice including partial blocks of a second block different from the first block and another slice that covers a third (different) block in its entirety (i.e., the third block is a complete / whole block). In this case, if the sub-picture contains more slices than blocks in the sub-picture, the slice_height_minus1[i][j] syntax element is in units of CTUs, and slice_width_minus1[i][j] is inferred to be equal to the sub-picture width in units of CTUs. Otherwise, the size of the slice is in units of blocks.
[0241] In another variation, the slice PPS syntax is as follows:
[0242]
[0243]
[0244] In this other variant, the same principle applies except that separate syntax elements define the slice width and height in tiles or the slice height in CTUs depending on whether the slice is defined in CTUs. When the slice is signaled in tiles, the slice width and height are defined by slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j], while when the width is expressed in CTUs, slice_width_in_ctu_minus1[i][j] is used. The variable sliceInCtuFlag[i] equal to 1 indicates that the i-th sub-picture contains only tile fragment slices (i.e., there are no whole / complete tile slices in the sub-picture). When equal to 0, it indicates that the i-th sub-picture contains a slice including one or more tiles.
[0245] For i in the range 0 to pps_num_subpic_minus1, the variable sliceInCtuFlag[i] is derived as follows:
[0246]
[0247] The determination of the sliceInCtuFlag[i] variable introduces a parsing dependency between slice and sub-picture partition information. Therefore, in a variant, when slice, block, and sub-picture partitions are signaled in different parameter set NAL units, sliceInCtuFlag[i] is signaled and not inferred.
[0248] In another variant, a sub-picture is allowed to contain block segment slices from several blocks without any specific constraints / restrictions. Therefore, a sub-picture containing a first block segment slice (which includes a partial block of the first block) and a second block segment slice (which includes a partial block of a second block different from the first block) and another slice (which covers the third (different) block as a whole (i.e., the third block is a complete / entire block)) can be defined. In this case, when the number of blocks in the sub-picture is greater than 1, the flag indicates whether the slice size is specified in CTU or block units. For example, the following syntax of the PPS signals the slice_in_ctu_flag[i] syntax element, which indicates whether the slice size of the i-th sub-picture is expressed in CTU or block units. slice_in_ctu_flag[i] equal to 0 indicates the presence of slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] syntax elements and the absence of slice_height_in_ctu_minus1[i][j], i.e. the slice size is expressed in block units. slice_in_ctu_flag[i] equal to 1 indicates the presence of slice_height_in_ctu_minus1[i][j] and the absence of slice_width_in_tiles_minus1[i][j] and slice_height_in_tiles_minus1[i][j] syntax elements, i.e. the slice size is expressed in CTU units.
[0249]
[0250]
[0251] Figure 8 An example of signaled sub-picture and slice partitioning using the embodiments / variants / described further above is shown. In this example, the picture 800 is partitioned into 6 sub-pictures labeled (1) to (6) and a 4×5 grid of blocks (block boundaries are shown with thick solid lines). For each sub-picture, the slice partitioning (the regions included in each slice are shown with thin solid lines just within the slice boundaries) is as follows:
[0252] Sub-picture (1): 3 stripes of row size 1 block, 2 blocks, and 3 blocks (i.e., 3 stripes consisting of a row of blocks arranged in the horizontal direction)
[0253] Sub-picture (2): 2 stripes of equal size, the size of 1 block (i.e., 2 stripes each consisting of a single block)
[0254] Sub-picture (3): 4 “block fragment” strips, i.e. 4 strips each consisting of a single partial block
[0255] Sub-image (4): 2 stripes of column size 2 blocks (i.e., 2 stripes each consisting of a column of 2 blocks arranged in the vertical direction)
[0256] Sub-image (5): 1 strip of 3 blocks
[0257] Sub-image (6): 2 stripes with a row size of 1 block and a row size of 2 blocks
[0258] For sub-picture (3), the number of slices in the sub-picture is equal to 4, and the sub-picture contains only two blocks. The width of the inferred slice is equal to the sub-picture width, and the height of the slice is specified in CTU units.
[0259] For sub-pictures (1), (2), (4) and (5), the number of slices is lower than the blocks in the sub-picture, so the width and height are specified in units of blocks when necessary (i.e., when they cannot be inferred / derived / determined from other information).
[0260] For sub-picture (1), the width and height of the first two slices are encoded, while the size of the last slice is inferred.
[0261] For sub-image (2), the width and height of the first two stripes are inferred because there are 2 stripes of 2 blocks in the sub-image.
[0262] For sub-picture (4), the width and height of the first slice and the size of the last slice are inferred from the sub-picture size.
[0263] For sub-picture (5), the width and height are inferred to be equal to the sub-picture size because there is a single stripe in the sub-picture.
[0264] For sub-picture (5), the slice height is inferred to be equal to 1 (because the sub-picture height in blocks is equal to 1), the width of the first slice is encoded, and the width of the last slice is inferred to be equal to the width of the sub-picture minus the size of the first slice.
[0265] Example 3
[0266] In this third embodiment (Embodiment 3), in the bitstream it is specified whether tile segment slices are enabled or disabled. The principle is to include in one of the parameter set NAL units (or non-VCL NAL units) a syntax element indicating whether the use of "tile segment" slices is allowed.
[0267] According to a variant, the SPS includes a flag indicating whether "block fragment" slices are allowed, so when the flag indicates that "block fragment" slices are not allowed, the signaling of "block fragment" can be skipped. For example, when the flag is equal to 0, it indicates that "block fragment" slices are not allowed. When the flag is equal to 1, "block fragment" slices are allowed. The NAL unit may include a syntax element to indicate the location of the "block fragment" slice. For example, the following SPS syntax elements and their semantics can be used to perform this operation.
[0268] SPS syntax: used to enable / disable block segment striping
[0269]
[0270]
[0271] SPS semantics
[0272] sps_tile_fraction_slices_enabled_flag specifies whether "tile slice" slices are enabled in the coded video sequence. sps_tile_fraction_slices_enabled_flag equal to 0 indicates that the slice should contain an integer number of blocks. sps_tile_fraction_slices_enabled_flag equal to 1 indicates that the slice may contain an integer number of blocks or an integer number of CTU rows from a block.
[0273] In another variation, the sps_tile_fraction_slices_enabled_flag is specified at the PPS level to provide more granularity for adaptively applying / defining the presence of "tile fragment" slices. In yet another variation, the flag may be located in the picture header NAL unit to allow / enable the presence of "tile fragment" slices to be adapted on a picture basis. The flag may be present in multiple NAL units with different values to allow overriding of configurations defined in higher level parameter sets. For example, the value of the flag in the picture header overrides the value in the PPS, which overrides the value in the SPS.
[0274] In alternative variations, the value of sps_tile_fraction_slices_enabled_flag may be constrained or inferred from other syntax elements. For example, sps_tile_fraction_slices_enabled_flag is inferred to be equal to 0 when subpictures are not used in the video sequence (ie, subpics_present_flag is equal to 0).
[0275] Variations of Embodiments 1 and 2 may consider the value of sps_tile_fraction_slices_enabled_flag in a similar manner to infer whether there is a signaled tile fragment slice. For example, the PPS presented above may be modified as follows:
[0276]
[0277]
[0278] When sps_tile_fraction_slices_enabled_flag is equal to 0, the signaling of slice height is inferred in units of tiles.
[0279] Example 4
[0280] In VVC7, tile fragment stripes are enabled only for rectangular stripe mode. Embodiment 4 described below has the advantage of enabling tile fragment stripes to be used also in raster scan stripe mode. This provides the possibility of being able to adjust the length of the coded stripe in bits more accurately, because the stripe boundaries are not constrained to be aligned with the tile boundaries as in VVC7.
[0281] The principle consists in defining the slice partition (or providing information about the slice partition) in two places. The parameter set defines the slice in units of tiles. The tile fragment slice is signaled in the slice header. In a variant, sps_tile_fraction_slices_enabled_flag is predetermined equal to 1, and "tile fragments" are always present, signaled in the slice header.
[0282] To achieve this, in practice, the semantics of a slice is modified from that of VVC7 and the aforementioned embodiments / variants / further variations: a slice is a collection of one or more slice segments that collectively represent a block or an integer number of complete blocks of a picture. A slice segment represents an integer number of complete blocks within a picture, or an integer number of consecutive complete CTU rows within a block (i.e., a "block fragment"), that is, one or more blocks or "block fragments". A "block fragment" slice is a collection of consecutive CTU rows of a block. A slice segment is allowed to contain all CTU rows of a block. In this case, a slice segment contains a single slice segment.
[0283] According to a variation, the PPS syntax of any previous embodiment is modified to include tile segment specific signaling. An example of such a PPS syntax modification is shown below.
[0284] PPS Syntax
[0285] The PPS syntax of any previous embodiment is modified to remove the block segment specific signaling. For example, the PPS syntax is as follows:
[0286]
[0287]
[0288] For syntactic elements, the same semantics apply.
[0289] Stripe Section Syntax
[0290]
[0291] The slice segment NAL unit consists of a slice segment header and slice segment data, which is similar to the VVC7 NAL unit structure of a slice. The slice header from the previous embodiment becomes a slice segment header with the same syntax elements as in the slice header, but as a slice segment header, it includes additional syntax elements for locating / identifying slice segments in a slice (e.g., described / defined in the PPS).
[0292]
[0293]
[0294] The slice segment header includes a signal to specify which CTU row in the slice the slice segment starts with. slice_ctu_row_offset specifies the CTU row offset of the first CTU in the slice.
[0295] When rect_slice_flag is equal to 0 (i.e., the slice mode is in raster scan slice mode), the CTU row offset is relative to the first row of the block with an index equal to slice_address. When rect_slice_flag is equal to 1 (i.e., in rectangular slice mode), the CTU row offset is relative to the first CTU of the slice with an index equal to slice_address in the sub-picture identified by slice_subpic_id. The CTU row offset is encoded using variable or fixed length coding. For fixed length, the number of CTU rows in the slice is determined from the PPS, and the length in bits of the syntax element is equal to log2(the number of CTU rows minus 1).
[0296] There are two ways to indicate the end of a slice segment.
[0297] In the first approach, a slice segment indicates the number of CTU rows in the slice segment (minus 1). The number of CTU rows is encoded using variable length coding or fixed length coding. For fixed length coding, the number of CTU rows in the slice is determined from the PPS. The length in bits of the syntax element is equal to log2 (the difference between the number of CTU rows in the slice minus the CTU row offset minus 1).
[0298] The syntax of a stripe header is as follows:
[0299]
[0300]
[0301] num_ctu_rows_in_slice_minus1 plus 1 specifies the number of CTU rows in the slice segment NAL unit. The range of num_ctu_rows_in_slice_minus1 is 0 to the number of CTU rows in the blocks contained in the slice minus 2.
[0302] When sps_tile_fraction_slices_enabled_flag is equal to 1 and num_tiles_in_slice_segment_minus1 is equal to 0, the variable NumCtuInCurrSlice, which specifies the number of CTUs in the current slice, is equal to the number of CTU rows multiplied by the width in CTUs of the blocks present in the slice.
[0303] In the second approach, the slice segment data includes a signal to specify whether the slice segment ends at the end of each CTU row. The advantage of this second approach is that the encoder does not have to predetermine the number of CTUs in a given slice segment. This reduces the delay of the encoder, and the encoder can output the slice header in real time, while for the first approach, the slice header must be buffered to indicate the number of CTU rows in the slice segment at the end of encoding of the slice segment.
[0304] Example 5
[0305] Embodiment 5 is a signaled modification of the sub-picture layout, which can achieve improvements in certain situations. In fact, increasing the number of sub-pictures or slices or blocks in a video sequence will limit / limit the effectiveness / efficiency of the time and intra-frame prediction mechanism. As a result, the compression efficiency of the video sequence can be reduced. For this reason, there is a high possibility that the sub-picture layout can be predetermined / determined / predicted / estimated according to the application requirements (such as the size of the ROI). The encoding process then generates a block partition that will best match the sub-picture layout. In the best case scenario, each sub-picture contains exactly one block. In order to limit the impact on compression efficiency, the encoder attempts to minimize the number of slices for each sub-picture by using a single slice for each sub-picture. Therefore, the best option for the encoder is to define one slice and one block for each sub-picture.
[0306] In this case, the sub-picture layout and the slice layout are the same. This embodiment adds a flag in the SPS to indicate this specific case / scenario / situation. When the flag is equal to 1, the sub-picture layout does not exist and can be inferred / derived / determined to be the same as the slice partition. Otherwise, when the flag is equal to 0, the sub-picture layout is explicitly signaled in the bitstream according to the content described above with respect to the previous embodiment / variant / further variant.
[0307] According to a variant, the SPS includes the sps_single_slice_per_subpicture flag for this purpose:
[0308] sps_single_slice_per_subpicture equal to 1 indicates that each sub-picture includes a single slice and subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i] and subpic_height_minus1[i] are not present (i is in the range of 0 to sps_num_subpics_minus1 (inclusive)). sps_single_slice_per_subpicture equal to 0 indicates that the sub-picture may or may not include a single slice and subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i] and subpic_height_minus1[i] are present (i is in the range of 0 to sps_num_subpics_minus1 (inclusive)).
[0309] According to yet another variation, the PPS syntax includes the following syntax elements to indicate that the sub-picture layout can be inferred from the slice layout:
[0310]
[0311] If pps_single_slice_per_subpic_flag or sps_single_slice_per_subpic_flag is equal to 1, there is a single slice for each sub-picture. When sps_single_slice_per_subpic_flag is equal to 1, from the SPS, there is no slice layout, and pps_single_slice_per_subpic_flag must be equal to 0. The PPS then specifies the slice partitioning. The i-th sub-picture has a size and position corresponding to the i-th slice (i.e., the i-th sub-picture and the i-th slice have the same size and position).
[0312] When sps_single_slice_per_subpic_flag is equal to 0, the slice layout exists in the SPS, and pps_single_slice_per_subpic_flag can be equal to 1 or 0. When sps_single_slice_per_subpic_flag is equal to 1, then the SPS specifies sub-picture partitioning. The i-th slice has a size and position corresponding to the i-th sub-picture (i.e., the i-th slice and the i-th sub-picture have the same size and position).
[0313] In order to maintain the same sub-picture layout of the encoded video sequence, the encoder may constrain all PPSs that reference an SPS with sps_single_slice_per_subpic_flag equal to 1 to describe / define / enforce the same slice partitioning.
[0314] In a variant, another flag (pps_single_slice_per_tile) is provided in the PPS to indicate that there is a single slice for each tile. When this flag is equal to 1, the slice partition is inferred to be equal to the tile partition (i.e., the same as the tile partition). In this case, when sps_single_slice_per_subpic_flag is equal to 1, the sub-picture and slice partition is inferred to be the same as the tile partition.
[0315] Implementation of the embodiment of the present invention
[0316] One or more of the aforementioned embodiments / variants may be implemented in the form of an encoder or decoder, which performs one or more of the method steps of the aforementioned embodiments / variants. The following embodiments illustrate such an implementation.
[0317] Figure 9a is a flowchart showing the steps of an encoding method according to an embodiment / variant of the present invention, and Figure 9b is a flow chart showing the steps of a decoding method according to an embodiment / variant of the present invention.
[0318] according to Figure 9a In the encoding method, sub-picture partition information is obtained at 9911, and slice partition information is obtained at 9912. At 9915, the obtained information is used to determine information for determining one or more of the following items: the number of slices in the sub-picture; whether only a single slice is included in the sub-picture; and / or whether the slice can include tile fragments. Then at 9919, data for obtaining the determined information is encoded, for example by providing the data in a bitstream.
[0319] according to Figure 9b In a decoding method, at 9961, data is decoded (e.g., from a bitstream) to obtain information for determining the following items: the number of slices in a sub-picture; whether only a single slice is included in the sub-picture; and / or whether the slice can include tile segments. At 9964, the obtained information is used to determine one or more of the following items: the number of slices in the sub-picture; whether only a single slice is included in the sub-picture; and / or whether the slice can include tile segments. Then at 9967, based on the determination and the result thereof, sub-picture partition information and / or slice partition information is determined.
[0320] It should be understood that any of the foregoing embodiments / variations may be Fig.10 an encoder in (e.g., when performing segmentation into blocks 9402, entropy coding 9409, and / or bitstream generation 9410) or Fig.11 Used by the decoder in (e.g., when performing bitstream processing 9561, entropy decoding 9562 and / or video signal generation 9569).
[0321] Fig.10 A block diagram of an encoder according to an embodiment of the invention is shown. The encoder is represented by connected modules, each module being suitable for implementing at least one corresponding step of a method for implementing at least one embodiment of encoding an image of an image sequence according to one or more embodiments / variants of the invention, e.g. in the form of programming instructions executed by a central processing unit (CPU) of the device.
[0322] The original sequence of digital images i0 to in 9401 is received as input by the encoder 9400. Each digital image is represented by a set of samples, sometimes also referred to as pixels (hereinafter, they are referred to as pixels). After the encoding process is implemented, a bitstream 9410 is output by the encoder 9400. The bitstream 9410 includes data of multiple coding units or image parts such as slices, each slice including a slice header for sending a coding value of a coding parameter for encoding the slice and a slice body including coded video data. The input digital images i0 to in 9401 are divided into pixel blocks by module 9402. A block corresponds to an image part (hereinafter, an image part means a part of any type of image, such as a block, a slice, a slice segment or a sub-picture) and can be of variable size (for example, 4×4, 8×8, 16×16, 32×32, 64×64, 128×128 pixels and several rectangular block sizes can also be considered). A coding mode is selected for each input block.
[0323] Two types of coding modes are provided: coding modes based on spatial prediction coding (intra-frame prediction) and coding modes based on temporal prediction (eg inter-frame coding, merge, skip). Possible coding modes are tested.
[0324] Module 9403 implements an intra prediction process in which a given block to be encoded is predicted by a predictor calculated from the neighboring pixels of the given block to be encoded. If intra coding is selected, an indication of the selected intra predictor and the difference between the given block and its predictor are encoded to provide a residual.
[0325] Temporal prediction is implemented by the motion estimation module 9404 and the motion compensation module 9405. First, a reference image is selected from the reference image set 9416, and a portion of the reference image (also referred to as a reference region or image portion, which is the region closest to the given block to be encoded (closest in terms of pixel value similarity)) is selected by the motion estimation module 9404. The motion compensation module 9405 then uses the selected region to predict the block to be encoded. The difference (also referred to as residual block / data) between the selected reference region and the given block is calculated by the motion compensation module 9405. The selected reference region is indicated using motion information (e.g., a motion vector).
[0326] Therefore, in both cases (spatial and temporal prediction), when not in SKIP mode, the residual is calculated by subtracting the predictor from the original block.
[0327] In the intra prediction implemented by module 9403, the prediction direction is encoded. In the inter prediction implemented by modules 9404, 9405, 9416, 9418, 9417, at least one motion vector or information (data) for identifying such a motion vector is encoded for temporal prediction.
[0328] If inter prediction is selected, information related to the motion vector and the residual block is encoded. To further reduce the bit rate, the motion vector is encoded by the difference relative to the motion vector predictor, assuming that the motion is uniform. The motion vector predictor from the motion information predictor candidate set is obtained from the motion vector field 9418 by the motion vector prediction and encoding module 9417.
[0329] The encoder 9400 also includes a selection module 9406 for selecting a coding mode by applying a coding cost criterion such as a rate-distortion criterion. To further reduce redundancy, a transform module 9407 applies a transform (such as DCT) to the residual block, and the obtained transform data is then quantized by a quantization module 9408 and entropy encoded by an entropy encoding module 9409. Finally, when not in SKIP mode and the selected coding mode requires encoding of the residual block, the encoded residual block of the current block being encoded is inserted into the bitstream 9410.
[0330] The encoder 9400 also performs decoding of the encoded image to generate a reference image for motion estimation of subsequent images (e.g., a reference image in a reference image / picture 9416). This enables the encoder and decoder receiving the bitstream to have the same reference frame (e.g., using a reconstructed image or a reconstructed image portion). The inverse quantization ("dequantization") module 9411 performs inverse quantization ("dequantization") of the quantized data, which is then inversely transformed by the inverse transform module 9412. The intra-frame prediction module 9413 uses the prediction information to determine which predictor to use for a given block, and the motion compensation module 9414 actually adds the residual obtained by module 9412 to the reference area obtained from the reference image set 9416. The reconstructed frame (image or image portion) of pixels is then filtered by module 9415 with post filtering to obtain another reference image of the reference image set 9416.
[0331] Fig.11 A block diagram of a decoder 9560 that can be used to receive data from an encoder according to an embodiment of the present invention is shown. The decoder is represented by connected modules, each module being suitable for implementing the corresponding steps of the method implemented by the decoder 9560, for example in the form of programming instructions executed by the CPU of the device.
[0332] The decoder 9560 receives a bitstream 9561, which includes coded units (e.g., data corresponding to image parts, blocks, or coding units), each coding unit consisting of a header containing information about coding parameters and a body containing coded video data. Fig.10As illustrated, the coded video data is entropy encoded on a predetermined number of bits for a given image portion (e.g., a block or CU), and motion information (e.g., an index to a motion vector predictor) is encoded. The received coded video data is entropy decoded by module 9562. The residual data is then dequantized by module 9563, and an inverse transform is then applied by module 9564 to obtain pixel values.
[0333] Mode data indicating the coding mode is also entropy decoded, and based on the mode, the coding block (unit / set / group) of the image data is intra-type decoded or inter-type decoded. In the case of intra-frame mode, the intra-frame prediction module 9565 determines the intra-frame predictor based on the intra-frame prediction mode specified in the bitstream (for example, the intra-frame prediction mode can be determined using the data provided in the bitstream). If the mode is an inter-frame mode, motion prediction information is extracted / obtained from the bitstream to find (identify) the reference area used by the encoder. For example, motion prediction information includes a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual to obtain a motion vector by the motion vector decoding module 9570.
[0334] The motion vector decoding module 9570 applies motion vector decoding to each image portion (e.g., current block or CU) encoded by motion prediction. Once the index of the motion vector predictor of the current block is obtained, the actual value of the motion vector associated with the image portion (e.g., current block or CU) can be decoded and used to apply motion compensation by module 9566. The reference image portion indicated by the decoded motion vector is extracted / obtained from the reference image set 9568 so that module 9566 can perform motion compensation. The motion vector field data 9571 is updated with the decoded motion vector for use in predicting subsequently decoded motion vectors.
[0335] Finally, a decoded block is obtained. Where appropriate, post-filtering is applied by post-filtering module 9567. Finally, a decoded video signal 9569 is obtained and provided by decoder 9560.
[0336] Fig.12The data communication system in which one or more embodiments of the present invention can be implemented is illustrated. The data communication system includes a transmission device (in this case, a server 9201) that is operable to transmit data packets of a data stream to a receiving device (in this case, a client terminal 9202) via a data communication network 9200. The data communication network 9200 can be a wide area network (WAN) or a local area network (LAN). Such a network can be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network consisting of several different networks. In a specific embodiment of the present invention, the data communication system can be a digital television broadcasting system in which the server 9201 sends the same data content to multiple clients.
[0337] The data stream 9204 provided by the server 9201 can be composed of multimedia data representing video and audio data. In some embodiments of the present invention, audio and video data streams can be captured by the server 9201 using a microphone and a camera, respectively. In some embodiments, the data stream can be stored on the server 9201 or received by the server 9201 from other data providers, or generated at the server 9201. The server 9201 is provided with an encoder for encoding video and audio streams, in particular to provide a compressed bit stream for transmission, which is a more compact representation of the data presented as the input of the encoder. In order to obtain a better ratio of the quality of the transmitted data to the amount of the transmitted data, the video data can be compressed, for example, according to the high efficiency video coding (HEVC) format or the H.264 / advanced video coding (AVC) format or the multi-purpose video coding (VVC) format. The client 9202 receives the transmitted bit stream and decodes the reconstructed bit stream to reproduce the video image on the display device and reproduce the audio data using a speaker.
[0338] Although a streaming scenario is considered in this embodiment, it will be appreciated that in some embodiments of the invention, data communication between the encoder and decoder may be performed using, for example, a media storage device such as an optical disk, etc. In one or more embodiments of the invention, video images may be transmitted along with data representing compensation offsets to be applied to reconstructed pixels of the image to provide filtered pixels in the final image.
[0339] Fig.13 Schematically illustrates a processing device 9300 configured to implement at least one embodiment / variant of the present invention. The processing device 9300 may be a device such as a microcomputer, a workstation, a user terminal, or a light portable device. The device / apparatus 9300 includes a communication bus 9313, which is connected to:
[0340] - a central processing unit 9311 denoted as CPU, such as a microprocessor, etc.;
[0341] - a read-only memory 9307, denoted ROM, for storing computer programs / instructions for operating the device 9300 and / or implementing the invention;
[0342] - a random access memory 9312 represented as a RAM for storing executable codes of the method of the embodiment / variant of the present invention, and registers suitable for recording variables and parameters required for implementing the method for encoding a digital image sequence and / or the method for decoding a bit stream according to the embodiment / variant of the present invention; and
[0343] - A communication interface 9302 connected to a communication network 9303, through which digital data to be processed are transmitted or received.
[0344] Optionally, the device 9300 may further include the following components:
[0345] - a data storage component 9304, such as a hard disk, for storing a computer program for implementing a method of one or more embodiments / variants of the present invention and data used or generated during the implementation of one or more embodiments / variants of the present invention;
[0346] a disk drive 9305 for a disk 9306 (e.g. a storage medium), the disk drive 9305 being adapted to read data from the disk 9306 or to write data to said disk 9306, or;
[0347] - A screen 9309 for displaying data and / or serving as a graphical interface for interaction with a user by means of a keyboard 9310, a touch screen or any other indication / input means.
[0348] The device 9300 may be connected to various peripheral devices such as a digital camera 9320 or a microphone 9308 , each of which is connected to an input / output card (not shown) to provide multimedia data to the device 9300 .
[0349] The communication bus provides communication and interoperability between the various elements included in or connected to the device 9300. The representation of the bus is not limiting, and in particular, the central processing unit 9311 is operable to communicate instructions to any element of the device 9300 directly or with the aid of other elements of the device 9300.
[0350] The disk 9306 may be replaced by any information medium, such as a rewritable or non-rewritable compact disk (CD-ROM), a ZIP disk or a memory card, and in general by an information storage component that can be read by a microcomputer or a processor, the disk 9306 being integrated into the device or not, possibly removable and suitable for storing one or more programs whose execution enables the implementation of the method for encoding a digital image sequence and / or the method for decoding a bit stream according to the present invention.
[0351] The executable code may be stored in a read-only memory 9306, on a hard disk 9304 or on a removable digital medium such as, for example, a disk 9306 as previously described, etc. According to a variant, the executable code of the program may be received via the interface 9302 by means of a communication network 9303 to be stored in one of the storage means of the device 9300 (e.g., the hard disk 9304) before execution.
[0352] The central processing unit 9311 is suitable for controlling and directing the execution of instructions or parts of software code of one or more programs according to the present invention, the execution of instructions stored in one of the above-mentioned storage means. At power-on, one or more programs stored in non-volatile memory (for example, on hard disk 9304, disk 9306 or in read-only memory 9307) are transferred to random access memory 9312 (which then contains the executable code of one or more programs) and registers for storing variables and parameters necessary for implementing the present invention.
[0353] In this embodiment, the device is a programmable device that implements the invention using software. Alternatively, however, the invention may be implemented in hardware (for example in the form of an application specific integrated circuit or ASIC).
[0354] Implementation of the embodiment of the present invention
[0355] Implementation of the embodiment of the present invention
[0356] It should also be understood that according to other embodiments of the present invention, a decoder according to the above-mentioned embodiment / variant is provided in a user terminal such as a computer, a mobile phone (cellular phone), a tablet or any other type of device (e.g., a display device) capable of providing / displaying content to a user. According to yet another embodiment, an encoder according to the above-mentioned embodiment / variant is provided in an image capture device, which also includes a camera, a video camera or a network camera (e.g., a closed-circuit television or video surveillance camera) for capturing and providing content for encoding by the encoder. See below Fig.14 and 15 Two such embodiments are provided.
[0357] Fig.14is a diagram illustrating a network camera system 9450 including a network camera 9452 and a client device 9454 .
[0358] The network camera 9452 includes a camera unit 9456, a coding unit 9458, a communication unit 9460, and a control unit 9462. The network camera 9452 and the client device 9454 are connected to each other via the network 9200 so as to be able to communicate with each other. The camera unit 9456 includes a lens and an image sensor (e.g., a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS)), and captures an image of an object and generates image data based on the image. The image may be a still image or a video image. The camera unit may also include a zoom component and / or a pan component respectively adapted for zooming or panning (optically or digitally). The coding unit 9458 encodes the image data by using the coding method described in the aforementioned embodiment / variation. The coding unit 9458 uses at least one of the coding methods described in the aforementioned embodiment / variation. For other examples, the coding unit 9458 may use a combination of the coding methods described in the aforementioned embodiment / variation.
[0359] The communication unit 9460 of the network camera 9452 transmits the encoded image data encoded by the encoding section 9458 to the client device 9454. In addition, the communication unit 9460 can also receive commands from the client device 9454. The commands include commands for setting parameters for encoding by the encoding section 9458. The control unit 9462 controls other units in the network camera 9452 according to the commands received by the communication unit 9460.
[0360] The client device 9454 includes a communication unit 2114, a decoding unit 2116, and a control unit 9468. The communication unit 2118 of the client device 9454 can transmit a command to the network camera 9452. In addition, the communication unit 2118 of the client device 9454 receives the encoded image data from the network camera 9452. The decoding unit 9466 decodes the encoded image data by using the decoding method described in one or more of the aforementioned embodiments / variants. For other examples, the decoding unit 9466 can use a combination of the decoding methods described in the aforementioned embodiments / variants. The control unit 9468 of the client device 9454 controls other units in the client device 9454 according to the user operation or command received by the communication unit 2114. The control unit 9468 of the client device 9454 can also control the display device 9470 to display the image decoded by the decoding unit 9466.
[0361] The control unit 9468 of the client device 9454 also controls the display device 9470 to display a GUI (Graphical User Interface) for specifying the values of the parameters of the network camera 9452 (e.g., the parameters for encoding of the encoding section 9458). The control unit 9468 of the client device 9454 may also control other units in the client device 9454 according to the user operation input to the GUI displayed by the display device 9470. The control unit 9468 of the client device 9454 may also control the communication unit 9464 of the client device 9454 according to the user operation input to the GUI displayed by the display device 9470 to transmit a command for specifying the values of the parameters of the network camera 9452 to the network camera 9452.
[0362] Fig.15 9500 is a diagram illustrating a smartphone 9500. The smartphone 9500 includes a communication unit 9502, a decoding / encoding unit 9504, a control unit 9506, and a display unit 9508.
[0363] The communication unit 9502 receives the encoded image data via the network. The decoding / encoding unit 9504 decodes the encoded image data received by the communication unit 9502. The decoding unit 9504 decodes the encoded image data by using the decoding method described in one or more than one of the aforementioned embodiments / variants. The decoding / encoding unit 9504 may also use at least one of the encoding or decoding methods described in the aforementioned embodiments / variants. For other examples, the decoding / encoding unit 9504 may use a combination of the decoding or encoding methods described in the aforementioned embodiments / variants.
[0364] The control unit 9506 controls other units in the smartphone 9500 according to a user operation or a command received by the communication unit 9502. For example, the control unit 9506 controls the display unit 9508 to display an image decoded by the decoding / encoding section 9504.
[0365] The smartphone may also include an image recording device 9510 (e.g., a digital camera and associated circuitry) for recording images or videos. Such recorded images or videos may be encoded by the decoding / encoding section 9504 under the instruction of the control unit 9506. The smartphone may also include a sensor 9512 suitable for sensing the orientation of the mobile device. Such a sensor may include an accelerometer, a gyroscope, a compass, a global positioning (GPS) unit, or a similar position sensor. Such a sensor 2212 may determine whether the smartphone changes orientation, and such information may be used when encoding the video stream.
[0366] Although the present invention has been described with reference to embodiments and variations thereof, it should be understood that the present invention is not limited to the disclosed embodiments / variations. It will be appreciated by those skilled in the art that various changes and modifications may be made without departing from the scope of the present invention as defined by the appended claims. All features disclosed in this specification (including any appended claims, abstracts and drawings), and / or all steps of any disclosed method or process, may be combined in any combination, except for at least some mutually exclusive combinations of such features and / or steps. Unless expressly stated otherwise, each feature disclosed in this specification (including any appended claims, abstracts and drawings) may be replaced by alternative features for the same, equivalent or similar purposes. Therefore, unless expressly stated otherwise, each feature disclosed is only an example of a general series of equivalent or similar features.
[0367] It should also be understood that any result of the above-mentioned comparison, determination, inference, evaluation, selection, execution, conduct or consideration (e.g., a selection made during encoding, processing or partition processing) can be indicated in data in the bitstream (e.g., a flag or information indicating the result) or can be determined / inferred from data in the bitstream, so that the indicated or determined / inferred result can be used in the processing instead of actually performing the comparison, determination, evaluation, selection, execution, conduct or consideration, such as during decoding or partition processing. It should be understood that when a "table" or "lookup table" is used, other data types such as arrays can also be used to perform the same function, as long as the data type can perform the same function (e.g., represent the relationship / mapping between different elements).
[0368] In the claims, the word "comprising" does not exclude other elements or steps and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage. Reference signs appearing in the claims are by way of illustration only and shall not have a limiting effect on the scope of the claims.
[0369] In the foregoing embodiments / variations, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or sent through a computer-readable medium as one or more instructions or codes, and may be executed by a hardware-based processing unit.
[0370] Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media or communication media including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementing the techniques described in the present invention. A computer program product may include a computer-readable medium.
[0371] As an example and not limitation, this computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, flash memory or any other medium that can be used to store the desired program code in the form of an instruction or data structure and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer-readable medium. For example, if a coaxial cable, optical fiber cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave, etc.) is used to send instructions from a website, server or other remote source, then the coaxial cable, optical fiber cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave, etc.) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connection, carrier wave, signal or other transient media, but are for non-transient tangible storage media. The disk and disc used here include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and blue disc, wherein the disc usually copies data magnetically, and the disc reproduces data optically by laser. Combinations of the above should also be included within the scope of computer-readable media.
[0372] Instructions may be executed by one or more processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Furthermore, the technology may be fully implemented in one or more circuits or logic elements.
Claims
1. A method for decoding image data of one or more images, each image consisting of one or more blocks and being divisible into one or more slices, wherein: A slice can include a portion of a block, and the image can be partitioned into one or more sub-pictures, and the method comprises: obtaining information indicating a width of a sub-picture and information indicating a height of the sub-picture; determining parameters associated with one or more slices included in the sub-picture using the obtained information indicating the width of the sub-picture and the obtained information indicating the height of the sub-picture; and The one or more images are decoded using parameters obtained from the determining.
2. The method according to claim 1, wherein: The portion of the block is an integer number of consecutive complete coding tree unit rows, ie, CTU rows, within the block.
3. The method according to claim 1 or 2, further comprising: The one or more slices are defined using an identifier of the sub-picture.
4. The method according to claim 1 or 2, further comprising: One or more than one slice is defined based on whether only a single slice is included in the sub-picture.
5. The method according to claim 1, wherein: The parameter associated with the one or more slices is the number of slices in the sub-picture.
6. The method according to claim 1, wherein: In case a slice includes one or more than one portion of a block, the slice is determined based on the number of rows or columns of coding tree units, ie, CTUs, to be included in the slice.
7. The method according to claim 1, further comprising: Obtain information from the bit stream for determining one or more of the following: whether only a single slice is included in the sub-picture; as well as The number of slices included in the sub-picture.
8. The method according to claim 7, wherein: The information used for determination is obtained from the picture parameter set, ie, PPS.
9. The method according to claim 7, wherein: The information used for determination is obtained from a sequence parameter set, namely, an SPS.
10. The method according to claim 1, wherein: The sub-picture includes two or more slices, and each slice includes one or more parts of the block.
11. The method according to claim 10, wherein: One or more parts of the block are from the same single block.
12. The method according to claim 1, wherein: A stripe consists of one or more than one block, which forms a rectangular area in the image.
13. The method according to claim 12, further comprising: Information for determining whether only a single slice is included in the sub-picture is obtained from a picture parameter set, ie, a PPS, from a bitstream.
14. The method according to claim 12, further comprising: Information for determining the number of slices included in the sub-picture is obtained from a picture parameter set, ie, a PPS, from a bitstream.
15. The method according to claim 12, further comprising: Information indicating whether a sub-picture is used in a video sequence is obtained from a sequence parameter set, ie, SPS, from the bitstream.
16. A method for encoding image data of one or more images, each image consisting of one or more blocks and being divisible into one or more slices, wherein: A slice can include a portion of a block, and the image can be partitioned into one or more sub-pictures, and the method comprises: providing information indicating a width of a sub-picture and information indicating a height of the sub-picture; determining parameters of one or more slices included in the sub-picture using the information indicating the width of the sub-picture and the information indicating the height of the sub-picture; and The one or more images are encoded using parameters obtained from the determining.
17. A device for decoding one or more images, the device comprising a memory, a computer or a processor, and one or more executable instructions stored on the memory, wherein when the one or more executable instructions are executed on the computer or the processor, the computer or the processor executes the decoding method according to any one of claims 1 to 15.
18. A device for encoding one or more images, the device comprising a memory, a computer or a processor, and one or more executable instructions stored on the memory, wherein when the one or more executable instructions are executed on the computer or the processor, the computer or the processor executes the encoding method according to claim 16.
19. A computer-readable storage medium carrying one or more executable instructions, which, when executed on a computer or a processor, causes the computer or the processor to perform the method according to any one of claims 1 to 15.
20. A computer-readable storage medium carrying one or more executable instructions, which, when executed on a computer or a processor, cause the computer or the processor to perform the method according to claim 16.
Citation Information
Patent Citations
Media information processing method and device
CN110035331A
Apparatus and method of coding of pictures
WO2019016287A1