Method and apparatus for encoding or decoding video data having frame portions

By dividing video frames into frame parts and encoding them using frame part identifiers and spatial information, the problems of encoding efficiency and decoding complexity in large video content processing of the HEVC mechanism are solved, achieving more efficient encoding and decoding operations.

CN116033151BActive Publication Date: 2025-10-24CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211641848.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-09
Filing Date
2019-03-25
Publication Date
2025-10-24
Estimated Expiration
2039-03-25

AI Technical Summary

Technical Problem

The existing HEVC mechanism struggles to effectively utilize blocks for flexible encoding and decoding when processing large video content. In particular, when combining different blocks or frame parts, data rewriting is required, which affects encoding efficiency and decoding complexity.

Method used

By dividing video frames into frame parts and using frame part identifiers and spatial information for encoding during the encoding process, independent encoding and decoding are allowed, reducing data rewriting.

Benefits of technology

It improves the flexibility and efficiency of encoding, simplifies frame operations, enhances compression performance, and adapts to the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116033151B_ABST
    Figure CN116033151B_ABST
Patent Text Reader

Abstract

The invention relates to a method for encoding a frame into a bitstream, said frame being spatially divided into frame portions, said method comprising: encoding a frame portion in said frame into said bitstream; representing in said bitstream an identifier of each of said frame portions in said frame; and representing spatial information related to the position of a frame portion within said frame, wherein said identifier and said spatial information are represented in a parameter set in said bitstream, and wherein the number of bits used to represent said identifier is further represented in said bitstream, wherein the number of bits used to represent said identifier is variable.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] (This application is a divisional application of application No. 2019800247039, filed on March 25, 2019, with the title "Method and apparatus for encoding or decoding video data with frame portions".) TECHNICAL FIELD

[0002] The present invention relates to a method and apparatus for encoding or decoding video data comprising spatial portions. BACKGROUND

[0003] HEVC tiles were introduced and designed for parallel encoding. However, for large size video contents, several use cases exist to use tiles in different ways. In particular, the need to stream individual tiles or sets of tiles has developed. Some applications also developed the need to combine different tiles from the same sequence or from different sequences to compose new video sequences.

[0004] Current mechanisms in HEVC were designed without taking into account such scenarios. Implementing these scenarios with current HEVC mechanisms implies adding encoding constraints on tiles and any composition of tiles at decoding involves a rewriting process of data. In particular, a rewriting of data including manipulation of slice segment headers is usually required. SUMMARY

[0005] The present invention is designed to solve one or more of the above problems. The present invention relates to the definition of frame portions and to the signaling of these frame portions in a bitstream. The purpose of the present invention is to facilitate the extraction and re-composition of these frame portions at decoding while limiting the rewriting process required to do so.

[0006] According to a first aspect of the present invention, there is provided a method for encoding video data comprising frames into a bitstream, said frames being spatially divided into frame portions, said method comprising:

[0007] - encoding at least one frame portion into one or more first encoding units;

[0008] wherein the method further comprises:

[0009] - signaling in said first encoding units at least one frame portion identifier, said frame portion identifier identifying one encoded frame portion; and

[0010] - providing frame portion configuration information, said frame portion configuration information comprising said frame portion identifier and spatial information related to said frame portion.

[0011] The first aspect of the invention has the advantage of providing greater flexibility and simpler manipulation, while enabling improved compression compared to known block-based designs such as HEVC blocks.

[0012] In an embodiment, the frame portion configuration information is provided in a second coding unit.

[0013] In an embodiment, the at least one frame portion is independently coded.

[0014] In an embodiment, the method further comprises providing a flag indicating that the frame portion has been independently coded.

[0015] In an embodiment, the one or more first coding units comprise a flag indicating for each frame portion that the frame portion has been independently coded.

[0016] In an embodiment, the one or more first coding units comprise a flag indicating that the at least one frame portion has been independently coded.

[0017] In an embodiment, the one or more first coding units comprise a flag indicating a level of coding constraint used for coding the frame portion.

[0018] In an embodiment, the frame portion is a slice and the first coding unit is a slice unit comprising a data portion, the flag being comprised in a slice segment header of the data portion of the slice unit.

[0019] In an embodiment, the frame portion is a slice and the first coding unit is a slice unit comprising a data portion, the frame portion identifier being comprised in a slice segment header of the data portion of the slice unit.

[0020] In an embodiment, the first coding unit comprises a header portion and a data portion, the data portion comprising the coded frame portion, the frame portion identifier being comprised in the header portion.

[0021] In an embodiment, a frame portion identifier is represented in all frame portion coding units, and a predefined frame portion identifier value indicates that the frame portion is not independently coded.

[0022] In an embodiment, the second coding unit is a parameter set dedicated to information related to one or more frames.

[0023] In an embodiment, the second coding unit is a parameter set dedicated to frame portion information.

[0024] In an embodiment, the first coding unit has a specific type indicating that the frame portion has been independently coded.

[0025] In embodiments, the frame portion identifier is encoded using a fixed predetermined number of bits.

[0026] In embodiments, the frame portion identifier is encoded using a variable number of bits.

[0027] In embodiments, the spatial information comprises a position of the frame portion given by a coding tree unit address.

[0028] In embodiments, the spatial information comprises a position of the frame portion given by a sample address.

[0029] In embodiments, the spatial information comprises a size of the frame portion.

[0030] In embodiments, the position of the frame portion is given with respect to the frame.

[0031] In embodiments, a plurality of parameter data units is represented in the bitstream for including different frame portion configurations for the same frame portion.

[0032] In embodiments, the second coding unit comprises a flag indicating whether a given post-filtering algorithm can be used for the frame portion.

[0033] In embodiments, a same frame portion identifier can be used to identify a plurality of frame portions defining a set of frame portions.

[0034] In embodiments, the header portion comprises a layer identifier, and the layer identifier is used to represent the frame portion identifier.

[0035] According to a second aspect of the application, there is provided a method for decoding video data comprising frames from at least one bitstream, the frames being spatially divided into frame portions, the method comprising:

[0036] - obtaining frame portion configuration information from the bitstream, the frame portion configuration information comprising a frame portion identifier and spatial information related to the frame portion;

[0037] - extracting at least a frame portion from one or more first coding units in the bitstream, the frame portion comprising the frame portion identifier;

[0038] - determining a position of the frame portion within the frame based on the spatial information; and

[0039] - decoding the frame portion according to the determined position to render the frame portion into a frame.

[0040] In embodiments, the frame portion configuration information is provided into one second coding unit.

[0041] In embodiments, the at least one frame portion is independently coded.

[0042] In embodiments, the method further comprises obtaining a flag indicating that the frame portion has been independently coded.

[0043] In embodiments, the one or more first coding units comprise a flag indicating for each frame portion that the frame portion has been independently coded.

[0044] In embodiments, the one or more first coding units comprise a flag indicating that the at least one frame portion has been independently coded.

[0045] In embodiments, the one or more first coding units comprise a flag indicating a level of coding constraint used for coding the frame portion.

[0046] In embodiments, the frame portion is a slice and the first coding unit is a slice unit comprising a data portion, the flag being comprised in a slice segment header of the data portion of the slice unit.

[0047] In embodiments, the frame portion is a slice and the first coding unit is a slice unit comprising a data portion, the frame portion identifier being comprised in a slice segment header of the data portion of the slice unit.

[0048] In embodiments, the first coding unit comprises a header portion and a data portion, the data portion comprising the coded frame portion, the frame portion identifier being comprised in the header portion.

[0049] In embodiments, a frame portion identifier is represented in all frame portion coding units and a predefined frame portion identifier value indicates that the frame portion is not independently coded.

[0050] In embodiments, the second coding unit is a parameter set dedicated to information related to one or more frames.

[0051] In embodiments, the second coding unit is a parameter set dedicated to frame portion information.

[0052] In embodiments, the first coding unit has a specific type indicating that the frame portion has been independently coded.

[0053] In embodiments, the frame portion identifier is coded using a fixed predetermined number of bits.

[0054] In embodiments, the frame portion identifier is coded using a variable number of bits.

[0055] In embodiments, the spatial information comprises a position of the frame portion given by a coding tree unit address.

[0056] In embodiments, the spatial information comprises a position of the frame portion given by a sample address.

[0057] In embodiments, the spatial information comprises a size of the frame portion.

[0058] In embodiments, the position of the frame portion is given with respect to the frame.

[0059] In embodiments, a plurality of parameter data units is obtained from the bitstream comprising different frame portion configurations for the same frame portion.

[0060] In embodiments, the second encoding unit comprises a flag indicating whether a given post-filtering algorithm can be used for the frame portion.

[0061] In embodiments, a same frame portion identifier can be used to identify a plurality of frame portions defining a set of frame portions.

[0062] In embodiments, the header portion comprises a layer identifier, and the layer identifier is used to represent the frame portion identifier.

[0063] According to a third aspect of the application, there is provided a method for generating a new bitstream, the new bitstream comprising video data comprising frames, the frames being spatially divided into frame portions, the method comprising:

[0064] - determining a plurality of frame portions to be extracted from a plurality of bitstreams and merged into the new bitstream, the plurality of bitstreams being encoded according to any of claims 1 to 25;

[0065] - determining a frame portion identifier of the frame portions to be extracted;

[0066] - generating frame portion configuration information for the new bitstream;

[0067] - extracting the plurality of frame portions to be extracted from the plurality of bitstreams; and

[0068] - embedding the plurality of frame portions and the generated frame portion configuration information into the new bitstream.

[0069] In embodiments, the method further comprises:

[0070] - determining a new frame portion identifier for the extracted frame portions; and

[0071] - replacing the frame portion identifier with the new frame portion identifier of the extracted frame portions.

[0072] In embodiments, extracting the plurality of frame portions comprises:

[0073] - parsing the plurality of bitstreams; and

[0074] - extracting a frame portion coded data unit, the frame portion coded data unit comprising one of the determined frame portion identifiers.

[0075] According to a fourth aspect of the application, there is provided an apparatus for encoding video data comprising frames into a bitstream, the frames being spatially divided into frame portions, the apparatus comprising circuitry configured to:

[0076] - encode at least one frame portion into one or more first coded units;

[0077] wherein the method further comprises:

[0078] - representing at least one frame portion identifier in the first coded units, the frame portion identifier identifying one coded frame portion; and

[0079] - providing frame portion configuration information, the frame portion configuration information comprising the frame portion identifier and spatial information relating to the frame portion.

[0080] According to a fifth aspect of the application, there is provided an apparatus for decoding video data comprising frames from at least one bitstream, the frames being spatially divided into frame portions, the apparatus comprising circuitry configured to:

[0081] - obtaining frame portion configuration information from the bitstream, the frame portion configuration information comprising frame portion identifiers and spatial information relating to the frame portions;

[0082] - extracting at least a frame portion from one or more first coded units in the bitstream, the frame portion comprising the frame portion identifier;

[0083] - determining a position of the frame portion within the frame based on the spatial information; and

[0084] - decoding the frame portion according to the determined position to render the frame portion into a frame.

[0085] According to a sixth aspect of the application, there is provided an apparatus for generating a new bitstream, the new bitstream comprising video data comprising frames, the frames being spatially divided into frame portions, the apparatus comprising circuitry configured to:

[0086] - determining a plurality of frame portions to be extracted from a plurality of bitstreams and merged into a new bitstream, the plurality of bitstreams being encoded according to any of claims 1 to 25;

[0087] - determining frame portion identifiers of the frame portions to be extracted;

[0088] - generating frame portion configuration information for the new bitstream;

[0089] - extracting a plurality of frame portions to be extracted from the plurality of bitstreams; and

[0090] - embedding the plurality of frame portions and the generated frame portion configuration information into the new bitstream.

[0091] According to a seventh aspect of the application, there is provided a computer program product for a programmable device, the computer program product comprising a sequence of instructions for implementing the method according to the application when loaded into and executed by the programmable device.

[0092] According to an eighth aspect of the application, there is provided a computer readable storage medium having stored thereon instructions of a computer program for implementing the method according to the application.

[0093] According to a ninth aspect of the application, there is provided a computer program which, when executed, causes the method of the application to be performed.

[0094] At least part of the method according to the application can be implemented by a computer. Thus, the application can take the form of an embodiment which is entirely hardware implemented; an embodiment which is entirely software implemented (including firmware, resident software, microcode, etc.); or an embodiment which combines software and hardware aspects, which can all be generally referred to herein as "circuitry", "modules" or "systems". Furthermore, the application can take the form of a computer program product embodied in any tangible medium with a computer-usable program code embodied therein.

[0095] Since the application can be implemented in software, the application can be embodied in computer-readable code which is provided on any suitable carrier medium. A tangible, non-transitory carrier medium can include a storage medium such as a floppy disk, a CD-ROM, a hard disk drive, a magnetic tape device or a solid state memory device, etc. A transitory carrier medium can include a signal such as an electrical signal, an electronic signal, an optical signal, an acoustic signal, a magnetic signal or an electromagnetic signal such as an electromagnetic wave or a RF signal. BRIEF DESCRIPTION OF DRAWINGS

[0096] Embodiments of the application will now be described, by way of example only, and with reference to the following drawings in which:

[0097] Figure 1 a system into which the application can be integrated is shown;

[0098] Figure 2 a picture encoding structure of a block-based video encoder (e.g. HEVC) is shown;

[0099] Figure 3 Illustration of the partitioning of a picture according to two types of partitions, referred to as slice segments and tiles in HEVC;

[0100] Figure 4 Illustration of the quadtree inference mechanism for CTUs crossing the boundaries of a picture in HEVC;

[0101] Figure 5 Illustration of the border extension mechanism used for example in HEVC;

[0102] Figure 6 Illustration of an example of HEVC bitstream organization;

[0103] Figure 7 Illustration of an example of HEVC partitioning for Region of Interest (ROI) streaming;

[0104] Figure 8a and 8b Illustration of two different use case examples for the combination of regions of interest;

[0105] Figure 9 Illustration of a typical encoding process of a video encoder in which the application is integrated;

[0106] Figure 10 Illustration of a typical decoding process of a video decoder in which the application is integrated;

[0107] Figure 11 Description of an exemplary use of the application;

[0108] Figure 12 Details related to the encapsulation step are provided;

[0109] Figure 13a , 13b and 13c illustrate a representation of the configuration of frame parts made by the encoding process;

[0110] Figure 14 Illustration of an example of non-grid based partitioning;

[0111] Figure 15 Illustration of an alternative embodiment to represent the CTile identifier;

[0112] Figure 16a Illustration of an XPS including a list of dependencies of each CTile;

[0113] Figure 16b Illustration of a first example of CTile dependencies;

[0114] Figure 16c Illustration of a second example of CTile dependencies;

[0115] Figure 17a and 17b Examples are provided of embodiments where CTiles can change position or size between successive encoded frames; and

[0116] Figure 18 is a schematic block diagram of a computing device for implementing one or more embodiments of the application. DETAILED DESCRIPTION

[0117] Encoding frames of a video sequence into spatial frame portions is particularly useful for example in schemes related to streaming of so-called 360-degree videos, which are in fact the result of projecting a 360-degree panoramic or spherical video onto a classical 2D video representation.

[0118] A 360-degree video (or simply 360 video) is a video that can have a very high resolution to provide a good user experience. When displayed inside a head-mounted display (or on a screen), only a spatial sub-portion of the 360 video content is presented to the user.

[0119] It is thus of interest to request spatial frame portions with high quality only for the area that the user is looking at, for example with a streaming protocol like Dynamic Adaptive Streaming over HTTP (DASH). For the areas that are not seen (i.e. the areas that the user is not looking at), the spatial frame portions can simply be skipped.

[0120] An application of the present application relates to a streaming method that adapts the streaming to the viewing direction of the user. In other words, the application refers to viewport-dependent streaming. For such a method, a good trade-off between storage cost, computation cost and user experience is to encode the sequence into independent spatial frame portions with various qualities. Then, frame portions can be accessed, extracted and / or combined with other frame portion sequences on the fly according to the needs and bandwidth constraints. This application does not require any additional encoding or transcoding. Reference is made to Figure 8a for an example of such a scheme.

[0121] An application relates to a video surveillance system in which spatial frame portions of multiple different videos are reorganized in a new video to match a structure requested from a system operator. For example, the operator can want only parts of the original videos. In particular, in Figure 8b this application is illustrated.

[0122] Finally, in another application, a new "video" that includes only a single frame portion extracted from a complete video sequence can involve rewriting the encoding parameters in case of a new position of the spatial frame portion in the new video.

[0123] When using HEVC, the encoding of the spatial frame portions is based on HEVC tiles. However, HEVC tiles, and more generally HEVC-like tiles, are not designed to address the above-mentioned applications.

[0124] Figure 1 Embodiments of a system (e.g. an interactive streaming video system) in which the present application can be integrated are described.

[0125] A video bitstream is sent from a server or proxy server 100 to a client 102 through a network 101. The server 100 uses a video stream (or video file) generated by a video encoder 103, which complies with the specifications of a block-based video codec (e.g. the HEVC video codec).

[0126] As described below, the encoder compresses a set of video sequences with different rate / distortion trade-offs while providing spatial random access to some spatial frame portions according to the present application.

[0127] The server 100 receives a request for a description of a video stream available for interactive streaming through a communication network 101. The communication network 101 is based on Internet Protocol standards. The standard protocol employed for sending a media presentation over the IP network 101 is preferably MPEG DASH: Dynamic Adaptive Streaming over HTTP. However, the present application can also be used with any other streaming protocol.

[0128] Figure 2 The partitioning of an image according to two types of partitions (tile segments and spatial frame portions) is illustrated. The image 206 is partitioned into three tile segments. A tile segment is a portion of an image or the entire image. Each tile segment contains an integer number of coding blocks (which can correspond to the coding units of HEVC). A coding block is composed of samples.

[0129] The two types of tile segments are independent tile segments 207 and dependent tile segments 208. Each tile segment is embedded in one NAL unit, which is a structure with a generic format for use in packet-oriented and bitstream-oriented transport systems. The difference between the two types of tile segments lies in the fact that the data specified in the independent tile segment header defines all the parameters needed to decode the coding blocks of the tile segment. On the other hand, the header of a dependent tile segment is reduced and the dependent tile segment relies on the previous first independent tile segment to infer the parameters not available in its header. The address of the first coding unit in the tile is specified in the independent tile segment header.

[0130] Figure 3 The other partitioning into spatial frame portions (SPF) is illustrated to allow splitting each frame into independently coded rectangular regions as illustrated in frame 305.

[0131] Like HEVC-like tiles, a spatial frame portion contains an integer number of coding blocks. Like a slice boundary, an SPF boundary 310 breaks all intra prediction mechanisms.

[0132] Like HEVC-like tiles, an SPF is defined in the picture parameter set included within a specific NAL unit used to initialize the decoding process. The PPS NAL unit includes syntax elements that can specify the number of tile rows and the number of tile columns in a picture and their associated size. Other parameter set NAL units (e.g., video parameter set or VPS, sequence parameter set or SPS) convey parameters used to describe the coding structure of the bitstream. In the present invention, any of these parameter sets is referred to as an XPS (X is used as a wildcard). The SPF location in a tile segment (e.g., offset in bits) is identified using syntax elements available at the end of the tile segment header.

[0133] SPFs and tile segments can be used jointly, but with some restrictions. One or both of the following conditions must be verified:

[0134] - all coding blocks of one slice (or tile segment) belong to the same SPF; or

[0135] - all coding blocks of one SPF belong to the same slice (or tile segment).

[0136] This means that one slice (or tile segment) can contain multiple complete SPFs, or just a sub-portion of a single tile. Secondly, an SPF can contain multiple complete slices (or tile segments), or just a sub-portion of a single slice (or tile segment).

[0137] Figure 4A quadtree inference mechanism for coding units crossing the boundaries of a picture in HEVC is schematically illustrated, for illustration purposes only. In HEVC, a picture is not limited to be a multiple of the width and height of the coding unit size. Then, the rightmost coding unit of a frame can cross the right boundary 401 of the picture and the lowermost coding unit of a frame can cross the lower boundary 402 of the picture. In these cases, HEVC defines a quadtree inference mechanism for coding units crossing the boundaries. The mechanism consists in recursively splitting any CU of a coding unit crossing a picture boundary until there are no more CUs crossing the boundary or until a maximum quadtree depth is reached for these coding units. For example, coding unit 403 is not automatically split, while coding units 404, 405 and 406 are automatically split. There is no representation of the inferred quadtree: the decoder has to infer the same quadtree on the picture boundary. However, for coding units inside the frame, such as 407, the automatically obtained quadtree can be further refined by representing the splitting information for these coding units (without reaching the maximum quadtree depth).

[0138] Figure 6 A typical video bitstream 600 is shown, sent from a server to a client. The bitstream is compliant with HEVC or block-based bitstream.

[0139] The bitstream 600 is organized as a series of Network Abstraction Layer (NAL) units. There are multiple kinds (types) of NAL units. Parameter set NAL units (e.g. VPS, SPS and PPS used by HEVC) describe the structure of the coding tools used to encode the sequence. These parameter set NAL units also describe some information about the characteristics of the pictures (resolution, frame rate, etc.).

[0140] The first NAL unit 601 is a video parameter set (VPS) providing information about the whole bitstream. In particular, the first NAL unit 601 indicates the number of scalable layers in the bitstream.

[0141] The next NAL unit 602 is a sequence parameter set (SPS). This NAL unit 602 provides sequence level parameters. It is followed by a picture parameter set (PPS) NAL unit 603 providing picture level parameters. Then a slice segment 604 can be provided. There is usually one slice segment per frame. The slice segment 604 can contain references to NAL units of various NAL unit types (CRA, IDR, BLA, RASL, RADL, STSA, TSA or TRAIL...). The NAL units containing the slice segment are indicated by NAL headers 605 (which will be described later) and the slice segment 604 is indicated by a slice segment header 606. The slice segment 604 contains the slice data 607. Figure 10The NAL unit 600 consists of a NAL header 605 (further description of the NAL header is provided in the description of the figure) and a raw byte sequence payload (RBSP) 606. The NAL header 605 contains information including the NAL unit type. The RBSP (i.e. the NAL unit data) contains information specific to the NAL unit type. In case of a slice segment, the RBSP contains a slice segment header 607 and thereafter slice segment data 608. The slice segment data is a series of encoded data of the coding tree units 609 in the raster scan order of the slice segment.

[0142] In an embodiment (not shown here), a parameter set called TPS (which is an abbreviation for tile parameter set) can be inserted in the bitstream before the tile segment NAL units. The respective parameters are valid until a new TPS is found. The TPS describes the partitioning structure of the frame.

[0143] In another embodiment, if no TPS is present in the bitstream, it is assumed that there is only one spatial frame part in the bitstream. The spatial frame part has the same size as the video frame and is located at its origin.

[0144] The bitstream can contain independent frame parts or regions of interest (ROI). Figure 7 An example of a region of interest is schematically shown here as a rectangular area within a frame. ROIs are well known in HEVC and block-based codecs.

[0145] Streaming a ROI or an independent frame part means a partitioning strategy. This has an impact on coding efficiency, because the introduction of tile boundaries breaks some of the HEVC prediction mechanisms.

[0146] In Figure 7 In this figure, the frame 700 is partitioned in a 4x4 SPF grid. To access the predefined ROI 701, the NAL units of the respective slice segments that have the SPF 6, 7, 10 and 11 embedded are selected and sent to the client.

[0147] Preferably, in the present invention, one independent slice segment and zero or more dependent slice segments are embedded in an SPF. The advantage is to guarantee access to the ROI in a way that is independent of other parts of the frame that include the ROI.

[0148] Indeed, it is reminded that, for HEVC and more generally for block-based codecs, the HEVC block or similar blocks break all the intra prediction mechanisms at the boundaries of these blocks (except the loop filtering process). Therefore, the spatial prediction mechanisms are not allowed. However, several prediction mechanisms rely on the temporal redundancy of data between frames of a video sequence to improve the compression performance. For example, one block in a HEVC block can be predicted from a predictor block that is partially or totally outside the current HEVC block boundary. Moreover, the predictor block can also be partially or totally outside the frame boundary, as HEVC provides a well-known picture padding mechanism to pad the picture borders to allow the predictor block to be partially or totally outside the reference picture.

[0149] Finally, the predictor block can be located at sub-pixel positions. This means that the reference block pixel values are the result of a sub-pixel interpolation filter that generates sub-pixel values from a range of up to four pixels that are outside the pixel block located at the full pixel coordinates corresponding to the predictor block. As a result, the temporal prediction can introduce coding dependencies between a block within a HEVC block and a set of pixel data located outside the HEVC block boundary.

[0150] A second HEVC mechanism involved in temporal prediction consists in predictively coding the motion vectors using motion vector predictors.

[0151] Finally, HEVC provides a set of loop filters that introduce dependencies between the pixels of consecutive blocks. These loop filters are the de-blocking filter and the SAO filter that remove some artifacts introduced in particular by the quantization of the residual blocks. HEVC provides a flag in the picture parameter set to indicate whether these loop filters are disabled at the block and / or slice boundaries. When disabled, these compression tools do not introduce coding dependencies between blocks.

[0152] To guarantee the decoding of the region of interest (which means to decode the region of interest independently), the solution is to disable some or all of the aforementioned prediction mechanisms.

[0153] This results in a less efficient compression of the resulting bitstream and a higher bit rate. The activation / deactivation of the prediction mechanisms can be adjusted according to the region of interest usage scheme in order to optimize the bit rate of the resulting bitstream.

[0154] Figure 8a and 8b Two different application examples are illustrated for the combination of regions of interest already explained above.

[0155] For example, in a first example, Figure 8aTwo frames 800 and 801 are shown representing two frames from two different video streams consisting of four regions of interest. The first video stream 800 has high quality encoding parameters and the second video stream 801 is low quality, thus a low bit rate version. The client efficiently combines the high quality version of region of interest #3 with the low quality regions of interest of regions 1, 2 and 4. This permits to emphasize the quality of region of interest #3 while maintaining the bit rate relatively low for the other less important regions.

[0156] In a second example, in Figure 8b A set of four video streams (803, 804, 805 and 806) is shown in

[0157] According to embodiments of the application, it is proposed to define spatial frame portions, hereinafter referred to as constrained tiles (shortcut CTile in the following description). A CTile refers to a spatial frame portion belonging to a sequence of frames which is divided into spatial frame portions that can be randomly accessed and decoded completely without decoding errors. The decoding of a CTile can be done independently of its spatial position and / or its neighborhood. In other words, a CTile is independently coded or coded in such a way that a decoder can always decode it without any error.

[0158] The encoding data corresponding to the samples forming a CTile are independently encoded. For example, the data are encoded into encoding units or NAL units forming a slice (or any other portion of a frame having similar characteristics) so that one parser can extract the samples corresponding to a CTile. As a result, two CTiles are encoded into two different sets of encoding units. In order to decode at any spatial position, a CTile cannot be a portion of a slice containing other CTiles. Thus, the encoding data of a slice corresponding to a CTile are independent of the encoding data from other slices.

[0159] The encoding data corresponding to a slice can be further divided into a plurality of slice segment encoding units.

[0160] In embodiments, a CTile is strictly independently decodable, which means that all the data needed to parse the encoding data forming a CTile are contained in said CTile. Moreover, the prediction mechanism uses prediction information computed from the encoding data of the same CTile. For INTER (inter-frame) prediction, the reference block is retrieved from the same CTile in another frame.

[0161] In other embodiments, the encoding constraints can be lifted.

[0162] In a first other embodiment, a border extension mechanism is used for image boundaries. This mechanism will be illustrated in Figure 5 This mechanism is used for CTile boundaries to allow unrestricted and more efficient motion compensation.

[0163] Figure 5 The border extension mechanism used for example in HEVC is illustrated in a simplified way. This mechanism allows motion compensation in INTER prediction (well-known prediction mode which allows using data outside the current frame) with reference to sample values outside the frame.

[0164] When predicting block 501 while encoding frame 502, it is useful to allow prediction from block 503 which crosses the boundary of reference frame 504. This allows for example to predict motion content from the same content which is partially outside the field of view of the previous frame. Preferably, a sample padding method is defined to allow accessing samples within the blank around the frame boundary of the reference picture.

[0165] In a second other embodiment, for CTile, a derivation mechanism of the motion vector prediction result (or any other prediction result) can be authorized in a way which is completely independent of any neighboring block information and block structure.

[0166] Figure 9 An example of encoding process implemented in a video encoder according to the application is illustrated.

[0167] First, for each input video sequence 900 under consideration, the encoder determines the partitioning of the frame into frame portions (i.e. frame portion configuration) in step 901. In some embodiments, the size of the frame portions is predetermined so that one frame portion covers a single region of interest or a portion of a region of interest. For example, the frame portions can have a size of 512x512 pixels.

[0168] Then, in step 902, the encoder determines which frame portions have to be encoded as CTiles. Such frame portions can correspond for example to regions of interest (ROIs) which the client can want to decode individually or which the client can want to constitute with one or more other regions of interest.

[0169] Then, in step 903, the encoder determines and assigns an identifier to each CTile. In a variant, the encoder can determine and assign an identifier to only a selection of CTiles. In a variant, the identifier of a CTile can be inferred.

[0170] When the same identifier is assigned to CTiles in multiple coded frames, it means that these CTiles belong to the same CTile sequence. A sequence of CTiles (or CTile sequence) can be decoded independently of other frame parts. Only data from the CTile sequence is needed to decode said CTile sequence. In other words, CTiles from a CTile sequence can have temporal dependencies.

[0171] After the CTiles and CTile identifiers are determined, the encoder compresses and encodes the frame parts 904 according to the encoding structure. The encoding of the frame parts ensures that any decoder can decode these frame parts as described before.

[0172] In step 905, the encoder generates frame part configuration information. The frame part configuration information consists in determining the description parameters of the frame partitioning into frame parts. The frame part configuration information also consists in determining the description parameters of the CTiles by associating the CTile identifiers with the position of the CTile identifiers in the frame or frame sequence. As described with reference to Figure 13a 、 Figure 13b and Figure 13c (first alternative) or in Figure 15 (second alternative), different representation alternatives are proposed. Step 905 comprises generating representation tiles parameters in a parameter set (XPS). In variants, step 905 is performed before encoding step 904 instead of after encoding step 904.

[0173] Step 906 comprises optionally encapsulating the NAL units of the XPS and the compressed CTile frame parts into a bitstream.

[0174] This step can also comprise, for example, encapsulating the bitstream inside a higher level video description format such as the ISO BMF file format, based on a streaming protocol. This step can also allow, for example, multiplexing video data with audio data.

[0175] Steps 901, 902 and 903 can be implemented by using one or more configuration files providing pre-determined frame part positions, thereby providing information on whether a frame part is a CTile and which identifiers must be used for the CTiles. In alternative embodiments, the frame parts and CTiles can be determined automatically by analysis of the video content, for example using a deep neural network or a simpler segmentation algorithm.

[0176] As mentioned in some embodiments, step 901 can be used to determine a tile that is constant throughout the video sequence (or at least for a plurality of consecutive frames within a video segment). This means that the position and size of the CTile is constant throughout the CTile sequence, at least within the plurality of frames containing a portion of the CTile sequence.

[0177] Alternatively, the determined frame portion can have a variable size and position between frames.

[0178] Figure 10 An example of a decoding process implemented in a video decoder is shown. The decoding process involves the use of a CTile as defined above.

[0179] First, the video decoder extracts the NAL unit containing the parameter set (XPS). In step 1000, the frame portion configuration information is obtained from the parameter set.

[0180] For each frame portion 1001 under consideration, the decoder determines in step 1002, from the frame portion configuration information, whether the frame portion is a CTile.

[0181] If the frame portion is represented as a CTile (branch "yes" after test 1003), the decoder extracts (or infers) the CTile identifier from the frame portion configuration information, and determines in step 1005 the decoding position of the CTile thanks to the CTile identifier and the CTile position information associated with this identifier. Otherwise, when the frame portion is not a CTile (branch "no" after test 1003), the decoder determines in step 1005 the decoding position of the frame portion from the positioning information described in the frame portion coded data and from the XPS information.

[0182] Finally, the decoder decodes the frame portion coded data 1006 taking into account whether the frame portion is a CTile or not, and puts the decoded sample values inside the rendering picture buffer.

[0183] Figure 11 An example of a merging process of two bitstreams generated by the encoding process described above (see Figure 9 and in the applications of Figure 8a and Figure 8b ) is described. The merging process means combining the extracted CTiles into a new video bitstream to send to a client.

[0184] The merging process starts with determining in step 1100 a set of CTiles to extract from one or more video bitstreams and merge into a new bitstream. For example, a graphical user interface allows a user to select the set of CTiles and to rearrange the set of CTiles in the frame. In another example, the selection is made automatically based on the content of the bitstreams. An application can select a set of CTiles containing moving content.

[0185] The process determines in step 1101 the new positions of the CTiles when merged into the new video bitstream.

[0186] Once the CTiles to extract are known, the new identifiers of these CTiles are determined in step 1102 by obtaining the current CTile identifiers of each of the determined CTiles to extract. According to embodiments of the application, the identifiers are represented in the frame configuration information. As mentioned previously, in alternative embodiments, the frame configuration information can be described using a file format for encapsulating the input bitstreams. The frame configuration information can be present in the form of an XPS and a file format.

[0187] In case of identifier collision, meaning more than one CTile has the same identifier, step 1101 further comprises determining new CTile identifiers to resolve the collisions.

[0188] The process then generates 1103 the frame part configuration information of the merged video sequence of the new video bitstream. The process comprises generating the parameters in one of the XPSs that associate the new positions of the CTiles in the merged bitstream with the new CTile identifiers of these CTiles.

[0189] In step 1104, the encoded frame part data of the set of CTiles determined in step 1100 is extracted or obtained. The step comprises retrieving the NAL units containing the encoded frame part data of the CTiles. This can be done by parsing all the NAL units in the input bitstreams to extract the NAL units having the CTile identifiers determined in step 1102. When the input bitstreams comply with a file format specification, all the NAL units corresponding to one frame part are encapsulated in one container (e.g. a video track of the ISOBMFF). Then, step 1104 comprises retrieving the data corresponding to the track of the selected frame part.

[0190] Finally, in optional step 1105, the new bitstream is generated by embedding the NAL units of the XPS and the NAL units containing the extracted CTile encoded frame part data into the new bitstream, and possibly encapsulating the bitstream into a higher level description format.

[0191] For the CTile whose new CTile identifier is determined in step 1101 due to a CTile identifier conflict, step 1105 further includes modifying the headers included in the NAL unit that include the original CTile identifier, and modifying these headers so that the original CTile identifier is replaced with the CTile identifier determined in step 1102.

[0192] In one example, Figure 11 The merging process consists in extracting a subset of CTiles from the same bitstream. In this case, there is no need to handle identifier conflicts.

[0193] Figure 13a 、 13b 13c show examples of representations of frame portion configurations performed by relevant encoding processes according to various embodiments of the present invention.

[0194] Figure 13a FIG. 4 shows the identification of CTile in a bitstream according to an embodiment of the present invention.

[0195] The CTile identifier, here named ctile_unique_identifier 1301, is indicated in the frame portion coded data. Preferably, this identifier is indicated in each data sequence (i.e., slice segment header) belonging to the frame portion coded data. Thus, this identifier allows:

[0196] - easily identify which parts of the bitstream belong to CTile, and

[0197] - Quickly access or extract these sections.

[0198] More precisely, in Figure 13a In the illustrated embodiment, the CTile identifier 1301 is indicated in a slice segment header (slice_segment_header) 1302 of each slice segment corresponding to the CTile having the identifier 1301 .

[0199] As previously described, the decoder parses the CTile identifier based on the frame portion configuration information to determine the associated position of the CTile. Figure 13b or Figure 13c As described above, frame portion configuration information is provided in a parameter set (eg, TPS).

[0200] For simplicity, in the following, unless explicitly mentioned or not applicable, no distinction will be made between CTile and CTile sequence. In addition, CTile identifiers can also be regarded as CTile sequence identifiers.

[0201] In an embodiment, to distinguish HEVC-type tiles, which do not necessarily need an identifier, from CTiles, information such as a ctile_flag 1303 can be used in the data sequence belonging to the frame part coded data, for example in the slice segment header. If the ctile_flag is inactive (for example set to "false"), parameters 1304 of the HEVC-type tile are provided. These parameters can include tile positioning information such as first_slice_in_pic_flag or CTU address (slice_segment_adress), or references to other bitstream elements such as slice_pic_parameter_set_id. These syntax elements depend on the frame partitioning and can vary from video sequence to video sequence.

[0202] When the ctile_flag is active, these parameters are omitted and instead CTile-specific information including a CTile unique identifier 1301 is provided. To allow the possibility of having multiple slices in a CTile, one solution is to provide information here named ctb_addr_offset_inside_tile 1305. This information 1305 is also used to specify the position of the start of the decoding of the slice segment relative to the CTile position of the frame under consideration. For example, this position is expressed in the original scan order numbering of the coding blocks (for example, CTBs which are the coding tree blocks of the HEVC standard) relative to the start of the CTile and its (in CTBs) width, so the ctb_addr_offset_inside_tile information is independent of the CTile coding / decoding position.

[0203] In another embodiment, the flag ctile_flag is not used. For example, the CTile identifier is present in all tiles (CTiles and other tiles (HEVC-type tiles)). A predetermined value (for example the value zero) can be used to identify the HEVC-type tiles.

[0204] In an embodiment, information is provided to identify whether a spatial frame part is a CTile.

[0205] In another embodiment, assuming that only CTiles are used as frame parts, information to identify whether a spatial frame part is a CTile is not provided.

[0206] Preferably, in a given frame, there is no more than one CTile with a given identifier. The same CTile identifier is used in all temporally dependent CTiles (for example, in a CTile sequence). Thus, if CTiles with the same CTile identifier in consecutive coded pictures are extracted, these CTiles will be correctly decoded.

[0207] In other words, the CTile identifier is a unique identifier that identifies a CTile within an encoded video sequence. In an embodiment, the CTile identifier is inserted in the slice header of the slice segment contained in the CTile. This means that, in the bitstream, the NAL unit (slice segment) corresponding to a CTile contains the CTile identifier. Therefore, any CTile can be easily parsed and extracted from the bitstream based on this CTile identifier.

[0208] It is advantageous to represent the CTile configuration information in the bitstream. For example, the CTile configuration information is defined by the number of CTiles, the associated CTile identifiers and the position of the CTiles in the frame.

[0209] Figure 13b A CTile configuration according to an embodiment of the application is illustrated.

[0210] In a first embodiment, the encoder specifies additional representation information related to the frame portion configuration information of the CTiles in the picture to be decoded. This representation information is provided in a parameter set (XPS), preferably in a tiling parameter set (TPS). Preferably, the additional representation information comprises the number of CTiles in the picture, here referred to as num_ctiles 1311. For each CTile, this additional representation information associates a unique identifier of the CTile 1312 and a CTile position 1313, here referred to as tile_ctb_addr, which means a decoding position inside the picture. The CTile position is provided in the picture as a decoding position. The CTile position can be expressed in CTB index number (e.g. with respect to the raster scan order).

[0211] In another embodiment, in the slice segment header, the parameter referred to in the part specified by 1304, here named slice_pic_parameter_set_id, refers to a unique identifier of a TPS. In a variant, the unique identifier refers to a PPS. In this other embodiment: Figure 13a The parameter referred to in the part specified by 1304 in the slice segment header, here named slice_pic_parameter_set_id, refers to a unique identifier of a TPS. In a variant, the unique identifier refers to a PPS. In this other embodiment:

[0212] - each TPS comprises a tile_parameter_set_id (not shown for simplicity) parameter for identifying the TPS (e.g. each time the CTile configuration changes in the picture, which means that the encoder can generate a new TPS), it is recommended to generate TPSs with the same TPS (or PPS) unique identifier to avoid rewriting each slice header of the frame;

[0213] - The slice_pic_parameter_set_id 1314 of the slice segment header is equal to the tile_parameter_set_id of the TPS applied to the slice. In this case, the naming of the slice_pic_parameter_set_id of the slice segment header can be renamed as slice_tile_parameter_set_id.

[0214] In one alternative, the TPS identifier is not specified in the slice data: the decoder infers that the last TPS NAL unit before the slice NAL unit contains the frame portion configuration of the current CTile.

[0215] Figure 13c A CTile configuration according to another embodiment of the application is illustrated. According to this embodiment, the TPS 1320 contains a parameter value indicating the number of tiles minus 1, for example "num_tiles_minus1" 1321. Alternatively, the TPS contains a parameter value, for example named "num_tiles", providing directly or indirectly the number of tiles in the frame.

[0216] In an embodiment, if the TPS indicates that there is only one frame portion, it is assumed that this frame portion is a CTile having the same size as the video frame and located at its origin. Otherwise (the TPS describes multiple frame portions), the frame portion locations are described as in the previous embodiment.

[0217] In another embodiment, if there is no TPS, it is assumed that there is one CTile having the same size as the video frame. Said one CTile is located at its origin.

[0218] According to another embodiment, the TPS can describe a spatial frame portion grid with a syntax similar to the HEVC grid: for example, "num_tile_rows_minus1", "num_tile_cols_minus1" and "uniform_spacing_flag" are specified. If "uniform_spacing_flag" is not set, the width of each column and the height of each row are also specified (except the size of the last row and the last column which can be inferred). If "uniform_spacing_flag" is set, the width and height of the CTiles are computed according to the picture width and height, for example as in the HEVC specification. In such an embodiment, since the grid index allows to locate the corresponding CTile, the CTile position can be expressed with the CTile number corresponding to the spatial frame portion grid index, for example using the raster scan order of the tiles.

[0219] According to an alternative embodiment, "ctile_flag" is replaced by "ctile_level" which can take several values, each value indicating a different level of coding constraints applied to CTile. For example, ctile_level equal to zero, which means that CTile is not constrained (as in HEVC-like CTile). ctile_level equal to "1", which means that CTile is constrained such that it can be extracted and properly decoded alone (without the original neighborhood of this CTile), or it can be decoded with the original neighborhood of this CTile, but in case of shuffling of CTile with other CTile, it can not be properly decoded. ctile_level equal to "2", which means that CTile is constrained such that it can be decoded anywhere and CTile can be shuffled with any neighborhood (equivalent to the previous embodiment where ctile_flag equal to 1).

[0220] In another embodiment, "ctile_level" only provides information to the encoder to make its coding decision to meet the level of constraints. Thus, the decoding process of CTile with any level of constraints can be achieved by the same decoding process as HEVC-like block (e.g. no padding at CTile boundary).

[0221] In another embodiment, the encoding and decoding process are not the same for all levels of constraints. For example, CTile with ctile_level equal to "1" uses the same decoding process as HEVC-like CTile (some restrictions used at the encoder do not impact the decoder), while CTile with ctile_level equal to "2" must be decoded using padding at CTile boundary and using a specific derivation process of the list of motion vector predictors.

[0222] According to another embodiment, even HEVC-like blocks can need to have an identifier (e.g. to associate parameters of these blocks in XPS), then this identifier is specified in the slice segment header in a similar way as CTile identifier. In a given frame, HEVC-like blocks do not have the same identifier as CTile or another HEVC-like block.

[0223] According to one embodiment, the encoder indicates that a spatial frame portion is a CTile by indicating the value of a "ctile_flag" for the spatial frame portion in one of the parameter sets (e.g., PPS or TPS). For example, the encoder generates a unique identifier for each tile of a frame. When describing the frame portion configuration, the encoder associates a flag (e.g., ctile_flag) with each tile unique identifier. The flag is true when the encoding of the corresponding tile (i.e., the tile whose identifier is equal to the associated unique identifier) is constrained to ensure independent decoding. Conversely, the flag is false when the encoding of the tile is not constrained enough to guarantee independent decoding.

[0224] According to a second embodiment, the encoder generates frame portion configuration information that includes another flag (e.g., all_ctile_flag). If this flag is set to "1", it means that all tiles described in the frame portion configuration are CTiles. The flags (e.g., each ctile_flag) indicating whether a spatial frame portion is a CTile are omitted and inferred to be equal to true. If this flag is set to zero, one of the previous embodiments is used to explicitly describe the CTiles. If the parameters are specific to HEVC-type tiles, these parameters are indicated in the XPS (e.g., in the TPS) instead of in the slice segment header. For example, slice_segment_address incorporated by reference in 1304 for another embodiment is specific to HEVC-type tiles. In an embodiment, the TPS indicates whether the TPS also indicates that a spatial frame portion is not a CTile. This embodiment allows to simplify the syntax and parsing of the slice segment header.

[0225] According to another embodiment, the encoder defines new NAL unit types for the slice data corresponding to a CTile instead of using "ctile_flag" in the slice segment header. For example, the encoder defines a CTILE IDR NAL unit for slice NAL units from an instantaneous decoding refresh (IDR) frame that is inside a CTile. The encoder defines as many new NAL unit types as the encoding format specifies for the NAL unit types for regular slice data. For example, HEVC defines the following NAL unit types for slice segments of clean random access (CRA) pictures: CRA NUT; for slice segments of random access decodable leading (RADL) IDR pictures: IDR W RADL; for slice segments of IDR pictures that do not have an associated leading picture in the bitstream: IDR N LP; for slice segments of broken link access (BLA) pictures: BLA W LP, BLA W RADL, BLA N LP; for slice segments of random access skipped leading (RASL) pictures: RASL N, RASL R; for slice segments of RADL pictures: RADL N, RADL R; for slice segments of step-wise temporal sub-layer access (STSA) pictures: STSA N, STSA R; for slice segments of temporal sub-layer access (TSA) pictures: TSA N, TSA R; and for slice segments of non-TSA, non-STSA trailing pictures: TRAIL N, TRAIL R.

[0226] W LP: can have an associated RASL or RADL picture; W RADL: no associated RASL picture; N LP: no associated leading picture; * N: picture is a sub-layer non-reference (SLNR) picture (otherwise picture is a sub-layer reference picture); * R: picture is only a sub-layer reference picture.

[0227] These HEVC NAL unit types can be extended with new corresponding NAL unit types CTILE BLA *, CTILE CRA *, CTILE IDR *, CTILE RASL *, CTILE RADL *, CTILE STSA *, CTILE TSA *, CTILE TRAIL * with the same purpose as for the constrained tile data. One of these new NAL unit types is used to indicate that the NAL unit belongs to a CTile.

[0228] This alternative simplifies the decoding process since the encoder only has to parse the beginning bits of each NAL unit to determine whether the slice data is inside a CTile.

[0229] Preferably, the uniqueness of the CTile identifier is guaranteed by construction at encoding time for a given sequence (meaning in a given bitstream). However, when shuffling CTiles from different sequences (meaning from different bitstreams), the uniqueness cannot be guaranteed. According to an embodiment, to facilitate the shuffling of spatial frame portions with CTiles possibly coming from various sequences, the CTile identifier is unique on a limited number of bits. This unique value can be a random value, for example a hash value or any other value that does not necessarily represent its position. Thus, the probability of having a collision of identifiers when taking CTiles of different bitstreams is reduced.

[0230] In an embodiment, when shuffling CTiles from multiple sequences, it is sufficient to replace the colliding CTile identifier in case of collision between two CTile identifiers. To implement this efficiently without having to regenerate all the slice segment headers, in a preferred embodiment, a fixed predetermined number of bits is used to encode the CTile identifier. For example, in Figure 13a and 13b the CTile identifier is encoded on 8 bits.

[0231] In an alternative embodiment, all the identifiers of a sequence or picture are encoded on the same number of bits. This number of bits is specified in a parameter of one parameter set (such as SPS, PPS or RPS, etc.), for example "uid_num_bits". In the slice segment header, there is preferably a byte alignment mechanism after the CTile identifier (when a number of bits that is not a multiple of 8 is needed). Alternatively, the number of bits can be expressed in number of bytes (8 bits): for example "uid_num_bytes". When shuffling CTiles from various sequences together, it can be necessary to change the CTile identifier when it is not all of the same number of bits. This would require changing multiple slice segment headers, but is easier than updating slice segment headers as only bytes would need to be added / removed or replaced.

[0232] In yet another alternative embodiment, a variable number of bits can be used to encode each CTile identifier. This number of bits is specified in the slice segment header. Alternatively, the number of bits can be determined automatically according to the code used for the CTile identifier: a variable byte length code is used, for example exponential Golomb coding (or equivalent variable length code followed by a byte alignment bit).

[0233] According to an embodiment, the CTile identifier is not represented in the dependent slice segment header, to reduce the representation size. The CTile identifier of the dependent slice segment header is then inferred from the previous independent slice segment header. According to an alternative embodiment, the CTile identifier is represented in the dependent slice header, to facilitate parsing and extraction of the sub-bitstream containing the CTile.

[0234] As an alternative to representing the CTile position with the coding unit address coded with "tile_ctb_addr[i]" 1313 or "slice_segment_address" 1306, fine-granularity CTiles with a finer granularity positioning are introduced. This granularity can be refined to the luma sample position, but in another embodiment, a granularity of a number of luma samples corresponding to a power of "2" (smaller than the CTU size) is sufficient. In some embodiments, the granularity can be predetermined. In alternative embodiments, the granularity is represented for example in the VPS, SPS or PPS. When using fine-granularity CTiles, the size of the CTile is not necessarily a multiple of the CTU size.

[0235] When the size of the CTile is not a multiple of the CTU size, as illustrated with Figure 4 The right and bottom edges of the CTile are using the same auto-split mechanism as the HEVC CTU of the right and bottom edges of the picture.

[0236] According to an alternative embodiment, the syntax will describe a complete coding even if the coding unit is not complete, leaving some room for rate-distortion optimization of the decomposition tree (e.g. quad-tree or QTBT) and allowing to eventually fill in information suitable to improve compression.

[0237] In HEVC-type tiles, the size is specified using a grid. Therefore, all HEVC-type tiles are aligned by rows and columns, and all HEVC-type tiles of a given row have the same height and all HEVC-type tiles of a given column have the same width. The width of each column and the height of each row are specified in the XPS. With fine-granularity HEVC-type tiles, it can be convenient to allow less strict configurations (e.g. to allow more efficient coding of multiple ROIs).

[0238] According to an embodiment, the size of the CTile can be specified in the slice segment header of the slice segment of the CTile. To reduce the size of the bitstream thus obtained, the size of the CTile is only specified in the first slice segment. The following slice segments reuse the same CTile size. As an alternative, the size of all CTiles is set in the XPS, for example together with the CTile positions.

[0239] As another alternative, the size is provided in both the first slice segment header of a CTile and in the XPS.

[0240] As another alternative, the size of a CTile is not provided, but is inferred from the ordering used to provide tile information (e.g., position or "ctile_flag") in the XPT and from the CTile position: e.g., the position of a CTile is declared in the XPS and the CTile positions are ordered such that the respective bottom-right corners of the CTiles are ordered in, e.g., a raster-scan order increasing. In the following, Figure 14 An example of providing such an ordering.

[0241] According to embodiments, in Figure 13a The dependent_slice_segment_enabled_flag used in the example of Fig. 6 has the same meaning as in HEVC: it is used to indicate whether dependent slice segments are allowed or not. In HEVC, dependent_slice_segment_enabled_flag is signaled in the PPS. According to a preferred embodiment, dependent_slice_segment_enabled_flag is signaled in the tiling parameter set (TPS) for each CTile (to allow CTiles coded with dependent slice segments and CTiles not coded with dependent slice segments to be used in the same bitstream). To reduce the syntax in the TPS for the common use case where all CTiles are coded either with or without dependent slice segments, another flag is used at the root of the TPS structure: dependent_slice_segment_enabled_flag_for_all_ctiles. When this flag is set to 1, dependent_slice_segment_enabled_flag is not signaled for each CTile. Instead, ctile_dependent_slice_segment_enabled_flag is also signaled at the root of the TPS structure and provides the value to be inferred for dependent_slice_segment_enabled_flag for each CTile. For HEVC-type tiles, dependent_slice_segment_enabled_flag can still be signaled in the PPS, but in the preferred embodiment, dependent_slice_segment_enabled_flag is signaled in the TPS.

[0242] According to an alternative embodiment, dependent_slice_segments_enabled_flag is not signaled at all and is always inferred to be true to simplify the syntax.

[0243] Figure 14 An example of non-grid based partitioning is shown. The frame 1401 is split into 15 CTiles numbered #1 to #15. This numbering provides an order to declare tile positions in the XPS such that the lower right corner of each tile is ordered in raster scan order. Using this ordering, the size of each CTile can be derived. For example, with the last tile (CTile #15), the size of this tile can be derived since as the last tile, its lower right corner is the last in raster order and thus the lower right corner of the frame. The size of slice #15 is then the size of the frame - the position of the slice: h#15 = h_frame - y#15; w#15 = w_frame - x#15. Tile #14 must have the last lower right corner before CTile #15 as its own lower right corner, so the lower right corner position is the lower most (below the frame) and right most (just to the left of the previous tile). The size of CTile #14 is then h#14 = h_frame - y#14, w#14 = x#15 - x#14. The same operation is repeated for tiles #13 and #12. Then, for tile #11, the new lower most position is y#14 since the lower most position is padded. Then, h#11 = y#14 - x#11, and so on until CTile #1.

[0244] According to an alternative embodiment, instead of specifying the CTile positions in the XPS, only the CTile sizes are specified and the CTile positions are computed from the CTile sizes using the CTile ordered according to the top left position (e.g. in increasing raster scan order). The algorithm to compute the positions from the sizes can be easily derived from the algorithm to compute the sizes from the positions described above.

[0245] According to an embodiment, the CTile parameters described in the XPS can provide the CTile positions (and / or CTile sizes) for non-existing CTile: there will be no slice segment for these CTile. In the embodiment where only the positions or sizes are provided and the sizes or positions are inferred, this CTile description is needed to allow proper inference.

[0246] For video rendering, the non-existing CTile is filled using a default sample value or padding method, or alternatively, the value or index of the padding method is set in the XPS parameters. This can be done, for example, by setting a default value for the padding method in the XPS parameters at the beginning of the decoding process (e.g. in the slice header) and overriding this value in the slice segment header. The padding method can be signaled in the slice segment header or in the slice segment data. Figure 9- a preliminary step is implemented before step 900 in the method, the preliminary step consisting in:

[0247] - the content of the frame in the rendering buffer is initialized with appropriate default sample values, and / or

[0248] - a new step is added after all frame parts have been decoded, the step consisting in filling all areas not covered by any tile or CTile, for example by using an inpainting method.

[0249] According to an embodiment, multiple CTiles at the same spatial position or overlapping CTiles can be handled. For each CTile identifier, there is an associated decoded CTile buffer (equivalent to the decoded picture buffer (DPB) in HEVC, but here containing only decoded CTile data). For a given frame, each CTile is decoded using the temporal data available in the associated decoded CTile buffer. Then, according to a first alternative, the rendering order of the CTiles is the same order as the CTile order in the bitstream. In another alternative, the CTiles are associated with a rendering order that can be determined from the XPS data. For both alternatives, the samples of the decoded result of each CTile are put in the frame of the rendering frame buffer in the rendering order of the CTiles (then, possibly erasing / masking the samples put previously by the preceding CTiles in sequence).

[0250] According to an embodiment, the CTile samples further comprise an alpha channel indicating the level of transparency that should be applied when rendering the CTile in the frame of the rendering frame buffer. Alternatively, the samples further comprise a binary mask value indicating which samples of the CTile must be rendered in the frame of the rendering frame buffer.

[0251] According to an embodiment, in case multiple CTiles at the same position or overlapping CTiles can be handled, both the CTile position and the CTile size must be specified in the XPS, as in this case one cannot derive the other from one of them.

[0252] According to embodiments, for any given post-filtering algorithm (e.g. deblocking filter, sample adaptive offset or adaptive loop filter), a CTile boundary post-filtering flag can be specified in the XPS to indicate whether the post-filtering algorithm can be used for CTiles. The CTile boundary post-filtering flag (e.g. "usable_for_post_filtering_flag") indicates that the given post-filtering algorithm can be applied to the CTile boundaries in the rendered frame of the rendered frame buffer (rather than the decoded picture buffer, as the decoded picture buffer can modify the temporal decoding). Advantageously, the purpose of the post-filtering algorithm is for example to improve the visual quality. The flag can be specified for the whole frame level and / or for each CTile. This flag can be used to prevent filtering of some edges that are known to be prone to introduce artifacts when post-filtered. For example, for CTile shuffle in the case of adaptive quality streaming, the flag would be true, but in the case of a CTile boundary between two faces of a cubemap projection of 360° content and these faces are not contiguous on this edge, the flag would be false. The post-filtered CTile border is the border of the CTile on both sides of the edge for which it is specified that post-filtering can be applied, or when the edge is between a HEVC-like tile and a CTile for which post-filtering is authorized.

[0253] According to alternative embodiments, in this embodiment, the CTile boundaries are post-filtered in the decoded picture buffer (DPB) to ensure that the decoding is correct in any decoding structure and that the samples used for prediction when using INTER prediction are unpost-filtered samples. Thus, the border extension mechanism is applied to the last samples on the border before the post-filtered samples using the border information, which means that the unfiltered samples are border extended.

[0254] According to alternative embodiments, more than one CTile can have the same CTile identifier. In this embodiment, the CTile identifier becomes a CTile set identifier. The set of CTiles forming a CTile set must stay together and have the same relative positioning to be decoded properly.

[0255] In these embodiments, the position and size of the CTile set are inferred from the XPS. The position and size of the CTile set correspond to the position and size of the bounding box of the set of CTiles belonging to the CTile set. Thus, in the XPS, the CTile set identifier is associated with one or more positions and sizes (one position and size exists for each CTile in the CTile set).

[0256] In these embodiments, the slice segment header "ctb_addr_offset_inside_tile" 1305 information may be replaced by "ctb_addr_offset_inside_tile_set" information. "ctb_addr_offset_inside_tile_set" allows to deduce which CTile belongs to the slice segment and therefore the geometry to use when decoding the slice segment.

[0257] In one such embodiment, temporal motion compensation can be performed using any sample from a set of CTiles. If motion compensation uses a sample value outside the set of CTiles, the sample value is set to the value of the spatially closest sample from any CTile in the set (equivalent to applying border extension, but only for CTile boundaries not shared by two CTiles). If any sample outside a CTile has more than one closest CTile sample, a simple rule is used to determine which sample to use (e.g., the sample with the smallest raster scan order).

[0258] Figure 15 Show Figure 13a 、 13b and 13c for representing an alternative embodiment of the CTile identifier. In current block-based codecs (typically HEVC), the NAL unit header 1501 contains the following fields:

[0259] - a bit set to 0: (False);

[0260] - Six bits containing the NAL unit type: (Type);

[0261] - six bits containing a layer identifier: (LayerID), which is always equal to zero in HEVC, but corresponds to a scalable layer index in scalable HEVC (SHVC), or a view index in, for example, multi-view HEVC (MV-HEVC); and

[0262] - Three bits representing the temporal layer identifier: (TID), which corresponds to the temporal layer index of temporal scalability in HEVC.

[0263] In an embodiment, based on the NAL unit header 1501, the encoder splits the video sequence into frame portions. The encoder uses one encoding or scalable layer per frame portion. This can be seen as a layer encoding based on spatial regions. The encoder can encode each spatial region layer independently from the other regions. In this case, each spatial region layer corresponds to one CTile. In this particular case, the ctile_flag is set to true for all slices of a spatial region layer at encoding time. The main difference is that each spatial region layer can be further divided into HEVC type tiles.

[0264] The encoder denotes the different spatial region layers with Layerld. The encoder sets the value of Layerld equal to the identifier of the CTile. As a result, the CTile identifier is not needed in the slice segment header. Since the CTile identifier has a fixed bit length, the handling of the CTile identifier remains simple when shuffling the frame portions of the video stream.

[0265] The encoder denotes the frame portion configuration in one of the parameter sets, e.g. the VPS. The VPS indicates the decoding position of each spatial region layer by associating the unique identifier of the spatial region layer with the decoding position using a syntax that can correspond to the syntax described in the previous embodiment.

[0266] The encoder also describes the dependencies between the different layers of the video stream. Then, the decoder determines the spatial region layers that are encoded independently from the other layers by analyzing the dependencies between the layers described in the parameter set NAL units.

[0267] The encoder compresses a subset of the spatial region layers as CTiles independently from the other spatial region layers. When one spatial region layer depends on another spatial region layer in a previous frame (i.e. a frame with the same CTile identifier) (the ctile_flag of the slices in this layer is set to false), the encoder adds the reference frame from this dependent layer to the decoded picture buffer of the current layer. When the size of the two layers is different, an up-sampling or down-sampling filter is applied so that the size of the reference frame is equal to the size of the current layer.

[0268] According to an embodiment, it is also possible to use Layerld to infer the ctile_flag: when Layerld is zero, the NAL unit belongs to a HEVC type tile. When Layerld is not zero, the NAL unit belongs to the CTile whose CTile identifier is equal to Layerld. Optionally, one bit of Layerld is reserved to represent the ctile_flag. The advantage of using Layerld to transmit the CTile identifier is that it greatly reduces the complexity of parsing the bitstream to extract the CTile identifier.

[0269] In another embodiment, spatial region scalability is defined similarly to temporal scalability in HEVC, i.e. different layer identifiers identify temporal, spatial regions from other scalable layers (e.g. SNR, resolution, multi-view). Indeed, this approach has the advantage that both spatial region scalable layers and SNR or resolution scalable layers can be used.

[0270] The NAL unit header 1502 is extended with a random access identifier (RAID) that now indicates a frame portion identifier. The Layerld semantics remains the same as in HEVC, i.e. the Layerld semantics indicates a multi-view, SNR or resolution scalable layer.

[0271] The encoder specifies the location of each spatial region layer by associating its RAID value with the decoding location in one of the parameter sets (e.g. VPS). The RAID of each NAL unit encoding a spatial region (including SPS, PPS and VCL NAL units) is equal to the frame portion identifier (CTile (set) identifier) corresponding to this spatial region.

[0272] As a result, the above-mentioned merging process (which consists in extracting CTiles from a set of video bitstreams and combining these CTiles into a new video bitstream) extracts the CTile identifiers of the spatial region layers to be merged from the frame portion configuration related to the video streams to be combined. Then, the merging process extracts all the NAL units whose RAID value is equal to the set of extracted identifiers.

[0273] In order to limit the risk of identifier collision when combining two video sequences, the encoder sets the RAID value to a random value. This includes the case where the video sequence contains a single frame portion.

[0274] According to an embodiment, the RAID specifies whether a spatial region is a CTile (the RAID replaces the representation of ctile_flag in the slice segment header): when the RAID is zero, the NAL unit belongs to a HEVC-type tile. When the RAID is not zero, the NAL unit belongs to a CTile whose identifier is equal to the RAID. Alternatively, the RAID of one bit is reserved to represent the ctile_flag. In alternative embodiments, the RAID identifier is 16 bits or 24 bits to allow more CTiles.

[0275] According to an embodiment, a CTile sequence is considered as an independent bitstream. For example, in some embodiments, the sequence order of CTiles having the same identifier can be different from the sequence order of CTiles having another identifier (i.e. the GOP structure can be different between two CTiles). Thus, two CTiles in the same frame can have different NAL unit Types or TIDs.

[0276] In another embodiment, the XPS includes additional information describing some dependencies between CTiles, enabling the decoder to process these CTiles without any error. It has been seen that the independence of CTiles can be considered at the level of a group of CTiles, rather than at the level of each CTile. In this structure, some CTiles within a set of CTiles can have certain dependencies.

[0277] For example, Figure 16a An XPS is shown that includes a list of dependencies for each CTile. The list of dependencies provides a CTile identifier 1601 that depends on a given CTile with identifier 1600. When a given CTile is indicated as having a dependency on another CTile, this means that the given CTile cannot be extracted without the other CTile.

[0278] Figure 16b A first example of CTile dependencies is shown. CTile #1 in the current frame 1602 uses sample values from CTile #2 when performing motion compensation from a previously encoded frame 1603; and CTile #2 uses sample values from CTile #1. In such an example, the XPS indicates that the CTile with identifier #1 has a dependency on the CTile with identifier #2, and the CTile with identifier #2 has a dependency on the CTile with identifier #1. This representation of mutual dependencies is an alternative to represent a set of CTiles.

[0279] Figure 16c A second example of CTile dependencies is shown. In such an example, CTile #3 present in frames 1605 and 1607 has a dependency on CTile #1 and CTile #2 in frames 1604 and 1607. In such an example, CTile #3 cannot be extracted without also extracting CTile #1 and CTile #2. But CTile #1 and CTile #2 have no dependencies and can be extracted individually. According to an embodiment, this scheme can be applied to facilitate extraction of CTiles at various frame rates, for example, in the case of temporal ordering of frames 1604-1606. Alternatively, this scheme can be used for scalable coding, for example, in the case where frame 1605 is a refinement layer of frame 1604, and frame 1607 is a refinement layer of frame 1606, to facilitate extraction of CTiles for layers of different quality.

[0280] According to some embodiments, a CTile can change spatial position or size between consecutive frames.

[0281] According to embodiments, a "Tile Parameter Set" or TPS is also introduced. The TPS allows to update the CTile parameters of a subset of CTiles (only the subset that is moved and / or changed in size), the TPS containing for example a "num_updated_tiles" value, then the TPS associates tile identifiers with the new attributes of the modified CTiles.

[0282] Conventionally, a motion vector gives the position of a prediction result block in a reference image relative to the block at the same position as the block to be encoded. For a given block to be encoded, the first step is to identify the co-located block in the reference image. The co-located block is defined as the block in the reference image having the same position, which means having the same origin (top-left position) and the same size as the block to be encoded. Then, the motion vector is applied to the origin of the co-located block to determine the origin of the prediction result block.

[0283] When considering CTiles, the determination of the co-located block is adapted to take into account blocks having the same position within the CTile but no longer having the same position within the frame when the CTile has been shuffled. When a CTile has been shuffled, this means that the position of the CTile in the frame at decoding has been modified relative to its position within the frame at encoding. However, considering that the prediction is limited within the CTile to ensure independent decoding, it is still possible to determine the correct prediction result block by applying the motion vector to the block co-located with the encoded block within the CTile. This is applicable as long as the CTile keeps its size and position between frames. When the CTile changes its position in the frame and / or its size between two consecutive frames, difficulties arise. In this case, the encoder and the decoder agree on a way to determine the position of the co-located block in the reference frame to which the motion vector is to be applied to correctly determine the prediction result block.

[0284] According to embodiments where a CTile can change its position or size between consecutive encoded frames, the relative position of the CTile in two consecutive frames can not be the same in the two different bitstreams. Figure 17aAn example is provided where the first bitstream contains frames 1700 of a video surveillance. In a first frame 1701, there are multiple spatial frame portions including a CTile 1702 for a moving region of interest with a given ctilejd. In another frame 1703, the CTile 1704 with the given ctilejd has moved and has changed size. The second bitstream contains a generated video 1705 by assembling the CTiles extracted from 1700 with generated CTiles containing a uniform color (e.g., black). In a first frame 1706, the CTile 1702 with the given ctilejd is extracted from the first bitstream and placed in the center of the frame 1707. In another frame 1708, the CTile 1704 with the given ctilejd is extracted from the first bitstream and placed in the center of the frame 1709. In the first bitstream, the CTile 1704 uses INTER prediction with a temporal reference to the CTile 1701. Thus, in the generated video 1705, the CTile 1709 uses INTER prediction with a temporal reference to the CTile 1707. The relative spatial positions between the CTiles 1702 and 1704 are not the same between 1707 and 1709. Thus, to properly decode regardless of the decoded relative positions, the encoded motion vectors do not take into account the CTile position changes (i.e., relative spatial positions between consecutive frames) when using INTER prediction mode.

[0285] According to a first alternative, the motion vectors are computed as if the predetermined reference points of the CTiles in two consecutive encoded frames are at the same spatial position (e.g., top-left, top-right, bottom-left, bottom-right, middle-top, middle-bottom, middle-left, middle-right, or center). Thus, the encoded motion vectors will also correspond to the following: the motion vector for a block in a frame reference minus the motion vector between the reference points of the CTiles in the frame reference: thus obtaining a motion vector relative to the motion vector in the reference of the CTile. The result is that the CTile can then be decoded independently of subsequent spatial changes.

[0286] Figure 17b It is shown that for the CTile 1710 encoded in frame 1711, the reference CTile 1712 with the same ctilejd in the reference frame 1713 is used. The block 1714 is encoded using a motion vector 1715 corresponding to the difference between the motion vector of the motion vector 1716 in the frame 1716 and the motion vector between the predetermined reference point 1717 (which is the top-left corner of the CTile in this example). Figure 17bIt is also shown that even if the relative temporal decoding position of a CTile is not the same as the encoding position, the encoding vector 1715 is still valid when the block 1718 is decoded by adding the motion vector between the predetermined reference points in the decoded frame 1719 to obtain the motion vector in the decoded frame 1720.

[0287] According to a second alternative, the motion vector is computed as if the given point of the CTile is at the same spatial position in the two successive encoded frames. The given point is represented in the CTile encoding data as an index in a list of predetermined points (e.g. top-left, top-right, bottom-left, bottom-right, middle-top, middle-bottom, left-middle, right-middle or center).

[0288] According to a third alternative, a fixed point (or alternatively a representative point) is considered and the motion vector is encoded in the CTile encoding data. When it is considered that the fixed (or representative) point of the CTile in the two successive encoded frames is at the same spatial position, the motion vector encoded in the CTile encoding data provides a motion vector to be added to each of the INTER motion vectors associated with the temporal prediction in the CTile. This allows the encoder to reduce the encoding cost of the motion vectors. For example, the encoder can select the average motion vector of the motion compensated blocks of the CTile. For example, as shown in Figure 17b the average motion vector would be subtracted to the vector 1715. The result of the subtraction is considered as the motion vector to be encoded. An alternative which is equivalent in terms of result is to provide, instead of (or in addition to) the fixed point or the fixed point index and the motion vector, the (sub)pixel position of the reference point in the encoded CTile.

[0289] According to a fourth alternative, a fixed point (or alternatively a representative point) is considered. Parameters of a motion field are encoded in the CTile encoding data. The motion field allows to determine, when it is considered that the fixed (or representative) point of the CTile in the two successive encoded frames is at the same spatial position, a motion vector to be added to each of the INTER motion vectors associated with the temporal prediction in the CTile. For example, the encoder can estimate the motion vectors of the blocks in the CTile and can estimate a motion field which minimizes the prediction of these blocks, thereby minimizing the residuals of these blocks and thereby reducing the encoding cost of these blocks. The motion compensation vectors of the INTER encoded blocks are then the result of subtracting, to the motion compensation vectors (e.g. 1715), the motion vectors computed from the motion field parameters.

[0290] According to an embodiment, the INTER prediction mode can refer to more than one previously encoded reference frame. In this embodiment, the aforementioned embodiments can be extended to take into account that the fixed (or representative) point of the CTile is aligned in the encoded frame and in each reference frame.

[0291] In embodiments where a motion vector or a motion vector field is represented, the extension to multiple frames can be done in two alternative ways:

[0292] - by representing as many motion vectors (or motion vector field parameters) as the number of reference frames, or

[0293] - by representing only one motion vector (or motion vector field) "x" which is used to derive one motion vector (or motion vector field) for each reference frame according to the time difference between this motion vector and the encoded frame.

[0294] For example, using a linear scaling: if the time position of the reference frame is "t-s" (where: "s" is the constant time sampling period between frames, and "t" is the time), and the time position of the encoded frame is "t", the scaling factor used is "(t) / s - (t-s) / s = 1", but if the time position of the reference frame is "t+2s", the scaling factor used is "(t) / s - (t+2s) / s = -2". The scaling factor is applied to compute the motion vector for each reference frame. For example, as shown in Figure 17b frame 1713 is "t+2s", the motion vector 1715 is subtracted by -2*"x". The result of the subtraction "y" is the value of the encoded motion vector (e.g. "y" is the motion vector predicted using the motion vector prediction result index if the encoding mode is the INTER prediction mode of HEVC). In other words, on the decoder side, the motion vector "y" is decoded for the motion compensated block, and then the motion vector "y" is added with -2*"x" to obtain the vector 1715. Further added is the motion vector 1717 to obtain the frame level motion vector.

[0295] Figure 12 Details related to the encoding side encapsulation step 906 or 1105 when the bitstream is being encapsulated into a higher level description format as described in the foregoing description are provided.

[0296] In a preferred embodiment, the video bitstream with CTile is encapsulated according to the ISO Base Media File Format (ISOBMFF, ISO / IEC 14496-12 and 14496-15). In the following description related to Figure 12 The term "sample" corresponds to a "frame" as defined for ISOBMFF, i.e. a set of NAL units from the video bitstream corresponding to an encoded picture, as in the following description related to

[0297] The encapsulation is handled by an ISOBMFF or mp4 writer. This writer contains a parser of the NAL unit header. It is able to extract the NALU type, the identifier and the corresponding compressed data. Typically, the extracted NALU data is placed in the media data container of the encapsulation file, the "mdat" box. The metadata used to describe the NAL units are placed in the structured hierarchy of boxes under the main "moov" box. One video bitstream is encapsulated in a video track described by a "trak" box with sub-boxes.

[0298] For partitioned video frames, there are different possible encapsulations depending on the intended use of the video. This use can be hardcoded in the mp4 writer application, for example in the initialization step 1200, or can be provided as an input parameter by a user or another program. In an embodiment, it can be convenient to encapsulate one frame part or a given set of frame parts in one video track, resulting in a multi-track encapsulation.

[0299] Once the initialization of the ISOBMFF writer is done, the encoder parses the video bitstream in step 1201 by reading the NALU types, in particular the NALU types corresponding to parameter sets (XPS). As already explained above, parameter sets are specific NAL units providing high level and general information about the encoding structure of the video bitstream, etc.

[0300] From the parsing of these parameter sets, the mp4 writer can judge in a test 1202 whether the video bitstream contains frame parts (e.g. whether a TPS or a specific partition structure is present in one of the parameter sets). If frame parts are present, the mp4 writer judges in the same test 1202 whether these frame parts are "constrained tiles", i.e. CTiles. If the bitstream does not contain frame parts or does not contain CTiles, the test 1202 is false and the video bitstream is encapsulated in one video track in step 1203.

[0301] The TPS (Tile Parameter Set) is considered as one NALU of parameter set information and can be embedded in the metadata providing decoder configuration or setup information as the DecoderConfigurationRecord box can be found in one of the boxes dedicated to sample description, etc. (e.g. the "stsd" box), typically in some codec specific sample entries.

[0302] Alternatively, according to embodiments of the application, the TPS can be processed as a NAL unit of video data (VCL NALU) and can be stored as one sample data in the "mdat" box. The TPS can also be present in the sample entry and in the sample data at the same time. If the frame portion structure changes along the video sequence, it is more convenient to store the TPS at the sample level (sample data) rather than at the sample description level (sample entry).

[0303] When the change of frame portion partition structure requires a decoder reset at the receiving side, the ISOBMFF writer preferably stores the TPS and CTile related information from the video bitstream in the sample entry. This decoder reset allows the device receiving or consuming the file to take into account the new partition structure. The new partition structure can for example contain an indication of the encoding tools to support (i.e. the profile) or the amount of data to process (i.e. the level). According to the profile and level values or other parameters of the partition structure, the device can or can not support the new partition structure. When not, the device can adjust the transmission or can select an alternative version of the video if available. When the new structure is supported, the device continues the decoding and rendering of the file.

[0304] When spatial access to the ROI is required (branch "yes" after test 1204), the ISOBMFF writer can have different encapsulation strategies depending on the use case. Spatial access refers for example to ROI or partition based display (i.e. only extracting and decoding a portion of data corresponding to the ROI or portion or frame portion or set of frame portions) or ROI or portion based streaming (i.e. only transmitting this portion of data and metadata corresponding to the ROI or portion or set of frame portions). If the use case expected is storage for local display, this corresponds to test 1205 being true (branch "yes"), it can be convenient to store the partitioned video bitstream in one track but including a NALU map to the ROI or frame portion or set of frame portions requiring spatial access. The NALU map is generated in step 1206. For the ISOBMFF writer, the NALU map consists in listing for each NALU in the video bitstream the NAL units related to a given CTile or having the same RAID reference (i.e. corresponding to a selectable and decodable frame portion or spatial region) in the RAID reference 1502. To be able to do this listing, the NALU parser module of the ISOBMFF writer checks according to embodiments of the bitstream generation the value of the identifier assigned to the ROI or frame portion or set of frame portions requiring spatial access in step 903. Figure 15

[0305] ​If the bitstream does not contain the NALU header specific identifier for the ROI or frame portion or set of frame portions for which spatial access is required, the ISOBMFF writer needs a slice header parser to obtain the value of the identifier of the CTile assigned in 903, e.g. ctile_unique_identifier (in Figure 13a

[0306] Then, in step 1206, a NALUMapEntry structure "nalm" is created as a box under the "trak" box hierarchy to store the list of NALUs and their mapping to frame portions or sets of frame portions. For each frame portion or set of frame portions, a SampleGroupDescriptionBox of type "trif" provides a description of the frame portion or set of frame portions, e.g. the parameters from the TileRegionGroupEntry of ISO / IEC 14496-15, for each frame portion or set of frame portions. The group_ID value of the frame portion descriptor "trif" is set to the value of the identifier of the CTile to be encapsulated.

[0307] Then, in step 1203, the data of all frame portions or sets of frame portions is encapsulated as a single track. When the use case is streaming, which corresponds to the test 1208 being true (branch "yes"), it can be convenient to split each frame portion or set of frame portions corresponding to a level of spatial access in the video into a dedicated track, and the single track encapsulation is done when the test 1208 is true.

[0308] For the streaming use case, frame portion descriptions are generated in step 1209 with respect to the NALU mapping for each frame portion or set of frame portions. The number of frame portions or sets of frame portions can be determined by parsing the TPS. The "trif" sample group is used, even the default sample group can be used, as there is one frame portion or set of frame portions for each frame portion track (track used to encapsulate the data related to one frame portion or set of frame portions) generated in step 1210. Then, according to ISO / IEC 14496-15, all samples are mapped to the same sample group description which is the frame portion descriptor "trif". The group_ID value of the frame portion descriptor "trif" is set to the value of the identifier of the CTile (or RAID if present) to be encapsulated.

[0309] ​Then, in step 1210, each frame portion or set of frame portions is inserted in a self track, namely a frame portion track. The frame portion track comprises specific sample entries indicating that the samples are actually spatial portions of a video and, when there are no more frame portions or sets of frame portions to encapsulate (test 1211), reference the frame portion base track created in step 1212. This frame portion base track contains specific NAL units corresponding to parameter sets including a timing parameter set (TPS). The frame portion base track sequentially references each frame portion track with a specific track reference type to allow an implicit reconstruction of any selection of frame portions or sets of frame portions. Step 1212 can be replaced by a composition track in which a NAL unit called an extractor provides an explicit reconstruction from one or more frame portion tracks.

[0310] Then, the extractor allows any configuration of frame portions or sets of frame portions (even different from the original configuration) by simply pointing the extractor to a given identifier of a frame portion or set of frame portions for each sample of the composition track, usually referencing the identifier of the corresponding CTile (or RAID if present).

[0311] When using a composition track in step 1212, the frame portion tracks in step 1210 can actually be decodable frame portion tracks, meaning each contains a frame portion description (generated in 1209) as well as parameter sets. The presence of a TPS in each frame portion track is optional since the extractor can recombine in different ways. Then, the sample description can indicate a sample entry compliant with the codec in use: "hvc1" or "hvc2" if HEVC is in use, "avcl" or "avc2" if AVC (Advanced Video Coding) is in use, or any reserved four-character code explicitly identifying the video encoder in use.

[0312] Figure 18 is a schematic block diagram of a computing device 1800 for implementing one or more embodiments of the application. The computing device 1800 can be a device such as a microcomputer, a workstation or a light portable device. The computing device 1800 comprises a communication bus connected to the following components:

[0313] - a central processing unit 1801, denoted CPU, such as a microprocessor;

[0314] - a random access memory 1802, denoted RAM, for storing the executable code of the method of the embodiments of the application and the following registers adapted to record the variables and parameters necessary for implementing the method according to the embodiments of the application, the memory capacity of the RAM 1802 being able to be extended, for example, with optional RAMs connected to an expansion port;

[0315] - a read-only memory 1803, denoted ROM, for storing computer programs for implementing embodiments of the application;

[0316] - a network interface 1804, which is typically connected to a communication network via which the transmission or reception of digital data to be processed takes place. The network interface 1804 can be a single network interface or comprise a set of different network interfaces (for example, a wired interface and a wireless interface or different kinds of wired or wireless interfaces). Under the control of a software application running in the CPU 1801, data packets are written to the network interface for transmission or read from the network interface for reception;

[0317] - a user interface 1805, which can be used to receive input from a user or to display information to a user;

[0318] - a hard disk 1806, denoted HD, which can be provided as a mass storage device;

[0319] - an I / O module 1807, which can be used for the transmission / reception of data with respect to external devices such as video sources or displays.

[0320] The executable code can be stored in the read-only memory 1803, on the hard disk 1806 or on a removable digital medium such as a disk. According to a variant, the executable code of the program can be received via the network interface 1804 using a communication network before execution, so as to be stored in one of the storage means of the communication device 1800, such as the hard disk 1806.

[0321] The central processing unit 1801 is adapted to control and direct the execution of the instructions or part of the software code of one or more programs according to embodiments of the application, which are stored in one of the above-mentioned storage means. After power-up, the CPU 1801 is able to execute, for example, the instructions relating to a software application from the main RAM memory 1802 after loading them from the program ROM 1803 or the hard disk (HD) 1806. This software application, when executed by the CPU 1801, makes it possible to carry out the steps of the flowchart of the application.

[0322] Any of the steps of the algorithms of the application can be implemented in software by execution of a set of instructions or program by a programmable computer such as a PC ("Personal Computer"), a DSP ("Digital Signal Processor") or a microcontroller, or in hardware by means of a machine or a dedicated component such as an FPGA ("Field-Programmable Gate Array") or an ASIC ("Application-Specific Integrated Circuit").

[0323] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to these specific embodiments, and modifications within the scope of the invention will be apparent to those skilled in the art.

[0324] Numerous other modifications and changes will be apparent to those skilled in the art when referring to the foregoing exemplary embodiments, which are given by way of example only and are not intended to limit the scope of the invention, and which are determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0325] The various embodiments of the present invention described above may be implemented individually or as a combination of multiple embodiments. In addition, features from different embodiments may be combined where necessary or where it is beneficial to combine elements or features from various embodiments into one embodiment.

[0326] Unless expressly stated otherwise, each feature disclosed in this specification (including any accompanying claims, abstracts, and drawings) may be replaced by alternative features serving the same, equivalent, or similar purpose. Therefore, unless expressly stated otherwise, each feature disclosed is merely an example of a general series of equivalent or similar features.

[0327] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

Claims

1. A method for encoding a frame into a bitstream, the frame being spatially divided into frame portions, the method comprising: encoding frame portions in the frame into the bitstream; representing in the bitstream an identifier of each of the frame portions in the frame; representing a first information corresponding to a number of the frame portions in the frame; and representing spatial information related to a position of a frame portion within the frame, wherein the identifier, the first information and the spatial information are represented in one parameter set in the bitstream, and a second information representative of a number of bits used to represent the identifier is further represented in the bitstream, the number of bits used to represent the identifier being variable, and wherein a flag related to loop filtering across a boundary of frame portions is represented in the bitstream. The frame portions are independently encoded.

2. The method of claim 1, wherein, A flag is provided indicating that the frame portions have been independently encoded.

3. The method of claim 2, further comprising: The position of the frame portions is given with respect to the frame.

4. The method of claim 1, wherein, The flag indicates whether loop filtering across a boundary of frame portions is enabled.

5. The method of claim 1, wherein, 6. A method for decoding video data from a bitstream, the video data comprising a frame, the frame being spatially divided into frame portions, the method comprising: obtaining from one parameter set in the bitstream an identifier of each of the frame portions in the frame, a first information corresponding to a number of the frame portions in the frame, and spatial information related to a position of a frame portion within the frame, wherein the identifier, the first information and the spatial information to be obtained are represented in the one parameter set in the bitstream, and a second information representative of a number of bits used to represent the identifier is further represented in the bitstream, the number of bits used to represent the identifier being variable, and wherein a flag related to loop filtering across a boundary of frame portions is represented in the bitstream; determining the position of the frame portion within the frame based on the spatial information; and decoding the frame portions from the bitstream. The frame portions are independently encoded.

7. The method of claim 6, wherein, A flag is obtained indicating that the frame portions have been independently encoded.

8. The method of claim 7, further comprising: The position of the frame portions is given with respect to the frame.

9. The method of claim 6, wherein, The flag indicates whether loop filtering across a boundary of frame portions is enabled.

10. The method of claim 6, wherein, 11. An apparatus for encoding a frame into a bitstream, the frame being spatially divided into frame portions, the apparatus comprising circuitry configured to: encode frame portions in the frame into the bitstream; the circuitry being further configured to: wherein represent in the bitstream an identifier of each of the frame portions in the frame; represent a first information corresponding to a number of the frame portions in the frame; and represent spatial information related to a position of a frame portion within the frame, wherein the identifier, the first information and the spatial information are represented in one parameter set in the bitstream, and a second information representative of a number of bits used to represent the identifier is further represented in the bitstream, the number of bits used to represent the identifier being variable, and wherein a flag related to loop filtering across a boundary of frame portions is represented in the bitstream. ​ representing spatial information related to the position of the frame portions within the frame, wherein the identifier, the first information and the spatial information are represented in one parameter set in the bitstream, and wherein a second information representative of the number of bits used to represent the identifier is further represented in the bitstream, the number of bits used to represent the identifier being variable, and wherein a flag related to the loop filtering across the boundaries of the frame portions is represented in the bitstream.

12. An apparatus for decoding a frame from a bitstream, the frame being spatially divided into frame portions, the apparatus comprising circuitry configured to: obtain, from one parameter set in the bitstream, an identifier of each of the frame portions in the frame, a first information corresponding to the number of frame portions in the frame, and spatial information related to the position of the frame portions within the frame, wherein the identifier, the first information and the spatial information to be obtained are represented in the one parameter set in the bitstream, and wherein a second information representative of the number of bits used to represent the identifier is further represented in the bitstream, the number of bits used to represent the identifier being variable, and wherein a flag related to the loop filtering across the boundaries of the frame portions is represented in the bitstream; determine the position of the frame portions within the frame based on the spatial information; and decode the frame portions from the bitstream.

13. A computer-readable storage medium having stored therein instructions of a computer program for implementing the method of any of claims 1 to 10.

Citation Information

Patent Citations

  • Video data stream concept

    CN104685893A