Encoding / decoding video picture data using picture blocks

By segmenting the video picture into picture chunks without intra-dependence relationships and decoding characteristics are defined in metadata, the problem of low encoding and decoding efficiency of video picture is solved, and flexible access and efficient decoding of video content is achieved.

CN120380754APending Publication Date: 2025-07-25BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081044.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-05-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video picture encoding and decoding technologies cannot effectively support flexible access and efficient decoding of video content, resulting in waste of resources and poor user experience.

Method used

The video picture is spatially divided into multiple picture blocks without intra-dependence, and partitioned metadata is written in the metadata structure, defining the decoding characteristics of each spatial partition, and notifying the client's spatial partitioning and decoding requirements of the video picture through signaling operations.

Benefits of technology

It realizes efficient encoding and decoding of video images, supports flexible access to video streams, avoids waste of resources, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380754A_ABST
    Figure CN120380754A_ABST
Patent Text Reader

Abstract

The application relates to encoding / decoding a video picture (PC1) using picture blocks (TL). The method for encoding a video picture comprises spatially segmenting the video picture (PC1) into a plurality of picture blocks (TL) without intra-frame dependencies; writing partition metadata of at least one spatial partition (SP) in a metadata structure, each spatial partition containing at least one picture block (TL) of the video picture (PC1), the partition metadata of each spatial partition (SP) defining decoding characteristics of video picture data of said at least one picture block (TL) of the spatial partition (SP); and encoding video picture data of at least one picture block (TL) of the at least one spatial partition (SP) according to the partition metadata of the at least one spatial partition (SP).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority based on and claims the benefit of European Patent Application No. 22306754.7, filed on November 29, 2022, the entire content of which is incorporated herein by reference. Technical field

[0003] This application generally relates to video picture encoding and decoding. In particular, but not limited to, the technical field of this application relates to encoding video picture data of picture blocks of a video picture, decoding such encoded video picture data, and corresponding encoding and decoding devices. Background art

[0004] This section aims to introduce to the reader various aspects of the field that may be related to various aspects of at least one exemplary embodiment of the present application described and / or claimed below. This discussion is considered to help provide background information for the reader to better understand the various aspects of the present application. Therefore, it should be understood that these statements are to be understood from this perspective and not as an admission of prior art.

[0005] It is well known that the streaming of video content is for all or part of the consumption (or access) by user applications (or users) of different specifications. Currently, the streaming of video can be achieved by using scalable video coding (SVC) that requires a dedicated system architecture, or by separating the full - resolution video into blocks before transmission.

[0006] However, the current video coding and decoding based on picture blocks is not yet satisfactory and efficient enough. In particular, once the video content is downloaded, the encoder may be enabled to decode the received video content, which results in poor quality of service and inconvenience to users. This is because even when the video content contains picture blocks, the encoder may not be able to decode the video content for partial consumption. Therefore, resources such as processing and bandwidth may be wasted, and the user experience may also be limited.

[0007] Therefore, there is a need for encoding and decoding techniques to ensure the encoding and / or decoding efficiency of video pictures. In particular, there is a need for an efficient encoding and / or decoding technique that allows flexibility in the way of accessing video pictures (such as video streams), for example, for full or partial consumption by clients or user applications.

[0008] Based on the above considerations, at least one exemplary embodiment of the present application has been designed. Summary of the invention

[0009] The following presents a simplified overview of at least one exemplary embodiment to provide a basic understanding of certain aspects of the present application. This overview is not a comprehensive review of the exemplary embodiments. It is not intended to identify the key or core elements of the exemplary embodiments. The following overview only presents certain aspects of at least one exemplary embodiment in a simplified form as a prelude to the more detailed description provided in other parts of this document.

[0010] According to a first aspect of the present application, there is provided a method for encoding a video picture, the method comprising:

[0011] - spatially dividing the video picture into a plurality of picture blocks without intra-frame dependencies;

[0012] - writing partition metadata of at least one spatial partition in a metadata structure, each spatial partition containing at least one picture block of the video picture, and the partition metadata of each spatial partition defining: the decoding characteristics of the video picture data of the at least one picture block of the spatial partition; and

[0013] - encoding the video picture data of the at least one picture block of the at least one spatial partition into a container according to the partition metadata of the at least one spatial partition.

[0014] According to a second aspect of the present application, there is provided a method for decoding a video picture implemented by a media player, the video picture being divided into a plurality of picture blocks without intra-frame dependencies, the method comprising:

[0015] - reading the partition metadata of at least one spatial partition from a metadata structure, each spatial partition containing at least one picture block of the video picture, and the partition metadata of each spatial partition defining: the decoding characteristics of the video picture data of the at least one picture block of the spatial partition;

[0016] - determining the supported spatial partitions by comparing the partition metadata of the at least one spatial partition with the decoding capabilities of the media player;

[0017] - decoding the video picture data of the at least one picture block of the supported spatial partitions from a first container.

[0018] In an exemplary embodiment, the partition metadata of the at least one spatial partition defines: a spatial region in the video picture.

[0019] In an exemplary embodiment, the spatial region is defined based on the partition metadata of at least one picture block of the video picture.

[0020] In an exemplary embodiment, a spatial region is defined by partition metadata based on at least one of the following: at least one dimension of the spatial partition; and a reference point indicating a position of the spatial partition relative to a picture chunk of a video picture.

[0021] In an exemplary embodiment, at least part of the metadata structure is included in a descriptor file that is separate from and points to a first container.

[0022] In an exemplary embodiment, the descriptor file is a media presentation descriptor.

[0023] In an exemplary embodiment, the descriptor file defines: the positioning of one or more second containers including the first container.

[0024] In an exemplary embodiment, the descriptor file defines: the partition metadata and positioning of one or more picture chunks of the at least one spatial partition.

[0025] In an exemplary embodiment, at least part of the metadata structure is included in the first container.

[0026] In an exemplary embodiment, the partition metadata is included in at least one of the following:

[0027] - in a bitstream conversion descriptor incorporated into the first container; and

[0028] - in the metadata information of the first container.

[0029] In an exemplary embodiment, a first part of the metadata structure is included in a descriptor file that is separate from and points to the first container, and the first part of the metadata structure includes a flag indicating the presence of the at least one spatial partition within the video picture,

[0030] wherein a second part of the metadata structure is included in the first container.

[0031] In an exemplary embodiment, the decoding characteristics of the partition metadata define: at least one codec parameter for decoding the at least one spatial partition.

[0032] In an exemplary embodiment, the partition metadata includes at least one of the following:

[0033] - an identifier of the at least one spatial partition;

[0034] - chunk information indicating at least one picture chunk of the at least one spatial partition;

[0035] - the resolution of the at least one spatial partition;

[0036] - Define region information of a spatial region occupied by the at least one spatial partition; and

[0037] - Define ratio information of an aspect ratio of the at least one spatial partition.

[0038] In one exemplary embodiment, the first container is an ISO BMFF file.

[0039] According to a third aspect of the present application, there is provided a container that is formatted to include encapsulated data obtained from the method according to the first aspect of the present application.

[0040] According to a fourth aspect of the present application, there is provided an apparatus including means for performing one of the methods according to the first and / or second aspects of the present application.

[0041] According to a fifth aspect of the present application, there is provided a computer program product that includes instructions which, when executed by one or more processors, cause the one or more processors to perform the methods according to the first and / or second aspects of the present application.

[0042] According to a sixth aspect of the present application, there is provided a non-transitory storage medium (or storage medium) that carries instructions for a program code for performing any one of the methods according to the first and / or second aspects of the present application.

[0043] The present application provides encoding and decoding techniques that ensure the encoding and / or decoding efficiency of video pictures. In particular, efficient encoding and / or decoding techniques can be achieved by allowing flexibility in the way of accessing video pictures (such as video streams), for example, for full or partial consumption by a client or user application. Thanks to the present invention, waste of resources and inconvenience to users can also be restricted or avoided.

[0044] The specific nature of at least one exemplary embodiment and other objectives, advantages, features, and uses of the at least one exemplary embodiment will become apparent from the following description of the examples in conjunction with the drawings. Description of the Drawings

[0045] Exemplary embodiments of the present application will now be described by way of example with reference to the drawings, where:

[0046] Figure 1 Schematically shows an example according to the prior art, a video picture divided into slices;

[0047] Figure 2 Schematically shows an example according to the prior art, the slice boundaries of a slice of a video picture;

[0048] Figure 3 Schematically shows a video picture for encoding according to at least one specific exemplary embodiment of the present application;

[0049] Figure 4 Schematically shows a video picture that is spatially divided into a plurality of picture blocks and includes spatial partitions according to at least one specific exemplary embodiment of the present application;

[0050] Figure 5 and Figure 6 Schematically shows a video picture that is spatially divided into a plurality of picture blocks and includes a plurality of spatial partitions according to a specific exemplary embodiment of the present application;

[0051] Figure 7 and Figure 8 Schematically shows a plurality of video pictures that are spatially divided into a plurality of picture blocks and include spatial partitions with different aspect ratios according to a specific exemplary embodiment of the present application;

[0052] Figure 9 Shows a schematic block diagram of steps of a method for encoding a video picture according to at least one specific exemplary embodiment of the present application;

[0053] Figure 10 Schematically shows a video picture that is spatially divided into a plurality of picture blocks and includes spatial partitions according to at least one specific exemplary embodiment of the present application;

[0054] Figures 11 - 13 Shows a schematic block diagram of a plurality of ISO BMFF files according to a specific exemplary embodiment of the present application;

[0055] Figure 14 Shows a schematic block diagram of steps of a method for decoding a video picture according to at least one specific exemplary embodiment of the present application;

[0056] Figure 15 Shows a schematic block diagram of steps of a method for decoding a video picture according to at least one specific exemplary embodiment of the present application;

[0057] Figure 16 Shows a schematic block diagram of an example of a system in which various aspects and exemplary embodiments are implemented.

[0058] Similar or identical elements are denoted by the same reference numerals. Detailed Description

[0059] As noted above, video streaming can be achieved by using scalable video coding which requires a dedicated system architecture, or by splitting the full-resolution video into picture chunks before transmission. However, it has been observed that if the picture chunks are not split (or formed) before transmission, there is a risk that the client may not be aware of the existence of these picture chunks, or may not be aware of the applicability of these chunks for partial consumption, without first downloading and analyzing the video content. There is a risk that the client may not download the video content based on the wrong assumption that it cannot be decoded and read, because the client is not aware that the content can be partially decoded. In addition, if the client downloads video content (e.g., a video stream) in the hope of extracting picture chunks for use during decoding, the decoding may fail, either because the relevant picture chunks cannot be successfully extracted, or because the video content is not encoded in an appropriate manner, resulting in a waste of processing resources, bandwidth, and time; even if the decoding is successful, the manner in which the chunk splitting is done may result in the actual content of the chunks being unsuitable for (partial) consumption. The video content may indeed need to be encoded in a specific manner to allow for full or partial decoding.

[0060] The present application thus provides various aspects of an encoding / decoding technique for encoding / decoding video content, such as a video stream, or more generally one or more video pictures comprising a plurality of picture chunks. In particular, one aspect of the technique relies on a new signaling operation that informs the client about the existence of the picture chunks and their decoding requirements, so that these information elements can be accessed before transmitting a video stream containing these picture chunks. In particular, the signaling for the streaming of video pictures can be used to indicate at least one spatial partition of the video picture and associated decoding characteristics. For example, such signaling can indicate spatial partitions with different decoding characteristics.

[0061] This approach can provide convenience in a content-independent manner, for example, in use cases where the video content is downloaded but only partially consumed (i.e., at a lower resolution - by omitting at least some of the frames). Partial consumption (or partial access) of the video content can be achieved by decoding only one or more plural regions (referred to as spatial partitions) of the video picture. In other words, the sub-parts of the video picture corresponding to its spatial partitions can be decoded, while other parts are left aside. The parts of the video frame that are decoded can also be referred to as derivative frames.

[0062] The technique of the present application specifically allows for signaling support for spatial partitions of video picture data, the existence of picture chunks in the video picture data (e.g., in a bitstream), and decoding characteristics (e.g., compliance points (e.g., profiles, levels, and / or tiers, etc.)) that can be used for the purpose of decoding the video picture data.

[0063] In the following, at least one exemplary embodiment will be described more fully with reference to the accompanying drawings, in which examples of at least one exemplary embodiment are depicted. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it is to be understood that the exemplary embodiments are not intended to be limited to the specific forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives falling within the scope of this application.

[0064] At least one aspect generally relates to video picture encoding, another aspect generally relates to transmitting a metadata structure including partition metadata, and yet another aspect relates to decoding encoded video picture data using such partition metadata.

[0065] At least one of the exemplary embodiments is described for encoding / decoding one video picture, but extends to encoding / decoding multiple video pictures (a sequence of pictures) since each video picture is encoded / decoded sequentially as described below.

[0066] In addition, at least one exemplary embodiment is not limited to the MPEG standard, such as AVC (ISO / IEC 14496-10 Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), but can also be applied to other standards and recommendations, such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ). The exemplary embodiments of the application described in the present disclosure can be applied to existing or future-developed standards and recommendations and their extensions. Unless otherwise stated or technically excluded, the various aspects described in this application can be used alone or in combination.

[0067] The following provides several definitions and observations applied to this application.

[0068] A video picture (or image) can be a video frame belonging to a video, that is, a time series of video frames. There is a temporal relationship between multiple video frames of a video.

[0069] A video picture (or image) can also be a still image.

[0070] Hereinafter, the exemplary embodiments of this application will be discussed by regarding the video picture (or image) as a still image or as a video frame of a video.

[0071] An image includes at least one component (also called a channel) determined by a specific picture / video format, which specifies all information related to the sampling values and all information that can be used by a display unit and / or any other device to display and / or decode the image data related to the image to generate pixel values.

[0072] An image includes at least one component, which is typically represented in the form of a 2D sampling array.

[0073] The image can be a monochrome or a color image. A monochrome image includes a single component, while a color image (also referred to as a texture image) can include multiple components, such as three components.

[0074] For example, when the image / video format is the well-known (Y, Cb, Cr) format, a color image can include a luma (or luminance) component and two chroma components; or when the image / video format is the well-known (R, G, B) format, a color image can include three color components (one component for red, one component for green, and one component for blue).

[0075] Each component of the image can include a number of samples related to the number of pixels of the display screen on which the image is to be displayed. For example, the number of samples contained in a component can be the same as or a multiple (or fraction) of the number of pixels of the display screen on which the image is to be displayed.

[0076] The number of samples contained in a component can also be a multiple (or fraction) of the number of samples contained in another component of the same image.

[0077] For example, in the case of an image / video format (such as the (Y, Cb, Cr) format) that includes a luma component and two chroma components, depending on the color format considered, the number of samples contained in the chroma components in terms of width and / or height can be half of that of the luma component.

[0078] A sample is the smallest visual information unit that makes up a component of an image. For example, a sample value can be a luminance value or a chroma value, or a color value of the red, green, or blue component in the (R, G, B) format.

[0079] For a monochrome image, the pixel value of the display surface can be represented by one sample; while for a color image, it can be represented by multiple co-located samples. Co-located samples associated with a pixel refer to the samples corresponding to the location of the pixel in the display screen.

[0080] An image (or video frame) is generally regarded as a set of pixel values, with each pixel represented by at least one sample.

[0081] In the following, exemplary embodiments of the present application are discussed by considering picture partitioning conforming to the HEVC standard (ISO / IEC 23008-2 High Efficiency Video Coding (HEVC) / ITU-T Recommendation H.265). However, the concepts of the present invention can be applied in an analogous manner to other codecs, particularly any codec that supports a coding structure with any intra-dependency, such as the VVC standard (ISO / IEC 23090-3 Versatile Video Coding (VVC) / ITU-T Recommendation H.266).

[0082] Each video picture coded according to HEVC can be divided into Coded Tree Blocks (CTBs). More coarsely, each video picture can be divided into slices and tiles. A slice is a sequence of one or more slice segments, which starts with an independent slice segment and contains all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same picture. A slice segment is a sequence of coding tree units (CTUs). Similarly, a tile is also a sequence of CTUs.

[0083] For example, a video picture can be divided into two slices, as schematically illustrated in Figure 1 . In this example, the first slice consists of an independent slice segment containing 4 CTUs, a dependent slice segment containing 32 CTUs, and another dependent slice segment containing 24 CTUs; and the second slice consists of a single independent slice segment containing the remaining 39 CTUs of the video picture.

[0084] As schematically illustrated by way of example in Figure 2 , all CTBs of a video picture are within a slice, and the slice boundary always aligns with the CTB boundary. These CTBs form rows and columns, which can be grouped in rectangular regions (referred to as picture tiles, or more simply tiles). The tile boundary always aligns with the CTB boundary. However, the slice boundary may not always align with the tile boundary.

[0085] Furthermore, the tile boundary indicates an interruption of coding dependencies including motion vector prediction, intra prediction, and context selection. The only exception to the dependency may be loop filtering, which can be disabled to make the decoding of a tile completely independent.

[0086] In the following, the exemplary embodiments of the present application under discussion use the scalable extension of HEVC (SHVC - AnnexH). However, the concept of the present application can be applied to other codecs, such as any codec that supports a scalable video coding mechanism, for example, the AVC standard (ISO / IEC 14496 - 10 Advanced Video Coding (AVC) / ITU - T Recommendation H.264), and its corresponding scalable extension (SVC).

[0087] In scalable video coding, the encoding process produces (at least) two encoded output streams. One is considered the "base layer", which contains the encoded video with specifications for low - end devices, and one (or more) "enhancement layers", which can be decoded and combined with the base layer to produce a higher - quality output. Each layer has different characteristics and can even be created using different but compatible codecs.

[0088] Two decoders can be used to produce the final (enhanced) output, which can be a significant limitation, especially if the base layer is encoded with a lower - complexity codec (such as AVC in the example), yet this may be required because the computational cost of reconfiguring a single decoder to decode the bitstreams of two different codecs during runtime is very expensive (if possible), and this can pose a severe limitation to those use cases that most desire the enhanced output.

[0089] In the following, the exemplary embodiments of the present application under discussion use dynamic adaptive streaming via HTTP (DASH) (see ISO / IEC 23009 - 1 Dynamic adaptive streaming over HTTP (DASH) — Part 1: Media presentation description and segment formats). However, the concept of the present application can be applied to other streaming protocols, such as any codec of any streaming protocol that supports multiple representations of video content, for example, HTTP Live Streaming (HLS).

[0090] DASH uses a Media Presentation Description (MPD) to inform the client about the codec characteristics of the media content before downloading, enabling the client to select the media content with the most suitable characteristics. The same principle applies to any other file / descriptor that contains the codec characteristics of the media content, such as the Session Description Protocol (SDP).

[0091] Video pictures can be encoded at various resolutions and bitrates to facilitate processing through various device, application, and network requirements. Additionally, video pictures can be split into segments in a temporal manner to enable transitions between resolutions and bitrates at segment boundaries. Information containing the source and characteristics of the video content can be stored in a file that is used as a reference when called by a client, which in the case of DASH is the Media Presentation Description (MPD).

[0092] In the MPD, the video content is contained within one or more period elements that indicate time groupings. A period can contain multiple adaptation sets, each of which indicates a different type of media type content (e.g., one adaptation set for video encoded in AVC and one adaptation set for video encoded in HEVC). Each adaptation set can then contain multiple representations, where each representation can be differentiated in terms of content resolution, bitrate, frame rate, etc.

[0093] In the exemplary embodiments below of the present application, chunking and adaptive streaming techniques can be advantageously combined to provide chunk-based adaptive streaming capabilities. To this end, video pictures can be split into multiple chunks, which can then be used for streaming and / or decoding.

[0094] In the exemplary embodiments below of the present application, scalable video coding and adaptive streaming techniques can be advantageously combined to provide a video bitstream encoded using a scalable scheme over an adaptive streaming platform. Adaptive streaming of scalable video coding enables low-performance devices to consume video-encoded content that can be consumed by more powerful devices at a higher quality.

[0095] It has been observed that SHVC (in general streaming of DASH and scalable video coding) cannot provide satisfactory results, especially in terms of encoding / decoding efficiency, as it requires the preparation of (at least) two separate streams (base layer and enhancement layer). Additionally, if non-SHVC content is required, another (third) version of the content should be prepared and stored. Further, if different codecs are used for the base layer and the enhancement layer, two decoders should be used on the client side. For decoding, SHVC decoders typically are based on two internal decoders (one for the base layer and one for the enhancement layer), which makes coordination and synchronization of the devices difficult and resource-consuming.

[0096] As pointed out previously, one aspect of the present application relies on signaling or indicating to a client (or decoder) that a video picture includes at least one spatial partition that can be consumed (or decoded) independently. Each spatial partition may include at least one picture chunk of the video picture. As further described in the specific embodiments below, encoding according to the present application may mean spatially splitting (or partitioning) the video picture into multiple picture chunks without intra-frame dependencies. Various examples of such chunking are provided below for illustrative purposes only.

[0097] Figure 3 Schematically shown is a video picture PC1 that can be encoded and / or decoded according to an exemplary embodiment of the present application. In this example, as already noted, the video picture PC1 can have various properties and can include, for example, at least one graphic object OB.

[0098] Figure 4 Schematically shown is an exemplary embodiment in which the video picture PC1 is spatially split into 9 picture chunks TL. In this example, the central picture chunk forms a spatial partition SP1, i.e., a spatial region of the video picture PC1. This spatial partition SP1 can represent the main content of the video picture PC1, i.e., can be regarded as the most important (or more useful) content in the video picture PC1. More generally, the spatial partition SP1 can represent any content of interest, i.e., any sub-part of the video picture PC1. The way the spatial partition SP1 is defined within the video picture PC1 may vary for each case and may be particularly adjusted by the client.

[0099] In Figure 4 the specific example, the spatial partition SP1 corresponds to one picture chunk TL, although other embodiments in which the spatial partition can include more than one picture chunk TL are also possible.

[0100] The video picture PC1 can include multiple spatial partitions. Figure 5 Shown is an exemplary embodiment in which the video picture PC1 is spatially split into 25 picture chunks TL. As illustrated, multiple spatial partitions SP, i.e., spatial partitions SP2 and SP3 in this example, can be defined within the video picture PC1. The spatial partition SP2 corresponds to the central picture chunk, while the spatial partition SP3 corresponds to 9 picture chunks TL including the one that is the spatial partition SP2.

[0101] In Figure 5 the example, the smaller spatial partition SP2 is part of the larger spatial partition SP3, although other embodiments are also possible. Each picture chunk TL of the spatial partition SP does not necessarily have to be included in the larger spatial partition.

[0102] Figure 6 An exemplary embodiment is schematically shown in which the video picture PC1 is spatially divided into 25 picture blocks TL, as Figure 5 shown. In this case, two partially overlapping spatial partitions SP4 and SP5 are defined within the video picture PC1. The spatial partition SP4 corresponds to 4 picture blocks TL, while the spatial partition SP corresponds to 12 picture blocks TL. It can be seen that only two picture blocks in the spatial partition SP4 are part of (shared with) the spatial partition SP5.

[0103] It should be noted that the aspect ratio of the spatial partition SP of the video picture PC1 can be different for each case.

[0104] Figure 7 An exemplary embodiment is schematically shown in which the video picture PC1 is spatially divided into 3 picture blocks TL with different aspect ratios. The video picture PC1 can include a spatial partition SP6 corresponding to the central picture block, which has an aspect ratio closer to the square ratio than the aspect ratio of the video picture PC1.

[0105] According to the above method, different aspect ratios can be considered to adapt to different use cases, such as changing the aspect ratio direction for video content or streaming for a smartphone. Figure 8 An example is schematically shown in which the direction change of the video picture PC1 (which is most suitable for presentation in a horizontal display) is spatially divided into 3 picture blocks TL, including a spatial partition SP7 (more suitable for presentation in a vertical display) corresponding to the central picture block with an aspect ratio different from that of the video picture PC1.

[0106] Exemplary embodiments of the method of the present application will be described below. The above description of Figures 3 - 8 is noteworthy in terms of the description of the picture blocks TL and the spatial partitions SP, which also applies to the embodiments described below.

[0107] Encoding

[0108] Figure 9 A schematic block diagram of steps 102 - 106 of a method 100 for encoding a video picture PC1 according to at least one exemplary embodiment of the present invention is schematically shown. By way of example, the video picture PC1 can be Figures 3 - 8 one of the video pictures shown in, although the nature of the video picture PC1 can be adjusted as already indicated. The steps of the method 100 can be implemented by an encoder (or encoding device or packager or transcoder) DV1, which will be described in detail in a specific exemplary embodiment later.

[0109] In the splitting step 102, the video picture PC1 is spatially split into a plurality of picture chunks TL that have no intra-frame dependencies. For example, it can be split into picture chunks according to the description of any one of the foregoing references Figures 3 - 8 as described.

[0110] As previously explained, a slice is a sequence of one or more slice segments that starts with an independent slice segment and includes all subsequent dependent slice segments (if any) up to the next independent slice segment (if any) within the same picture. The spatial splitting is performed in step 102 such that the picture chunks TL are independently decodable parts of the video picture PC1. This means that each chunk TL can be independently decoded without relying on any other chunk TL of the video picture PC1 for encoding or decoding. Thus, the chunk boundaries can indicate the interruption of encoding / decoding dependencies (including motion vector prediction, intra-prediction, and context selection).

[0111] In the writing step 104, the partition metadata MD1 of at least one spatial partition SP is written (or incorporated) into the metadata structure MS1, where the nature of the metadata structure MS1 can be adjusted according to each case, as further described below. The writing step 104 constitutes a signaling operation, as will be further discussed below. The partition metadata MD1 can define one or more spatial partitions SP. The spatial partition SP can be, for example, any one of the spatial partitions SP1 - SP7, as described in the reference Figures 3 - 8 as described.

[0112] As already indicated, the spatial partition SP defined by the partition metadata MD1 can represent the main content of the video picture PC1, i.e., the most important (or more useful) content in the video picture PC1. More generally, the spatial partition SP can represent any content of interest, i.e., any sub-part of the video picture PC. The way the spatial partition SP is defined within the video picture PC1 may vary according to each case and may be specifically adjusted by the user.

[0113] By way of example, assume that the partition metadata MD1 then defines a plurality of spatial partitions SP (e.g., Figure 5 SP2 and SP3 in Figure 6 or

[0114] SP4 and SP5 in Figures 5 - 6Taking the example shown, each spatial partition SP can correspond to one or more picture tiles TL. In such an instance, the tiles TL of each spatial partition SP can be aligned with the boundaries of the spatial partition SP. As further described below, other embodiments are possible where each spatial partition SP as defined in the partition metadata MD1 does not necessarily coincide with the tile segmentation of the video picture PC1. For example, the spatial partition SP can include at least part of the picture TL1 such that the tiles TL of the spatial partition SP are not aligned with the boundaries of the spatial partition.

[0115] For each spatial partition SP, the partition metadata MD1 defines: the decoding characteristics PR2 ( Figure 9 ) of the video picture data of one or more picture tiles TL of the spatial partition SP. For example, the partition metadata MD1 can indicate the presence of at least one spatial partition representing a given content within the video picture PC1, such as an area considered interesting (by the picture creator), which may be the most important part of the video picture PC1.

[0116] In a specific example, the partition metadata MD1 further defines for each spatial partition SP: the spatial region PR1 in the video picture PC1. The spatial region PR1 of each spatial partition SP in the video picture PC1 can be defined in various ways, depending on each case. The partition metadata MD1 can, for example, include the tile identifiers of each tile TL of the corresponding spatial partition SP. However, it should be noted that embodiments are possible where the partition metadata MD1 does not define the spatial region PR1.

[0117] Figure 10 An example is shown where the video picture PC1 (divided into 25 picture tiles TL, as Figures 5 - 6 shown) contains the spatial partition SP8 (the partition SP8 contains a complete picture tile TL and only sub - parts of other picture tiles TL). The partition metadata MD1 can define the spatial region PR1 of the spatial partition SP8 based on: at least one dimension of the spatial partition SP8 (such as dimensions DM1 and DM2, as Figure 10 shown) and a reference point RF1 indicating the position of the spatial partition SP8 relative to the picture tiles TL of the video picture PC1. Based on the dimension, the reference point RF1 (such as the upper - left corner) can be used as a reference within the video picture PC1 to identify the spatial partition SP8. This way of defining the spatial region PR1 of the spatial partition SP can be used, for example, if the spatial partition SP at hand covers an area that does not match (or is inconsistent with) the tile segmentation performed in step 102. Thus, greater flexibility in encoding the spatial partition in this way can be achieved. During subsequent decoding, the spatial partition can be decoded while the remaining picture tiles TL can be ignored, thereby improving the encoding / decoding efficiency.

[0118] The decoding characteristic PR2, as defined in the partition metadata MD1, defines the characteristics (or parameters) of the corresponding spatial partition SP, which can be used later to identify a video picture PC1 containing the compatible spatial partition SP and to decode the spatial partition SP during the decoding phase. These decoding characteristics PR2 can, for example, define compatible interoperability points (e.g., profiles, levels, tiers) and / or the codec characteristics (e.g., resolution, bit rate, bit depth) of one or more spatial partitions of the video picture PC1. The nature of the decoding characteristics and the way they are defined in the partition metadata MD1 can be adjusted, and some implementations will be described by way of example later.

[0119] The decoding characteristic PR2 of the spatial partition MD (or more generally the partition metadata MD1) can, for example, include the presentation characteristic PR3 of the video picture data of one or more picture blocks of the spatial partition MD ( Figure 9 ).

[0120] The nature, configuration, and use of the partition metadata MD1, in particular the decoding characteristic PR2 and the presentation characteristic PR3 (if any), will become more apparent in the exemplary embodiments described further below.

[0121] In a specific example, the spatial segmentation in step 102 is based on the block characteristics of at least one spatial partition SP to be encoded. This allows the way the video picture PC1 is spatially segmented to be adjusted according to the spatial partition SP used for encoding (and later for decoding).

[0122] In the encoding step 106 ( Figure 9 ), the video picture data DT0 of the picture blocks TL of the spatial partition SP is encoded into encoded video picture data DT1 according to the partition metadata MD1 of the spatial partition MD1. In this example, the video picture data DT0 is encoded (or encapsulated) into a container CT1 (also referred to as the first container), which can have any suitable form. The container CT1 can be one or more files, one or more network data packets, or any binary structure capable of containing a video bitstream. In other words, the encoded video picture data DT1 is generated by encoding one or more picture blocks TL (as defined previously in step 104) contained (at least partially) in each spatial partition SP of the video picture PC1 according to the partition metadata MD1 of the spatial partition SP. For this purpose, the video picture data DT0 of the relevant picture blocks can be obtained in any suitable way. The encoder DV1 can, for example, extract the video picture data DT0 from a source bitstream or from any suitable container or storage component.

[0123] Encoder DV1 can encode video picture data D0 into a bitstream BT1 incorporating encoded video picture data DT1( Figure 9 ), and this bitstream BT1 is encapsulated (or incorporated) into a container CT1.

[0124] In a specific example, descriptive metadata defining a picture chunk TL of a video picture PC1 is incorporated into the container CT1, for example, in at least one ISO BMFF box and / or at least one supplemental enhancement information (SEI) message.

[0125] In encoding method 100( Figure 9 ), the structural metadata MS1 for carrying (or storing, or transmitting) partition metadata MD1 can have various properties. In a specific example, the structural metadata MS1 (or at least a part thereof) can be included (or incorporated) in a separate descriptor file FL1, that is, a descriptor file separate from the container CT1 containing the encoded video picture data DT1. In a specific example, the structural metadata MS1 (or at least a part thereof) can be included (or incorporated) in the container CT1. In a specific example, a first part of the structural metadata MS1 can be included (or incorporated) in a separate descriptor file FL1, and a second part of the structural metadata MS1 can be included (or incorporated) in the container CT1. The first part and the second part can together form the structural metadata MS1.

[0126] Then the container CT1 and the metadata structure MS1 can be transmitted to a client (decoder) to decode the encoded video picture data DT1 if necessary.

[0127] It should be noted that the writing step 104( Figure 9 ) can be executed before or after the encoding step 106. For example, the video picture data DT0 can be encoded first, and then the partition metadata MD1 can be written. Alternatively, the partition metadata MD1 can be written first, and then the video picture data DT0 can be encoded.

[0128] As will be further described in the decoding stage later, the encoding method 100 allows signaling support for at least one spatial partition SP, which is used to decode the encoded video picture data DT1 and for decoding characteristics (such as compliance points (such as profiles, levels, and / or tiers, etc.)) that can be used for decoding purposes. For this purpose, the metadata structure MS1 containing the partition metadata MD1 can be used in the decoding stage to detect: partial consumption (or access) of one or more spatial partitions SP of the video picture PC1 can be achieved.

[0129] The concept of the present application thus lies in extending the current DASH syntax with additional information to indicate to the client that the video stream (or more generally, the video picture) includes at least one spatial partition that can be consumed independently, where the at least one spatial partition constitutes a lower compatibility point. This additional information (i.e., partition metadata MD1) can be signaled in the form of a metadata structure MS1, as described. Thus, the present application introduces a new signaling operation point that can advantageously be used to indicate a spatial partition within the video picture PC1 or in the video stream with specific decoding characteristics PR2. These alternative decoding characteristics PR2 of the spatial partition CP can be signaled, for example, with a lighter codec profile or a lower resolution. This additional information can be obtained in any suitable manner during the decoding phase by extracting it from the metadata structure MS1 as described above (e.g., from the container CT1 and / or from a descriptor file separate from the container CT1), as will be further described below.

[0130] As an exemplary use case, high-quality content can be provided by a user using his / her mobile phone (or similar device) for live streaming (e.g., using DASH). Since such devices have limited resources (processing power, battery, etc.), it is typically only produced in one quality (i.e., in one representation) (e.g., a single representation MPD), and this should be the highest available quality. Producing more representations, enhancement layers, or using other techniques would increase resource consumption and cause delays. However, this situation is problematic because the client may not be able to support the produced high-quality content during the decoding phase. Thus, in order to also support devices with lower specifications, the producer can perform the encoding by implementing the encoding method 100. Specifically, the content can be directly encoded using picture tiling TL (i.e., according to the partition metadata MD1), and the second metadata MD1 can be signaled in the metadata structure MS1, as previously described in the specific example.

[0131] As another use case, the encoding method 100 can also be applied to multi-party conferences. A user can, for example, record (or produce) high-resolution video content using his / her mobile phone or similar device. However, one or some of the participants in the multi-party conference session may not be able to support this high resolution. Accordingly, thanks to the present application, the producer can continue to produce the content in high resolution for all possible participants (possibly the majority), rather than providing a low resolution to all users, while users with lower processing power can still access a lower-quality version of the content, as will be further described below.

[0132] The present application thus allows expanding the range of supported devices for a given video content without expanding the number of representations.

[0133] As already pointed out, scalable video coding (SVC) can be performed for streaming video content to be consumed (or accessed) in whole or in part by user applications (or users) of different specifications. SVC can provide the same content with different coding parameters in order to be able to support both high-performance and light (less performant) clients at the same time. Conventionally, SVC requires generating (and storing and sending) (at least) two different streams, one for the base layer and one (or more) for the enhancement layer. Therefore, dedicated system architectures and specific processing are required.

[0134] However, by implementing the encoding method 100, picture chunking can be advantageously used in the encoding phase, and signaling of the partition metadata MD1 can be performed such that it may be possible to encode only one version (or a limited number of versions) of the content.

[0135] In addition, the partition metadata MD1 can be advantageously used in at least two different ways in the decoding phase. First, the client can make an informed decision about whether to consume the full or lighter content and decide whether to download the video content even before the decoding phase (e.g., initialization). The client (or decoder) can thus decide whether to download or not download the video content based on the partition metadata MD1 according to the available space partitions in the encoded video picture data DT1. In this first instance, signaling the partition metadata MD1 in an identifier file (e.g., MPD) separate from the container CT1 can be particularly effective.

[0136] Second, the client (or decoder) can download the video content (encoded video picture data DT1) and decide which version of the video content to decode based on the partition metadata MD1. Specifically, the decision of which version to decode can be made according to the available space partitions in the video content. In this second instance, signaling the partition metadata MD1 in the container CT1 and / or the bitstream BT1 can be particularly effective.

[0137] Compared to the SVC method (for the production and distribution chain of only one stream), the concept of the present application assumes significantly less infrastructure on the production side. In addition, by configuring the encoder and introducing the above signaling, an already deployed system can easily add support for the present application. Similarly, already deployed receiving devices can also be updated with new software, e.g., repurposing a video bitstream that exceeds its decoding capabilities for a device that matches it, since repurposing does not require decoding and re-encoding but only rewriting the bitstream syntax.

[0138] As already noted, for each spatial partition SP of the video picture PC1, the partition metadata MD1 may include decoding characteristics PR2, which may be applied by the client (or encoder) for decoding purposes. The content of these partition metadata MD1 can be adjusted for each case, and some examples are provided below for illustrative purposes.

[0139] In a specific example, the decoding characteristic PR2 of the partition metadata MD1 defines at least one codec parameter for decoding the spatial partition SP of the encoded video picture PC1.

[0140] In a specific example, the decoding characteristic PR2 of the partition metadata MD1 defines the resolution of the spatial partition SP.

[0141] In a specific example, the partition metadata MD1 includes at least one of the following:

[0142] - An identifier of the spatial partition SP;

[0143] - Block information indicating at least one picture tile TL of the spatial partition SP;

[0144] - The resolution of the spatial partition SP;

[0145] - Region information defining the spatial region PR1 occupied by the spatial partition SP; and

[0146] - Ratio information defining the aspect ratio of the spatial partition SP.

[0147] Various embodiments of the encoding method 100 will be described below. Hereinafter, an exemplary implementation of signaling the partition metadata MD1 written into the metadata structure MS1 in the writing step 104 ( Figure 9 ) will be described. Specific examples will be provided to show how to modify the DASH Part 1 (MPD) syntax and semantics (the latest available version) to apply this application.

[0148] Partition metadata MS1 in container CT1

[0149] In an exemplary embodiment, in the writing step 104 ( Figure 9 ), the partition metadata MD1 (or at least a part thereof) is written (and thus included) in the container CT1. As a result of the encoding method 100, the container CT1 thus contains the encoded video picture data DT1 associated with the partition metadata MD1 (e.g., in the form of a bitstream). In other words, the metadata structure MS1 carrying (or storing) the partition metadata MD1 and the encoded video picture data DT1 are incorporated (or encapsulated) together into the container CT1. Various implementations of the exemplary embodiment are described below.

[0150] The partition metadata MD1 (or at least a part thereof) can be incorporated or included, for example, in a bitstream transformation descriptor (BTD) within the container CT1. To this end, a new descriptor can be introduced. Such a new BTD can be an attribute of a media presentation descriptor (MPD), i.e., an existing MPD element of the bitstream. In this instance, the BTD can be defined at different levels of the MPD. The BTD does not have to be bound (or associated) with the current assumptions and / or any conventional signaling.

[0151] The BTD can be in the same descriptor as the descriptor of the main content of the video picture PC1 (e.g., in the same AdaptationSet) or in a dedicated descriptor for spatial partitioning (e.g., in an AdaptationSet that contains / describes only the spatial partitions and their characteristics).

[0152] As already pointed out, for each spatial partition SP of the video picture PC1, the partition metadata MD1 can include decoding characteristics PR2. In an exemplary embodiment, these decoding characteristics PR2 define at least one codec parameter for decoding the spatial partition SP of the video picture PC1. The at least one codec parameter can have various properties depending on the case.

[0153] The at least one codec parameter can include, for example, a codec identifier that identifies the codec that can be used to decode the spatial partition SP of the video picture PC1.

[0154] The at least one codec parameter can include the profile, level, and / or tier (code PLT) of the codec that can be used to decode the spatial partition SP of the video picture PC1. Thanks to the BTD, different configurations of the encoder can thus be signaled.

[0155] In an exemplary embodiment, the decoding characteristics PR2 in the BTD can signal a given codec profile (e.g., a codec profile different from the codec profile of the full video picture PC1). The BTD can be used as an attribute of an existing MPD element. An example of the syntax of the new BTD is provided in Table 1 below, where the "Common Attributes and Elements" semantic table is from DASH (it is chosen because it can advantageously be located at the AdaptationSet, Representation, or SubRepresentation levels of the MPD and can also use existing elements):

[0156] Table 1

[0157]

[0158]

[0159] In a specific example, derivative_codecs can be a comma-separated list of codecs supported by the corresponding element content.

[0160] An example of the syntax of an MPD with derivative_codecs is provided in Table 2 below:

[0161] Table 2

[0162]

[0163] For example, in Table 2 above, the main content has a @codecs value of HEVC, for the profile "Main", level 5.1, and "Main" tier (unconstrained), which has the value "hev1.1.6.L153.00"; and has spatial partitioning, for the profile "Main", level 3.1, and "Main" tier (unconstrained), which can have a @derivative_codecs attribute value of "hev1.1.6.L93.00".

[0164] One aspect of the present application is to use existing codec (e.g., HEVC chunking) and streaming (e.g., DASH MPD) constructs to extend the range of supported devices by adding support for low-end clients of the same encoded video stream.

[0165] The resolution of the spatial partition SP is lower than that of the main content (i.e., the full video picture PC1), but the resolution may not be the only factor for selection. Most video codecs have different "profiles" (and "levels" and "tiers" in AVC / HEVC / VVC, etc.), and different "profiles" have different requirements (such as supported frame rate, resolution, bit rate, etc.). Therefore, a spatial partition SP with a resolution different from that of the main content can result in the spatial partition SP being compatible with a different, lighter profile. Since the client can decide on support for the content by identifying representations with compatible profiles, the corresponding profile of the spatial partition SP can be signaled with partition metadata MD1, which may be associated with the corresponding resolution of the spatial partition SP.

[0166] In an exemplary embodiment, the coding characteristic PR2 of the metadata MD1 may define a codec for the spatial partition SP of the video picture PC1. In this way, it can be shown that bitstream conversion is possible, for example, to convert the bitstream BT1 of the encoded data DT1 ( Figure 9 ).

[0167] The following example of the syntax in Table 3 shows the same example as in Table 2, using BTD as a supplementary characteristic and adopting an optional characteristic method with MPD, that is, the SupplementalPropertyBTD scheme of "urn:mpeg:dash:btd:2022", to indicate the spatial partition codec value (for the main content, use hev1.1.6.L93.00 instead of hev1.1.6.L153.00):

[0168] Table 3

[0169]

[0170]

[0171] In an exemplary embodiment, the partition metadata MD1 in BTD contains a flag indicating the presence of a spatial partition within the video picture PC1. This flag may be associated with at least one codec parameter as previously described, for example, as part of the decoding characteristic PR2. In this instance, the client may need to retrieve additional decoding information from within the bitstream BT1 during the decoding phase.

[0172] In an exemplary embodiment, in BTD, the partition metadata MD1 for each spatial partition SP includes at least one of the following:

[0173] - An identifier of the spatial partition SP;

[0174] - Block information indicating at least one picture tile TL of the spatial partition SP;

[0175] - The resolution of the spatial partition SP;

[0176] - Region information defining the spatial region PR1 occupied by the spatial partition SP; and

[0177] - Ratio information defining the aspect ratio of the spatial partition SP.

[0178] The following provides an example of the syntax of BTD in Table 4

[0179] Table 4

[0180]

[0181] Each MPD element can incorporate multiple descriptors for defining partition metadata MD1 of multiple corresponding spatial partitions SP of video picture PC1.

[0182] It should be noted that since BTD is newly introduced according to the specific embodiments of the present application, the mere presence of descriptors in the metadata structure MS1 ( Figure 9 ) can already be interpreted by the client (or decoder) during or before the decoding stage as indicating that the spatial partitioning function of the present application is implemented. In this instance, a separate flag as described above may not be required, as long as the decoding attribute PR2 is signaled in the partition metadata MD1.

[0183] In an exemplary embodiment, the container CT1 (step 106, Figure 9 ) into which the video picture data DT0 is encoded is a binary file, called an ISO Base Media File Format (BMFF) file, as defined in the standard ISO / IEC 14496-12 (Coding of Audiovisual Objects – Part 12: ISO Base Media File Format). However, other file instantiations can be used without any limitation to the scope of the present application.

[0184] The metadata structure MS1 can thus be written (or incorporated) into the ISO BMFF file that serves as the container CT1.

[0185] In a specific example, the metadata structure MS1 is written as a BTD, and the BTD is included in the ISO BMFF file that serves as the container CT1.

[0186] Figure 11 An example of an ISO BMFF file containing the BTD (and thus the partition metadata MD1) is schematically shown.

[0187] As explained in the ISO BMFF specification, an ISO BMFF file is formed as a series of objects (referred to as "boxes" in this document). All data is contained in the boxes, and there is no other data within the file. This includes any initial signatures required for a specific file format. ISO BMFF allows multiple sequences of continuous samples (audio, video, subtitles, etc.) to be stored into the so-called track concept. Tracks are distinguished by their media processors.

[0188] In addition to media processors and track references, ISO BMFF also allows metadata defined in ISO / IEC 23002-2 to be stored as metadata items in the tracks.

[0189] In an exemplary embodiment, the video bitstream BT1 containing the picture tiles TL ( Figure 9) can be stored in an ISO Base Media File Format (ISO BMFF) track. To this end, the track ID can be used (instead of or together with the identifier of the picture chunk TL). This new category can be added to the current ISO BMFF standard.

[0190] More specifically, one of the following two options can be used to perform signaling of the partition metadata MD1: BTD or (1) in the MPD (at a higher (presentation) level) or (2) in the ISO BMFF (at a lower (file format) level).

[0191] For the first option (1) above, the BTD may contain PLT (and / or resolution) information, such as shown in Table 3 above, so that the client can know the existence of the spatial partition SP during the decoding phase. Alternatively, a new attribute (e.g., @id) can be introduced, which includes the ISO BMFF track ID containing the spatial partition SP. In either instance, the parsing of the chunk sampling process (after obtaining the content) can be the same.

[0192] For the second option (2) above, the BTD can be placed in a box (e.g., a box with a new entry type "btds") to indicate that the box contains the BTD. This box can be the Sample Description Box ("stsd") for the spatial partition SP.

[0193] Figure 12 An example of an ISO BMFF file containing a BTD box "btds" that defines a spatial partition SP is schematically shown.

[0194] Figure 13 An example of an ISO BMFF file containing a BTD box "btds" that defines two spatial partitions (the first spatial partition consists of chunk TLs, and the second spatial partition consists of two chunk TLs) is schematically shown.

[0195] An example implementation of a box containing BTS information can be as shown in Table 5 below:

[0196] Table 5

[0197]

[0198]

[0199] Within the "btds", there can be sampling entries of the bitstream containing the spatial partition SP.

[0200] This definition of the box can be extended to include other relevant parameters, such as group ID (for chunked NALUs), spatial partition width / height, spatial partition position, etc. Table 6 below shows an example of such a configuration box:

[0201] Table 6

[0202]

[0203]

[0204] In either case of Table 5 and Table 6, the PLT information can be stored in the HEVCConfigurationBox, which is a partial sample entry of HEVC (for example, the sample entry name is "hev1").

[0205] In a specific example, the btds box describes the spatial partition position within the frame of the current track. The btds box does not describe the spatial dependencies between the possible multiple tracks present in the same file.

[0206] Another way to encapsulate using a single track is by grouping the samples of the chunked TL. Different groups can be identified using the groupID, and their details can be incorporated into the SampleGroupDescriptionBox. Thus, the BTD can refer to the groups, and the btds box can be extended to include the groupID that will be referenced by the corresponding NAL unit, since each NALU has a NALUMapEntry indicating which group it belongs to.

[0207] In an exemplary embodiment, content (such as video picture PC1) arrives at a streaming service that has been packaged in a container CT1 (such as ISOBMFF) and is to be prepared for streaming. The streaming service can extract the BTD from the file format metadata (such as from an ISO BMFF box) and insert the partition metadata MD1 into the presentation descriptor (such as MPD) when creating the presentation descriptor.

[0208] Partition metadata MS1 in a separate description file

[0209] In an exemplary embodiment, in write step 104( Figure 9 ), the partition metadata MD1 (or at least a part thereof) is written to a separate descriptor file FL1, i.e., a descriptor file separate from the container CT1. This descriptor file can reference (or point to) the container CT1.

[0210] As a result of the encoding method 100, the encoded video picture data DT1 can thus be included in the container CT1, while the partition metadata MD1 is included in a separate descriptor file FL1. Then, this descriptor file FL1 can be transmitted to the client (or encoder), either together with the container CT1 or independently of the container CT1, for decoding purposes. As already pointed out, based on the partition metadata MD1 incorporated in the descriptor file FL1, the client can, for example, detect the presence of at least one spatial partition SP in the encoded video picture PC1 and determine whether to decode the encoded video picture data DT1 and / or how to decode the encoded video picture data DT1. Various implementations of this exemplary embodiment are described below.

[0211] In a specific example, the descriptor file FL1 is a Media Presentation Descriptor (MPD). This MPD can have various properties. The descriptor file FL1 can be, for example, a manifest file (e.g., in the case where the encoded video picture data DT1 is transmitted to the client by streaming).

[0212] In a specific example, the descriptor file FL1 defines the location (or identification) of one or more containers (also referred to as second containers) including the container CT1. In other words, the descriptor file FL1 defines the location (or identification) of one or more containers within a group of at least one container, where the container CT1 belongs to that group. In this instance, the descriptor file FL1 can, for example, include the respective identifiers (e.g., URLs) for identifying each of the one or more containers in the group.

[0213] In a specific example, the descriptor file FL1 defines the respective locations (or identifiers) of the partition metadata MD1 and one or more picture tiles TL of the spatial partition SP of the encoded video picture PC1.

[0214] Hybrid implementations are also possible, where the partition metadata MD1 is signaled to the client (or encoder) by using the container CT1 and the separate descriptor file FL1 as described above. In the exemplary embodiment, a first part of the structural metadata MS1 can thus be included (or incorporated) in the descriptor file FL1 separate from (or independent of) the container CT1, and a second part of the structural metadata MS1 can be included (or incorporated) in the container CT1.

[0215] The first part of the partition metadata MS1 (which is included in the descriptor file FL1) may, for example, contain flags (as described above) to indicate the presence of at least one spatial partition SP in the encoded video picture PC1. This flag enables the client (or encoder) to know that the spatial partitioning function of the present invention is being implemented. The client may read or extract the remaining partition metadata MD1 (such as all or part of the decoding characteristics PR2) from the container CT1.

[0216] Spatial partitioning

[0217] As already pointed out, the picture tile TL containing the spatial partition SP may be aligned with the boundaries of the spatial partitions in the video picture PC1. However, this may not always be the case. To have better encoding and decoding efficiency, larger tiles TL may be used, and thus a trade-off should be made between the size of the tile and the quality of the image.

[0218] To address this situation, the spatial region PR1 defined in the partition metadata MD1 for each spatial partition SP may indicate one or more dimensions of the spatial partition SP (such as DM1 and DM2, see Figure 10 ) and the reference point RF1 (such as the upper left corner) according to which the frame to be rendered should be measured.

[0219] As already pointed out, Figure 10 The figure shows an example where the area covered by 9 tiles TL at the center of the video picture PC1 is larger than the spatial partition SP8. The dimensions DM1 and DM2 associated with the reference point RF1 (which indicates a point in the video picture PC1) may thus be signaled as the spatial region PR1 in the partition metadata MD1. Once decoding is complete, the client may crop and discard the unrendered regions.

[0220] In an exemplary embodiment, the main stream (such as the bitstream BT1) contains a plurality of (semantic) views aligned with the tile boundaries, such that each tile TL may, for example, correspond to a participant in a video conference call. Thus, by signaling the presence of the spatial partition SP aligned with the view of the receiving client, the decoding requirements may be reduced, for example, by reducing the number of views to be rendered.

[0221] The same ability to signal the derived codec characteristics from the original bitstream to the new derived bitstream may also be applied to real-time transmission when the client negotiates the codec for the communication session.

[0222] Decoding

[0223] Figure 14Schematic block diagram showing steps 202-206 of method 200 for decoding video picture PC1 according to at least one exemplary embodiment of the present application. The steps of method 200 can be implemented by any client, namely media player (or decoder, or decoding device) DV2, which will be described in more detail in specific exemplary embodiments. For example, media player DV2 can correspond to a reading application or a computing device.

[0224] By way of example, assume that media player DV2 implements decoding method 200 to decode encoded video picture data DT1 into decoded video picture PC2. Hereinafter, it is considered that the encoded video picture data DT1 is generated as the output of method 100 as previously described. Therefore, all the information provided above regarding encoding method 100 can be applied to decoding method 200 in a similar manner ( Figure 14 ).

[0225] Specifically, the encoded video picture data DT1 includes an encoded version of video picture PC1. The decoded video picture PC2 constitutes a decoded version of video picture PC1, which may correspond to all or part of video picture PC1.

[0226] In reading step 202, partition metadata MD1 of at least one spatial partition SP of video picture PC1 is read from metadata structure MS1. As already pointed out, each spatial partition SP contains at least one picture block TL of video picture PC1. The partition metadata MD1 of each spatial partition SP defines the decoding characteristics PR1 of the video picture data of the at least one picture block TL of spatial partition SP.

[0227] All the relevant information provided above regarding encoding, especially all the relevant information regarding picture block TL, spatial partition SP, partition metadata MD1, metadata structure MS1 and descriptor file FL1 (if any), can be applied to decoding method 200 in a similar manner.

[0228] Specifically, the partition metadata MD1 of each spatial partition SP may also define a spatial region PR1 in video picture PC1.

[0229] The metadata structure MS1 containing the partition metadata MD1 can be provided to or accessed by media player DV2 in any suitable manner.

[0230] In determination step 204 ( Figure 14) In [the above], the supported spatial partitions SP are determined (or identified) by comparing the partition metadata MD1 of the at least one spatial partition SP (i.e., the partition metadata MD1 read in step 202) with the decoding capabilities of the media player DV2. These decoding capabilities can define the type and / or characteristics of the video content that the media player DV2 is capable of reading (e.g., in terms of quality and / or resolution and / or profile).

[0231] In decoding step 206, the video picture data DT1 of at least one (e.g., one, multiple, or all) picture chunk TL in the supported spatial partition SP is decoded into the decoded video picture PC2. To this end, the video picture data DT1 can be extracted from the container CT1 (e.g., from the bitstream BT1), which is generated by the encoder DV1 according to the encoding method 100 ( Figure 9 ). The container CT1 can be provided to the media player DV2 or accessed by the media player DV2 in any suitable way.

[0232] Once the decoding 206 is completed, the media player DV2 can render the decoded video picture PC2 in any suitable way, such as by displaying the corresponding video content on the display unit.

[0233] As already described, the metadata structure MS1 can be transmitted to the media player DV2 in various ways. All or part of the metadata structure MS1 can be included in the container CT1. Alternatively, all or part of the metadata structure MS1 can be included in a separate descriptor file FL1. In a specific example, part of the metadata structure MS1 is included in the separate descriptor file FL1, while the rest is included in the container CT1. The media player DV1 can thus access and read the descriptor file FL1 (if any) to obtain all or part of the partition metadata MD1. The media player DV2 can also extract all or part of the partition metadata MD1 from the container CT1 (e.g., from the bitstream transform descriptor (BTD) within the container CT1).

[0234] Figure 15 An exemplary implementation of the decoding method 200 implemented by the media player DV2 as previously described is schematically shown.

[0235] In the receiving step 302, the partition metadata MD1 (including the spatial region PR1 of at least one spatial partition SP and the decoding characteristics PR2) is received. To this end, the partition metadata MD1 can be extracted from the metadata structure MS1. The metadata structure MS1 can be independent of or a part of the container CT1 of the encoded video picture data DT1 as previously described.

[0236] In determination step 304, it is determined what media codec schemes the media player DV2 can support. To this end, the media player DV2 can, for example, query a non-transitory memory storing the supported media codec schemes.

[0237] In determination step 306, based on the result of determination step 304, it is determined whether the media player DV2 supports a codec scheme for decoding the complete video picture PC1 (its main content). To this end, the media player DV2 can compare the supported media codec schemes identified in determination step 304 with at least one codec scheme identified based on the partition metadata MD1 as being allowed to decode the complete video picture PC1.

[0238] If the media player DV2 supports a codec scheme for decoding the complete video picture PC1, the decoding method 300 ends, and decoding can be performed in any suitable manner to decode the complete video picture PC1, thereby allowing complete consumption of the content by the client.

[0239] However, if the media player DV2 does not support any codec scheme for decoding the complete video picture PC1, the decoding method 300 continues with determination step 308 ( Figure 14 ) corresponding to the previously described determination step 204 ( Figure 14 ).

[0240] In identification step 310, which picture block(s) TL of the video picture PC1 contain the supported spatial partition SP identified in step 308 is / are identified by parsing the bitstream BT1 within the container CT1 generated by the encoding method 100. To this end, the media player DV2 can obtain the bitstream BT1 in any suitable manner. The media player DV2 can, for example, obtain the bitstream BT1 as the stream with the most relevant spatial partition SP from among multiple streams.

[0241] In extraction step 312, the multiple picture blocks TL identified in step 310 are extracted from the bitstream BT1 contained in the container CT1.

[0242] In decoding step 314 (corresponding to the decoding step 206 in Figure 14 ), the identified picture blocks TL are decoded into the decoded video picture PC2.

[0243] System

[0244] Figure 16 FIG. shows a schematic block diagram illustrating an example of a system 1800 in which various aspects and exemplary embodiments are implemented.

[0245] System 1800 may be embedded as one or more devices, including various components described below. In various exemplary embodiments, System 1800 may be configured to implement one or more aspects described in the present application. For example, System 1800 is configured to perform an encoding method (e.g., encoding method 100) and / or a decoding method (e.g., decoding methods 200 and / or 300) according to any of the previously described exemplary embodiments. System 1800 may thus constitute an encoder DV1 and / or a media player DV2 in the sense of the present application.

[0246] Examples of equipment that may form all or part of System 1800 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or network servers), static cameras or video cameras, encoding or decoding chips, or any other communication device. The elements of System 1800 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of System 1800 may be distributed across multiple ICs and / or discrete components. In various exemplary embodiments, System 1800 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0247] System 1800 may include at least one processor 1810 configured to execute instructions loaded therein for implementing various aspects described, for example, in the present application. The processor 1810 may include an embedded memory, input / output interfaces, and various other circuits known in the art. System 1800 may include at least one memory 1820 (e.g., a volatile memory device and / or a non-volatile memory device). System 1800 may include a storage device 1840, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 1840 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0248] System 1800 may include an encoder / decoder module 1830 configured to, for example, process data to provide encoded / decoded video picture data, and the encoder / decoder module 1830 may include its own processor and memory. The encoder / decoder module 1830 may represent the (one or more) modules that may be included in a device as previously described to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1830 may be implemented as a separate element of System 1800 or may be incorporated within the processor 1810 as a combination of hardware and software known to those skilled in the art.

[0249] The program code to be loaded into the processor 1810 or the encoder / decoder 1830 to execute various aspects described in the present application may be stored in the storage device 1840 and subsequently loaded into the memory 1820 for execution by the processor 1810. According to various exemplary embodiments, during the performance of the processes described in the present application, one or more of the processor 1810, the memory 1820, the storage device 1840, and the encoder / decoder module 1830 may store one or more of various items. Such stored items may include but are not limited to video picture data, information data for encoding / decoding video picture data, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic processing.

[0250] In several exemplary embodiments, the memory internal to the processor 1810 and / or the encoder / decoder module 1830 can be used to store instructions and provide a working memory for the processing that can be performed during encoding and / or decoding in accordance with the present application.

[0251] However, in other exemplary embodiments, a memory external to the processing device (e.g., the processing device can be the processor 1810 or the encoder / decoder module 1830) is used for one or more of these functions. The external memory can be the memory 1820 and / or the storage device 1840, such as, for example, dynamic volatile memory and / or non-volatile flash memory. In several exemplary embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one exemplary embodiment, a fast external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), AVC, HEVC, EVC, VVC, AV1, etc.

[0252] As indicated in block 1890, input to the elements of the system 1800 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF portion that can receive, for example, an RF signal transmitted over the air by a broadcast device, (ii) composite input terminals, (iii) USB input terminals, (iv) HDMI input terminals, (v) a bus when the present invention is implemented in the automotive field, such as a Controller Area Network (CAN), a Controller Area Network Flexible Data-Rate (CAN FD), a FlexRay (ISO 17458), or an Ethernet (ISO / IEC 802-3) bus.

[0253] In various exemplary embodiments, the input device of block 1890 has associated respective input processing elements, as known in the art. For example, the RF section can be associated with elements necessary for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to within a band), (ii) down-converting the selected signal, (iii) band-limiting the band again to a narrower band to select a signal band that can be referred to as a channel in some exemplary embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various exemplary embodiments can include one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband.

[0254] In one set-top box embodiment, the RF section and its associated input processing elements can receive an RF signal transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired band.

[0255] Various exemplary embodiments re-order the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0256] Adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various exemplary embodiments, the RF section can include an antenna.

[0257] In addition, the USB and / or HDMI terminals can include respective interface processors for connecting the system 1800 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or within the processor 1810 when necessary. Similarly, various aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 1810 when necessary. The demodulated, error-corrected, and de-multiplexed stream can be provided to various processing elements, including, for example, the processor 1810 and the encoder / decoder 1830, which operate in conjunction with memory and storage elements to process the data stream as necessary for presentation on an output device.

[0258] Various components of system 1800 can be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 1890, such as an internal bus (including I2C bus), wiring, and printed circuit board known in the art, can be used to interconnect the various components and transfer data between them.

[0259] System 1800 can include a communication interface 1850 that enables communication with other devices via a communication channel 1851. The communication interface 1850 can include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 1851. The communication interface 1850 can include, but is not limited to, a modem or a network card, and the communication channel 1851 can be implemented, for example, within a wired and / or wireless medium.

[0260] In various exemplary embodiments, data can be streamed to system 1800 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these exemplary embodiments can be received via the communication channel 1851 and the communication interface 1850 suitable for Wi-Fi communication. The communication channel 1851 of these exemplary embodiments can generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.

[0261] Other exemplary embodiments can use a set-top box to provide streamed data to system 1800, and the set-top box delivers data through an HDMI connection of input block 1890.

[0262] Still other exemplary embodiments can use an RF connection of input block 1890 to provide streamed data to system 600.

[0263] The data being streamed can be used as a way for system 1800 to signal information such as partition metadata MD1 (as described previously). The data being streamed can constitute or contain all or part of a metadata structure MS1 carrying partition metadata DT1.

[0264] As described previously, signaling can be implemented in various ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. can be used to signal information to a corresponding decoder.

[0265] System 1800 can provide output signals to various output devices, including a display 1861, a speaker 1871, and other peripheral devices 1881. In various exemplary examples of the embodiments, the other peripheral devices 1881 can include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices based on the output providing function of system 1800.

[0266] In various exemplary embodiments, control signals can be communicated between system 1800 and the display 1861, the speaker 1871, or other peripheral devices 1881 using signaling of communication protocols such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0267] The output devices can be communicatively connected to system 1800 via dedicated connections through corresponding interfaces 1860, 1870, and 1880.

[0268] Alternatively, the output devices can be connected to system 1800 via a communication interface 1850 using a communication channel 1851. The display 1861 and the speaker 1871 can be integrated into a single unit in an electronic device (such as, for example, a television) with other components of system 1800.

[0269] In various exemplary embodiments, the display interface 1860 can include a display driver, such as, for example, a timing controller (T Con) chip.

[0270] For example, if the RF portion of the input terminal 1890 is part of a separate set-top box, then the display 1861 and the speaker 1871 can alternatively be separate from one or more of the other components. In various embodiments where the display 1861 and the speaker 1871 can be external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP output terminals.

[0271] In Figures 1 - 16 this document, various methods are described, and each method includes one or more steps or actions to implement the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0272] Some examples are described with respect to block diagrams and / or operational flowcharts. Each block represents a circuit element, a module, or a portion of code that includes one or more executable instructions for implementing the specified logic function(s). It should also be noted that in other embodiments, the function(s) noted in the block(s) may not occur in the order indicated. For example, depending on the functionality involved, two blocks shown in succession may actually be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order.

[0273] The embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed may be implemented in other forms (e.g., an apparatus or a computer program).

[0274] A method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.

[0275] In addition, a method may be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) may be stored on a computer-readable storage medium, such as, for example, storage device 1840( Figure 16 ). The computer-readable storage medium may take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executable by a computer implemented thereon. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, a computer-readable storage medium as used herein may be considered a non-transitory storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be recognized that the following, while providing more specific examples of computer-readable storage media to which this exemplary embodiment may be applied, are merely illustrative and not exhaustive as would be readily recognized by a person of ordinary skill in the art: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.

[0276] The instructions may form an application program tangibly implemented on a processor-readable medium.

[0277] For example, the instructions may be in hardware, firmware, software, or a combination thereof. For example, the instructions may be found in an operating system, a separate application, or a combination of both. Thus, a processor may be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of the instructions, the processor-readable medium may store data values generated by an implementation.

[0278] The apparatus may be implemented in, for example, suitable hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, and any other device for processing video frames, or other communication devices. It should be clear that the equipment may be mobile and even installed in a mobile vehicle.

[0279] Computer software may be implemented by the processor 1810 or by hardware, or by a combination of hardware and software. As a non-limiting example, the exemplary embodiments may also be implemented by one or more integrated circuits. The memory 1820 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 1810 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0280] As will be apparent to those of ordinary skill in the art based on this application, the implementations may generate various signals formatted to carry information such as may be stored or transmitted. The information may include, for example, instructions for executing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described exemplary embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted via various different wired or wireless links. The signal may be stored on a processor-readable medium.

[0281] The terms used herein are for the purpose of describing particular exemplary embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also be intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Also, when an element is referred to as "responsive to" or "connected to" or "associated with" another element, it may be directly responsive or connected to or "associated with" the other element, or there may be intervening elements. In contrast, when an element is referred to as "directly responsive to" or "directly connected to" another element, or "directly associated with" another element, there are no intervening elements.

[0282] It should be recognized that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any of the symbols / terms " / ", "and / or", and "at least one of" may be intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0283] Various numerical values may be used in this application. Specific values may be used for illustrative purposes and the aspects described are not limited to these specific values.

[0284] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No order is implied between the first element and the second element.

[0285] References to "an exemplary embodiment" or "exemplary embodiments" or "an implementation" or "implementations" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrases "in an exemplary embodiment" or "in exemplary embodiments" or "in an implementation" or "in implementations" and any other variations thereof that occur throughout this application do not necessarily all refer to the same embodiment.

[0286] Similarly, references herein to "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in connection with the exemplary embodiment / example / implementation) may be included in at least one exemplary embodiment / example / implementation. Thus, the appearances throughout this application of the expressions "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" do not necessarily all refer to the same exemplary embodiment / example / implementation, nor do the individual or alternative exemplary embodiments / examples / implementations have to be mutually exclusive of other exemplary embodiments / examples / implementations.

[0287] The reference numerals that appear in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, the exemplary embodiments / examples and variations can be employed in any combination or sub - combination.

[0288] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0289] Although some figures include arrows on communication paths to indicate the main direction of communication, it should be understood that communication can occur in a direction opposite to that of the depicted arrows.

[0290] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of the process of performing on a received video picture (which may include a bitstream encoding one or more video pictures) to produce a final output suitable for display or further processing in a reconstructed video domain. In various exemplary embodiments, such processes include one or more of the processes typically performed by a decoder. In various exemplary embodiments, for example, such processes also include or alternatively include the processes performed by the decoders of the various implementations described in this application.

[0291] As a further example, in one exemplary embodiment, "decoding" may refer only to dequantization, in one exemplary embodiment, "decoding" may refer to entropy decoding, in another exemplary embodiment, "decoding" may refer only to differential decoding, and in another exemplary embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed that this will be well understood by those skilled in the art.

[0292] Various implementations involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in the present application may cover, for example, all or part of the process performed on an input video picture to produce an encoded bitstream. In various exemplary embodiments, such processes include one or more of the processes typically performed by an encoder. In various exemplary embodiments, such processes also include or alternatively include the processes performed by the encoders of the various embodiments described in the present application.

[0293] As a further example, in one exemplary embodiment, "encoding" may refer only to quantization, in one exemplary embodiment, "encoding" may refer only to entropy encoding, in another exemplary embodiment, "encoding" may refer only to differential encoding, and in another exemplary embodiment, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and it is believed that this will be well understood by those skilled in the art.

[0294] In addition, the present application may refer to "obtaining" each piece of information. Obtaining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or the process of retrieving information from a memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, etc.

[0295] In addition, the present application may also refer to "receiving" each piece of information. Receiving information may include, for example, one or more of accessing information or receiving information from a communication network.

[0296] Moreover, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder, etc. For example, in certain exemplary embodiments, an encoder signals specific information, such as codec parameters or encoded video picture data. In this way, in exemplary embodiments, the same parameters can be used on the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicitly signal) specific parameters to a decoder such that the decoder can use the same specific parameters. Conversely, if a decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various exemplary embodiments. It should be recognized that signaling can be accomplished in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the previous discussion related to the verb form of the term "signal", the term "signal" can also be used as a noun herein.

[0297] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to yield other embodiments. In addition, those of ordinary skill in the art will understand that other structures and processes can replace the disclosed structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, the present application contemplates these and other embodiments.

Claims

1. A method for encoding a video picture (PC1), the method comprising: - spatially partitioning (102) the video picture into a plurality of picture chunks (TL) without intra-dependency; - writing (104) partition metadata (MD1) of at least one spatial partition (SP) into a metadata structure (MS1), each spatial partition comprising at least one picture chunk of the video picture, and the partition metadata of each spatial partition defining decoding characteristics (PR2) of the video picture data of the at least one picture chunk of the spatial partition; and - encoding (106) the video picture data (DT0) of the at least one picture chunk of the at least one spatial partition into a container (CT1) according to the partition metadata (MD1) of the at least one spatial partition (SP).

2. A method (200) for decoding a video picture (PC1) implemented by a media player (DV2); 300), the video picture is partitioned into a plurality of picture chunks (TL) without intra-dependency, the method comprising: - reading (202) partition metadata (MD1) of at least one spatial partition (SP) from a metadata structure (MS1), each spatial partition comprising at least one picture chunk (TL) of the video picture, and the partition metadata of each spatial partition defining decoding characteristics (PR2) of the video picture data of the at least one picture chunk of the spatial partition; - determining (204) a supported spatial partition (SP) by comparing the partition metadata (MD1) of the at least one spatial partition with the decoding capabilities of the media player (DV2); and - decoding (206) the video picture data (DT1) of at least one picture chunk of the supported spatial partition from a first container.

3. The method according to claim 2, wherein the partition metadata of the at least one spatial partition (SP) defines a spatial region (PR1) in the video picture.

4. The method according to claim 3, wherein the spatial region (PR1) is defined based on the partition metadata (MD1) of at least one picture chunk of the video picture.

5. The method according to claim 3 or 4, wherein the spatial region (PR1) is defined based on the partition metadata (MD1) of the at least one spatial partition (SP) as follows: at least one dimension of the spatial partition; and a reference point indicating the position of the spatial partition relative to the picture chunks of the video picture.

6. The method according to any one of claims 2 to 5, wherein at least part of the metadata structure is contained in a descriptor file, the descriptor file being separate from the first container and pointing to the first container.

7. The method according to claim 6, wherein the descriptor file is a media presentation descriptor (MPD).

8. The method according to claim 6 or 7, wherein the descriptor file defines: the positioning of one or more second containers including the first container.

9. The method according to any one of claims 6 to 8, wherein the descriptor file defines: the partition metadata (MD1) and positioning of one or more picture tiles (TL) of the at least one spatial partition (SP).

10. The method according to any one of claims 2 to 5, wherein at least part of the metadata structure is contained in the first container.

11. The method according to claim 10, wherein the partition metadata (MD1) is contained in at least one of the following: - in a bitstream conversion descriptor incorporated into the first container; and - in the metadata information of the first container.

12. The method according to any one of claims 6 to 11, wherein a first part of the metadata structure is contained in a descriptor file that is separate from and points to the first container, the first part of the metadata structure contains a flag that indicates the presence of the at least one spatial partition within the video picture, wherein a second part of the metadata structure is contained in the first container.

13. The method according to any one of claims 2 to 12, wherein the decoding characteristic (PR2) of the partition metadata defines: at least one codec parameter for decoding the at least one spatial partition.

14. The method according to any one of claims 2 to 13, wherein the partition metadata (MD1) contains at least one of the following: - an identifier of the at least one spatial partition; - chunk information indicating at least one picture tile of the at least one spatial partition; - the resolution of the at least one spatial partition; - region information defining the spatial region occupied by the at least one spatial partition; and - ratio information defining the aspect ratio of the at least one spatial partition.

15. The method according to any one of claims 2 to 14, wherein the first container (CT1) is an ISO BMFF file.

16. A data structure (MS1) formatted to include partition metadata (MD1) obtained from the method according to claim 1.

17. An apparatus (DV1) for encoding a video picture (PC) into a bitstream (BT1) of encoded video picture data (DT1), the apparatus comprising means for performing the method according to claim 1.

18. An apparatus (DV2) for video decoding a video picture (PC1) from a bitstream (BT1) of encoded video picture data (DT1), the apparatus comprising means for performing one of the methods according to any one of claims 2 to 15.

19. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform one of the methods according to any one of claims 1 to 15.

20. A non-transitory storage medium carrying instructions for executing program code for performing one of the methods according to any one of claims 1 to 15.