Spatially unequal streaming

By adopting spatial uneven encoding and decoding strategies in VR streaming, the quality uneven caused by viewport changes during user interaction is solved, and the user experience and bandwidth utilization efficiency are improved.

CN120358337APending Publication Date: 2025-07-22FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510525007.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-07-08
Filing Date
2017-10-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing VR streaming technology cannot respond quickly to viewport changes during user interaction, causing users to see a mixture of high-quality and low-quality areas in the viewport, affecting visible quality and bandwidth consumption.

Method used

By selecting and encoding media clips to stream video content in a spatially uneven manner based on predetermined relationships and signal prompts on the server side, it is ensured that the high-quality area always covers the user's observation section, and the decoder concentrates the decoding capabilities in the focus area, reducing computing complexity and bandwidth requirements.

Benefits of technology

It improves the user's visible quality in the VR environment, reduces the complexity and bandwidth consumption of streaming processing, and enhances the response speed to user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358337A_ABST
    Figure CN120358337A_ABST
Patent Text Reader

Abstract

Various concepts for streaming media content are described. Some concepts allow for streaming spatial scene content in a spatially unequal manner such that the user's visible quality is increased, or the processing complexity or necessary bandwidth at streaming extraction sites is reduced. Other concepts allow streaming of spatial scene content in a manner that increases applicability to other application contexts.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the applicant Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V., with an application date of October 11, 2017, an application number of 202210671217.X, and an invention title of "Spatial Unequal Streaming". Technical Field

[0002] This application relates to spatial unequal streaming such as occurs in virtual reality (VR) streaming. Background Art

[0003] VR streaming typically involves the transmission of extremely high-resolution videos. The resolving power of the human fovea is about 60 pixels per degree. Considering the transmission of a full sphere of 360°×180°, the transmission can be ended by sending a resolution of about 22k×11k pixels. Since sending this high resolution will result in extremely high bandwidth requirements, another solution is to only send the viewport shown at the head-mounted display (HMD), which has a field of view of 90°×90°, thus resulting in a video of about 6k×6k pixels. The compromise between sending the full video at the highest resolution and only sending the viewport is to send the viewport at a high resolution and send some adjacent data (or the rest of the spherical video) at a lower resolution or lower quality.

[0004] In the context of DASH, omnidirectional videos (also known as spherical videos) can be provided in the way that the DASH client controls the hybrid-resolution or hybrid-quality videos described previously. The DASH client only needs to know the information on how the content is provided.

[0005] An example can be to provide different representations with different projections, which have asymmetric characteristics, such as different qualities and distortions for different parts of the video. Each representation will correspond to a given viewport and will encode the viewport at a higher quality / resolution than the rest of the content. Knowing the orientation information (the direction of the viewport where the content has been encoded at a higher quality / resolution), the DASH client can dynamically select one or another representation to match the user's viewing direction at any time.

[0006] A more flexible option for the DASH client to select this asymmetric characteristic for omnidirectional videos is when the video is split into several spatial zones, where each zone can be obtained at a different resolution or quality. One option can be to split the video into rectangular zones (also known as tiles) based on a grid, but other options are foreseeable. In this case, the DASH client will need some signaling on the different qualities at which it provides different zones, and the DASH client can download different zones at different qualities, such that the quality of the viewport shown to the user is better than the other unshown content.

[0007] In any of the previous scenarios, when user interaction occurs and the viewport has changed, the DASH client takes some time to react to the user's movement and download content in a way that matches the new viewport. During the time between the user's movement and the DASH client adapting its requests to match the new viewport, the user will see some regions of high quality and low quality simultaneously in the viewport. Although the acceptable quality / resolution differences are content-dependent, the quality seen by the user is reduced in any case.

[0008] Therefore, a concept that would have the effect of alleviating or more effectively manifesting or even increasing the visible quality of the user with respect to the partial rendering of spatial scene content streamed via adaptive streaming could be beneficial. Summary of the Invention

[0009] Accordingly, an object of the present invention is to provide a concept for streaming spatial scene content in a spatially non-uniform manner such that the visible quality of the user is increased, or the processing complexity or the required bandwidth at the streaming extraction site is reduced, or to provide a concept for streaming spatial scene content in a way that increases its applicability to other application scenarios.

[0010] This object is achieved by the apparatus for streaming media content of a spatial scene with respect to temporal variations, the streaming server for media content of a spatial scene with respect to temporal variations, the media presentation description, the signal defining the media presentation description, the video bitstream, the decoder for decoding the video from the video bitstream, the streaming server for streaming media content of a spatial scene with respect to temporal variations, the streaming server for allowing an apparatus to stream media content of a spatial scene with respect to temporal variations from the server, the method for streaming media content of a spatial scene with respect to temporal variations, the method for decoding the video from the video bitstream, the method for allowing an apparatus to stream media content of a spatial scene with respect to temporal variations from the server, and the computer program having program code.

[0011] A first aspect of the present application is based on the following discovery: If the selected and extracted media segments and / or the signals obtained from the server provide the extraction device with a hint of a predetermined relationship to be followed by the quality used to encode different parts of a spatially varying scene over time into the selected and extracted media segments, streaming media content (such as video) of a spatially varying scene over time in a spatially non-uniform manner can be improved in terms of visible quality and / or computational complexity at the streaming reception site under comparable bandwidth consumption. Otherwise, the extraction device may not know in advance how the juxtaposition of parts encoded with different qualities into the selected and extracted media segments negatively affects the overall visible quality experienced by the user. The information contained in the media segments and / or the signals obtained from the server (such as in a manifest file (media presentation description) or in additional streaming-related control messages from the server to the client (such as SAND messages)) enables the extraction device to make an appropriate selection among the media segments provided at the server. In this way, virtual reality streaming or partial streaming of video content can become more robust with respect to quality degradation that occurs due to insufficient distribution of available bandwidth over this spatial section of the spatially varying scene presented to the user.

[0012] On the other hand, the present invention is based on the discovery that streaming media content (such as video) of a spatio-temporal scene that changes over time in a spatially non-uniform manner (such as using a first quality for a first part and a lower second quality or leaving a second part un-streamed in a second part) by determining the size and / or position of the first part depending on information contained in the media segment and / or a signal obtained from a server can result in an improvement in visible quality and / or the bandwidth consumption and / or computational complexity at the extraction side of the streaming becomes less complex. For example, for tile-based streaming, it is envisaged that a spatio-temporal scene that changes over time can be provided at the server in a tile-based manner, i.e., the media segment can represent a spectral-temporal part of the spatio-temporal scene, each of which can be a temporal segment of the spatio-temporal scene within a corresponding tile of the distribution of tiles into which the spatio-temporal scene is subdivided. In this case, the extraction device (client) makes a decision on how to distribute the available bandwidth and / or computational power in the spatio-temporal scene (i.e., at tile granularity). The extraction device can perform a selection of the media segment such that a first part of the spatio-temporal scene (which respectively follows an observation section that tracks the temporal changes of the spatio-temporal scene) is encoded into the selected and extracted media segment at a predetermined quality, which can be, for example, the highest quality achievable under current bandwidth and / or computational power conditions. For example, a second part of the spatio-temporal scene that is spatially adjacent may not be encoded into the selected and extracted media segment, or may be encoded into the media segment at another quality that is lower than the predetermined quality. In this case, counting the number of adjacent tiles is computationally complex or even infeasible, and the aggregation of these tiles completely covers the observation section of the temporal change, regardless of the orientation of the observation section. Depending on the projection selected to map the spatio-temporal scene onto individual tiles, the angular scene coverage of each tile can vary in this scene, and the fact that individual tiles can overlap each other even makes it more difficult to calculate the count of adjacent tiles that are sufficient to cover the observation section in terms of space (regardless of the orientation of the observation section). Therefore, in this case, the aforementioned information can indicate the size of the first part as the count N of tiles or the number of tiles, respectively. By this measure, the device will be able to track the observation section of the temporal change by selecting those media segments that have a co-located aggregation of N tiles encoded at a predetermined quality. The fact that the aggregation of these N tiles sufficiently covers the observation section can be ensured by the information indicating N. Another example can be the information contained in the media segment and / or the effect of the signal obtained from the server, which indicates the size of the first part relative to the size of the observation section itself. For example, this information can set a "safety zone" or prefetch zone around the actual observation section to some extent to account for the movement of the observation section of the temporal change. The greater the speed at which the observation section of the temporal change moves across the spatio-temporal scene, the larger the safety zone should be.Accordingly, the foregoing information can indicate the size of the first portion in a manner that varies with the size of the observation segment over time (such as in an incremental or scaled manner). An extraction device that sets the size of the first portion based on this information will be able to avoid quality degradation that could otherwise occur due to unextracted or low-quality portions of the spatial scene being visible in the observation segment. Here, it is irrelevant whether this scene is provided in a tile-based manner or in some other way.

[0013] Related to the aspect just mentioned of the present application, a video bitstream encoding a video can be decoded with increased quality, provided that the video bitstream has a signaling regarding the size of the focused region within the video, and the decoding capabilities for decoding the video should be concentrated on the focused region. By this measure, a decoder that decodes a video from the bitstream can concentrate or even limit its decoding capabilities for the decoded video to a portion having the size of the focused region signaled in the video bitstream, thereby knowing (for example) that the portion decoded in this way is decodable by the available decoding capabilities and spatially covers the desired segment of the video. For example, the size of such signaled focused region can be chosen to be large enough to cover the size of the observation segment and the movement of this observation segment, thereby taking into account the decoding latency when decoding the video. Or, in other words, the signaling of the recommended preferred observation segment region of the video contained in the video bitstream can allow the decoder to process this region in a better way, thereby allowing the decoder to concentrate its decoding capabilities accordingly. Whether or not region-specific decoding capability concentration is performed, the region signaling can be forwarded to the platform where it is selected which media segments are to be downloaded, i.e., where to place the quality-increased portions and how to size the quality-increased portions.

[0014] The first and second aspects of the present application are closely related to the third aspect of the present application. According to the third aspect, leveraging the fact that a large number of extraction devices stream media content from a server to obtain information, the information can subsequently be used to appropriately set the aforementioned type of information, thereby allowing the size or size and / or position of the first part to be set, and / or appropriately setting the predetermined relationship between the first quality and the second quality. Thus, according to this aspect of the present application, the extraction device (client) issues a log message that records one of the following: an instantaneous measurement result or statistical value of measuring the spatial position and / or movement of the first part; an instantaneous measurement result or statistical value of measuring the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section; and an instantaneous measurement result or statistical value of measuring the quality of the first part or the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section. The instantaneous measurement result and / or statistical value can be provided with time information related to the time when the corresponding instantaneous measurement result or statistical value was obtained. The log message can be sent to the server where the media segment is located, or sent to some other device that evaluates the incoming log message, in order to update the current setting of the aforementioned information used to set the size or size and / or position of the first part based on the log message, and / or to derive the predetermined relationship based on the log message.

[0015] According to another aspect of the present application, in particular, it is more effective in avoiding useless streaming trials by providing a media presentation description for streaming media content (such as video) of a spatio-temporal scene that changes over time in a tile-based manner. The media presentation description includes: at least one version, for tile-based streaming, the spatio-temporal scene that changes over time is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming of the corresponding version of the spatio-temporal scene. By this measure, the extraction device can match the beneficial requirements of the at least one version with the device capabilities of the extraction device itself or another device that interacts with the extraction device regarding tile-based streaming. For example, the benefit requirements may be related to decoding capability requirements. That is, if the decoding capability for decoding the streamed / extracted media content is not sufficient to decode all the media segments required to cover the observation section of the spatio-temporal scene that changes over time, then attempting to stream and present the media content will waste time, bandwidth, and computing power, and thus, it may be more effective not to attempt to stream and present the media content under any circumstances. For example, if (for example) the media segments related to a specific tile form a media stream (such as a video stream) separate from the media segments related to another tile, the decoding capability requirements may, for example, indicate the number of decoder instances required for the corresponding version. For example, the decoding capability requirements may also be regarding other information, such as a specific portion of the decoder instances necessary for a predetermined decoding profile and / or level, or may indicate a specific minimum capability of the user input device to move the viewport / section for viewing the scene fast enough. Depending on the scene content, low mobility may not be sufficient for the user to view the concerned part of the scene.

[0016] Another aspect of the present invention relates to an extension of streaming media content of a spatio-temporal scene that changes over time. In particular, the idea according to this aspect is that the spatio-temporal scene can actually not only change over time, but also change with respect to at least one other parameter (for example, view and position, viewing depth, or some other physical parameter). The extraction device can use adaptive streaming in this context by: calculating the addresses of media segments that describe the spatio-temporal scene that changes over time and in the at least one parameter depending on the viewport direction and at least one other parameter; and extracting the media segments from the server using the calculated positions.

[0017] The aspects outlined above in the present application and the advantageous implementations that are the subject of the dependent claims can be combined individually or together. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The preferred embodiments of the present application are described below with reference to the accompanying drawings, in which:

[0019] Figure 1 A demonstration schematic diagram, which illustrates a system of a client and a server for virtual reality applications as an example of a situation where the embodiments described in the following drawings can be advantageously used;

[0020] Figure 2 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment of the present application for describing possible operation modes of the client device, wherein the server 10 provides the device with information about acceptable or tolerable quality changes in the media content presented to the user;

[0021] Figure 3 Show Figure 2 modification, the part with increased quality does not care about the part of the viewing section of the tracking viewport, but cares about the region of interest of the media scene content signaled from the server to the client;

[0022] Figure 4 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment, wherein the server provides information on how to set the size or size and / or position of the part with increased quality, or the size or size and / or position of the actually extracted section of the media scene;

[0023] Figure 5 Show Figure 4 a variant, wherein the information sent by the server directly indicates the size of part 64, rather than scaling the part depending on the expected movement of the viewport;

[0024] Figure 6 Show Figure 4 a variant, according to which the extracted section has a predetermined quality and its size is determined by information from the server;

[0025] Figures 7a to 7c Show an illustration of Figure 4 and Figure 6 the manner in which information 74 increases the size of the part extracted with a predetermined quality by corresponding magnification of the size of the viewport;

[0026] Figure 8a Show a schematic diagram of an embodiment in which the client device sends log messages to the server or a specific evaluator for evaluating these log messages in order to (for example) obtain appropriate settings for the type of information discussed regarding Figures 2 to 7c ;

[0027] Figure 8bSchematic diagram of a tile - based cube projection of a 360 - degree scene onto tiles, and examples of how some of the tiles are covered by exemplary positions of the viewport. The small circles indicate positions in the isotropically - distributed viewports, and the shaded tiles are encoded in the downloaded segment at a higher resolution than the non - shaded tiles;

[0028] Figure 8c and Figure 8d Schematic diagram showing how the buffer fill levels (vertical axis) of different buffers of a client can evolve along a time axis (horizontal), where Figure 8c it is assumed that the buffer will be used to buffer the representation of a specific tile, and Figure 8d it is assumed that the buffer will be used to buffer the omnidirectional representation of a scene encoded into it with non - uniform quality (i.e., increasing in a certain direction specific to the corresponding buffer);

[0029] Figure 8e and Figure 8f 3D diagram showing different pixel density measurements within the viewport 28, differing in terms of uniformity in the sense of a sphere or an observation plane;

[0030] Figure 9 Block diagram of a client device, and a schematic illustration of a media segment selection procedure when the device detects information from a server to evaluate whether a particular version of tile - based streaming provided by the server is acceptable for the client device;

[0031] Figure 10 Schematic diagram showing a plurality of media segments provided by a server according to an embodiment, allowing the media scene to depend not only on time but also on another non - temporal parameter (i.e., here, exemplarily, the scene center position);

[0032] Figure 11 Schematic diagram showing a video bitstream containing information for manipulating or controlling the size of a focus area within a video encoded into the bitstream, and an example of a video decoder capable of utilizing this information. Detailed Description

[0033] For ease of understanding the description of the embodiments of the present application with respect to various aspects of the present application, Figure 1 an example of an environment in which the subsequently - described embodiments of the present application can be applied and advantageously used is shown. In particular, Figure 1Disclosed is a system consisting of a client 10 and a server 20 that interact via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) can be used for the communication 22 between the client 10 and the server 20. However, the embodiments outlined subsequently should not be construed as being limited to the use of DASH, and likewise, terms such as Media Presentation Description (MPD) should be understood broadly so as to also cover manifest files that are different from those in DASH.

[0034] Figure 1 Described is a system configured to implement a virtual reality application. That is, the system is configured to present, to a user wearing a head-up display 24 (i.e., via an internal display 26 of the head-up display 24), an observation segment 28 of a spatially varying scene 30 that changes over time, the segment 28 corresponding to the orientation of the head-up display 24 exemplarily measured by an internal orientation sensor 32 (such as an inertial sensor of the head-up display 24). That is, the segment 28 presented to the user forms a segment of the spatially varying scene 30, the spatial position of which corresponds to the orientation of the head-up display 24. In Figure 1 the case where the spatially varying scene 30 that changes over time is depicted as an omnidirectional video or a spherical video, but Figure 1 the description and the embodiments explained subsequently can also be easily transferred to other examples, such as presenting a segment in a video, where the spatial position of the segment 28 is determined by the intersection of face access or eye access with a virtual or real projection wall or the like. Additionally, the sensor 32 and the display 26 can be included in different devices respectively (e.g., a remote control and a corresponding television), or the sensor and the display can be part of a handheld device (such as a mobile device, such as a tablet computer or a mobile phone). Finally, it should be noted that some of the embodiments described later can also be applied to the situation where the area 28 presented to the user always covers the entire spatially varying scene 30 that changes over time, where the non-uniformity during the presentation of the spatially varying scene 30 is related to, for example, an unequal distribution of quality in the spatial scene.

[0035] Additional details regarding the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are described in Figure 1 and are described below. However, these details should not be considered as limiting the embodiments explained subsequently, but should actually serve as examples of how to implement any of the embodiments explained subsequently.

[0036] In particular, as Figure 1As shown in, server 20 may include a memory 34 and a controller 36, such as a suitably programmed computer, application specific integrated circuit, etc. The memory 34 has media segments stored thereon, and the media segments represent a spatially varying scene 30 that changes over time. Specific examples will be outlined in more detail in the description of Figure 1 . The controller 36 answers requests sent by the client 10 by re - sending the requested media segments to the client 10, and the media presentation description may send information about itself to the client 10. Details about this are also stated below. The controller 36 may extract the requested media segments from the memory 34. Other information may also be stored in this memory, such as a media presentation description or parts thereof, which are sent from the server 20 to the client 10 in other signals.

[0037] As Figure 1 shown in, the server 20 may optionally further include a stream modifier 38 that modifies media segments sent from the server 20 to the client 10 in response to a request from the client 10 so as to produce a media data stream at the client 10, the media data stream forming a single media stream that can be decoded by an associated decoder, but (for example) the media segments extracted by the client 10 in this way are actually aggregated from several media streams. However, the presence of this stream modifier 38 is optional.

[0038] Figure 1 The client 10 of is illustratively depicted as including a client device or controller 40 and one or more decoders 42 and a reprojection unit 44. The client device 40 may be a suitably programmed computer, microprocessor, programmed hardware device (such as an FPGA or application specific integrated circuit), etc. The client device 40 is responsible for selecting the segments to be extracted from the server 20 from among a plurality of 46 media segments provided at the server 20. For this purpose, the client device 40 first extracts a manifest or media presentation description from the server 20. From the manifest or media presentation description, the client device 40 obtains calculation rules for calculating the addresses of the media segments corresponding to a specific desired spatial portion of the spatially varying scene 30 among the plurality of 46 media segments. The client device 40 extracts the so - selected media segments from the server 20 by sending corresponding requests to the server 20. These requests contain the calculated addresses.

[0039] The media segments extracted by the client device 40 in this way will be forwarded by the client device 40 to one or more decoders 42 for decoding. In Figure 1In the example of, the media segments extracted and decoded in this way represent only the spatial section 48 in the spatial scene 30 that changes over time for each time unit, but as already indicated above, this may be different depending on, for example, the viewing section 28 to be presented that always covers other aspects of the entire scene. The reprojection unit 44 may optionally reproject the viewing section 28 to be displayed to the user and cut out the viewing section from the extracted and decoded scene content of the selected, extracted, and decoded media segments. For this purpose, as Figure 1 shown in, the client device 40 may, for example, continuously track the spatial position of the viewing section 28 and update the spatial position in response to user orientation data from the sensor 32, and notify the reprojection unit 44, for example, of this current spatial position of the viewing section 28 and the reprojection mapping to be applied to the extracted and decoded media content in order to be mapped to the area forming the viewing section 28. The reprojection unit 44 may accordingly apply the mapping and interpolation to, for example, a regular grid of pixels to be displayed on the display 26.

[0040] Figure 1 Illustrates the case where the spatial scene 30 has been mapped to the tiles 50 using cube mapping. The tiles are thus depicted as rectangular sub-regions of the cube onto which the scene 30 in the form of a sphere has been projected. The reprojection unit 44 reverses this projection. However, other examples may also be applied. For example, instead of cube projection, a projection onto a truncated cone or a non-truncated cone may be used. Furthermore, although Figure 1 the tiles are depicted as non-overlapping with respect to covering the spatial scene 30, the subdivision into tiles may involve mutual tile overlap. And as will be outlined in more detail below, it is also not mandatory for the scene 30 to be spatially subdivided into tiles 50 (as will be further explained below, each tile forms a representation).

[0041] Therefore, as Figure 1 depicted in, the entire spatial scene 30 is spatially subdivided into tiles 50. In Figure 1 the example of, each of the six faces of the cube is subdivided into 4 tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52, as Figure 1 depicted in. To be more precise, the server 20 provides more than one video 52 for each tile 50, and these videos have different qualities Q#. Even further, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# respectively form or are encoded into one of the media segments of the plurality of media segments 46 stored in the memory 34 of the server 20.

[0042] Even to emphasize again, Figure 1The examples of tile-based streaming media described herein are merely examples that may deviate significantly from it. For example, although Figure 1 may seem to indicate that the media segments of the higher-quality representation of scene 30 are for tiles that are consistent with the tiles to which the media segment belongs, and that the tiles encode scene 30 at quality Q1 therein, this consistency is not required and tiles of different qualities may even correspond to tiles of different projections of scene 30. Additionally, although not discussed so far, it is possible that Figure 1 the media segments corresponding to different quality levels depicted in

[0043] differ in spatial resolution and / or signal-to-noise ratio and / or temporal resolution, etc.

[0044] Finally, different from the tile-based streaming media concept (according to which the media segments that can be individually extracted from server 20 by device 40 are spatially subdivided into tiles 50 for scene 30), the media segments provided at server 20 may alternatively (for example) each encode scene 30 therein in a spatially complete manner with a spatially varying sampling resolution, where the sampling resolution reaches a maximum at different spatial positions in scene 30. For example, this situation can be achieved by providing at server 20 a sequence of segments 54 related to the projection of scene 30 onto a truncated cone, the frustum of which can be oriented in mutually different directions, resulting in differently oriented resolution peaks.

[0045] After the systems of server 20 and client 10 have been explained more generally, the functionality of client device 40 will be described in more detail with respect to an embodiment according to the first aspect of the present application. For this purpose, reference is made to Figure 2 which shows device 40 in more detail. As explained above, device 40 is used for streaming media content regarding a spatially varying scene 30 over time. As regarding Figure 1As explained, the apparatus 40 may be configured such that the streamed media content is spatially continuous with respect to the entire scene, or only with respect to a section 28 of the scene. In any case, the apparatus 40 includes: a selector 56 for selecting an appropriate media segment 58 from a plurality of 46 media segments available on the server 20; and an extractor 60 for extracting the selected media segment from the server 20 by a corresponding request (such as an HTTP request). As described above, the selector 56 may use the media presentation description to calculate the addresses of the selected media segments, and the extractor 60 uses these addresses when extracting the selected media segment 58. For example, the calculation rules indicated in the media presentation description for calculating the addresses may depend on the quality parameter Q, the tile T, and a certain time segment t. For example, the address may be a URL.

[0046] As also discussed above, the selector 56 is configured to perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene encoded therein. The spatial section may cover the entire scene spatially continuously. Figure 2 An illustrative case where the apparatus 40 adapts a spatial section 62 of the scene 30 to overlap and surround the viewing section 28 is illustrated at 61. However, as mentioned above, this may not necessarily be the case, and the spatial section may cover the entire scene 30 continuously.

[0047] In addition, the selector 56 performs the selection such that the selected media segment has a section 62 encoded therein in a spatially non-uniform quality manner. More precisely, a first part 64 of the spatial section 62 (at Figure 2The middle part (indicated by the shaded line) is encoded to the selected media segment with a predetermined quality. This quality can be, for example, the highest quality provided by the server 20 or can be a "good" quality. For example, the device 42 moves or adjusts the first part 64 in a manner that spatially follows the observation segment 28 that changes over time. For example, the selector 56 selects the current time segment 54 of those tiles that inherit the current position of the observation segment 28. After such selection, as explained below with respect to other embodiments, the selector 56 can optionally keep the number of tiles that make up the first part 64 constant. In any case, the second part 66 of the segment 62 is encoded to the selected media segment 58 with another quality, such as a lower quality. For example, the selector 56 selects the media segment corresponding to the current time segment of the tiles that are spatially adjacent to the part 64 and belong to the lower quality tiles. For example, in order to address the possible moment when the observation segment 28 moves too fast and leaves the part 64 before the end of the time interval corresponding to the current time segment and overlaps with the part 66 and the selector 56 will be able to reconfigure the part 64 spatially, the selector 56 mainly selects the media segment corresponding to the part 66. In this case, the part of the segment 28 that protrudes into the part 66 can still be presented to the user (i.e., with reduced quality).

[0048] The device 40 is not possible to evaluate for the user which negative quality degradation can be caused by presenting the content of the reduced-quality scenario to the user in advance and the content of the scenario within the higher-quality part 64. In particular, a transition between these two qualities that can be clearly visible to the user is generated. At least, this transition depends on the current scenario content within the segment 28 being visible. The severity of the negative impact of this transition within the user's field of view is a characteristic of the scenario content provided by the server 20 and may not be predicted by the device 40.

[0049] Therefore, according to Figure 2In an embodiment, the apparatus 40 includes a derivator 66 that derives a predetermined relationship that will be satisfied between the quality of portion 64 and the quality of portion 66. The derivator 66 derives this predetermined relationship from information that may be included in a media segment (such as within a delivery block in media segment 58) and / or included in a signaling from the server 20 (such as in a media presentation description or in an exclusive signal sent from the server 20 (such as within a SAND message, etc.)). Examples of what the information 68 may look like are presented below. The predetermined relationship 70 derived by the derivator 66 based on the information 68 is used by the selector 56 to perform the selection appropriately. For example, compared to a completely independent selection of the quality of portions 64 and 66, the restrictions in selecting the quality of portions 64 and 66 affect the distribution of the available bandwidth for extracting media content of section 62 onto portions 64 and 66. In any case, the selector 56 selects a media segment such that the quality used to encode portions 64 and 66 into the last extracted media segment satisfies the predetermined relationship. Examples of what the predetermined relationship may look like are also stated below.

[0050] The media segment selected and finally extracted by the extractor 60 is finally forwarded to one or more decoders 42 for decoding.

[0051] For example, according to a first example, the signaling mechanism embodied by the information 68 relates to information 68 that indicates to the apparatus 40 (which may be a DASH client) which quality combinations are acceptable for the provided video content. For example, the information 68 may be a list of quality pairs that indicates to the user or the apparatus 40 the maximum quality (or resolution) difference by which different regions 64 and 66 can be mixed. The apparatus 40 may be configured to inevitably use a particular quality level (such as the highest quality level provided at the server 10) for portion 64 and derive from the information 68 the quality level to be used to encode portion 66 into the selected media segment, where the information is included in the form of (for example) a list of quality levels for portion 68.

[0052] The information 68 may indicate a tolerable value for a measure of the difference between the quality of portion 68 and the quality of portion 64. The quality index of the media segment 58 may be used as a "measure" of the quality difference, and media segments may be distinguished in the media presentation description by the quality index, and the address of the media segment is calculated using the calculation rules described in the media presentation description by the quality index. In MPEG-DASH, the corresponding attribute indicating quality may be (for example) @qualityRanking. The apparatus 40 may consider the restrictions in the available quality level pairs that can be used to encode portions 64 and 66 into the selected media segment when performing the selection.

[0053] However, instead of this difference metric, the quality difference may alternatively be measured, for example, as a bitrate difference (i.e., a tolerable difference in the bitrates used to encode portions 64 and 66 into corresponding media segments, respectively), assuming that bitrate generally increases monotonically with increasing quality. Information 68 may indicate the allowed pairings of options for the qualities used to encode portions 64 and 66 into the selected media segment. Alternatively, information 68 simply indicates the allowed quality for encoding portion 66, thereby indirectly indicating an allowed or tolerable quality difference, assuming that main portion 64 is encoded using some default quality (such as the highest quality possible or achievable). For example, information 68 may be a list of acceptable representation IDs or may indicate a minimum bitrate level for the media segment associated with portion 66.

[0054] However, alternatively, a more gradual quality difference may be desired, wherein, instead of a quality pair, a quality group (more than two qualities) may be indicated, wherein, depending on the distance from the segment 28 (viewport), the quality difference may increase. That is, the information 68 may indicate tolerable values for a measure of the difference between the qualities of the portions 64 and 66 in a manner that depends on the distance from the observation segment 28. This may be done by a list of pairs of respective distances from the observation segment and corresponding tolerable values for a measure of the quality difference exceeding the respective distance. Below the respective distance, the quality difference must be lower. That is, each pair may indicate for a respective distance that a portion of the portion 66 that is further from the segment 28 than the respective distance may have a quality difference with the quality of the portion 64 that exceeds the respective tolerable value of this list entry.

[0055] The tolerable value may increase with increasing distance from the observation segment 28. The acceptability of the quality differences just discussed often depends on the time at which these different qualities are displayed to the user. For example, content with a high quality difference may be acceptable if the content is only displayed for 200 microseconds, while content with a lower quality difference may be acceptable if the content is displayed for 500 microseconds. Therefore, according to another example, in addition to the aforementioned quality combinations, or in addition to the allowed quality differences, the information 68 may also include time intervals in which the combination / quality difference may be acceptable. In other words, the information 68 may indicate the tolerable or maximum allowed difference between the quality of the portions 66 and 64, as well as an indication of the maximum allowed time interval in which the portion 66 may be displayed simultaneously with the portion 64 in the observation segment 28.

[0056] As previously mentioned, the acceptability of quality differences depends on the content itself. For example, the spatial position of different tiles 50 affects acceptability. Quality differences in a uniform background region with low-frequency signals are expected to be more acceptable than those in foreground objects. In addition, due to content changes, the temporal position also affects acceptability. Thus, according to another example, the signals forming information 68 are sent intermittently (such as, in each representation or period in DASH) to device 40. That is, the predetermined relationship indicated by information 68 can be updated intermittently. Additionally and / or alternatively, the signaling mechanism implemented by information 68 can vary in space. That is, information 68 can be made spatially dependent, such as, by the SRD parameter in DASH. That is, for different spatial regions of scene 30, different predetermined relationships can be indicated by information 68.

[0057] As described with respect to Figure 2 the embodiments of device 40 relate to the fact that device 40 wishes to keep the quality degradation due to pre-fetched portion 66 within the extracted portion 62 of video content 30 that is briefly visible in portion 28 as low as possible before being able to change the positions of portions 62 and 64 so as to adapt said portions to the position change caused by portion 28. That is, in Figure 2 portions 64 and 66 are different portions of portion 62, the quality of which is restricted until their possible combination is taken into account by information 68, and the transition between the two portions 64 and 66 is continuously shifted or adapted so as to track or overtake the moving viewing portion 28. According to Figure 3 the alternative embodiment shown in Figure 3 device 40 uses information 68 to control the possible combinations of the quality of portions 64 and 66, however, according to

[0058] the embodiments of Figure 4 portions 64 and 66 are defined as being different or distinct from each other in a manner defined, for example, in a media presentation description (i.e., in a manner independent of the position of viewing portion 28). The positions of portions 64 and 66 and the transition therebetween can be constant or vary in time. If they vary in time, such variations are due to changes in the content of scene 30. For example, portion 64 can correspond to a region of interest worthy of consuming higher quality, while portion 66 is the portion for which quality degradation should be considered first (e.g., due to low bandwidth conditions) before considering the quality degradation of portion 64.

[0058] In the following, another embodiment of an advantageous implementation of device 40 is described. In particular, Figure 4 device 40 is shown, which is structurally identical to Figure 2 and 3 but the operating mode is changed so as to correspond to the second aspect of the present application.

[0059] That is, device 40 includes selector 56, extractor 60, and inferrer 66. Selector 56 selects from among a plurality of media segments 58 provided by server 20, and extractor 60 extracts the selected media segment from the server. Figure 4 Assume that device 40 operates as depicted and illustrated with respect to Figure 2 and Figure 3 i.e., selector 56 performs the selection such that the selected media segment 58 encodes a spatial segment 62 of scene 30 in such a way that the spatial segment follows an observation segment 28 whose spatial position varies in time. However, a variant corresponding to the same aspect of the present application will subsequently be described with respect to Figure 5 wherein, for each time instant t, the selected and extracted media segment 58 encodes the entire scene or a constant spatial segment 62 therein.

[0060] In any case, similar to the description with respect to Figure 2 and 3 selector 56 selects media segment 58 such that a first portion 64 within segment 62 is encoded into the selected and extracted media segment at a predetermined quality, while a second portion 66 of segment 62 (which is spatially adjacent to first portion 64) is encoded into the selected media segment at a quality reduced relative to the quality of portion 64. A variant is described in Figure 6 where selector 56 restricts the selection and extraction of media segments with respect to a moving template to tracking the position of viewport 28, and wherein the media segment has a segment 62 encoded therein at a predetermined quality such that first portion 64 completely covers segment 62 while being surrounded by an uncoded portion 72. In any case, selector 56 performs the selection such that first portion 64 follows observation segment 28 whose spatial position varies in time.

[0061] In this case, it is also not easy for client 40 to predict how large segment 62 or portion 64 should be. Depending on the scene content, most users may perform similar actions when moving observation segment 28 across scene 30, and thus, the same actions apply to the intervals of observation segment 28, which is likely to move across scene 30 at approximately this speed. Therefore, according to the embodiment of Figures 4 to 6 information 74 is provided by server 20 to device 40 to assist device 40 in setting the size or size and / or position of first portion 64, or the size or size and / or position of segment 62, respectively, depending on information 74. Regarding the possibility of transmitting information 74 from server 20 to device 40, as described above with respect to Figure 2 and Figure 3The described situation applies. That is, the information may be included within the media segment 58, such as within the event block of the media segment, or for this purpose, a media presentation description or a transmission within an exclusive message (such as a SAND message) sent from the server to the device 40 may be used.

[0062] That is, according to Figures 4 to 6 the embodiment of, the selector 56 is configured to set the size of the first portion 64 depending on the information 74 originating from the server 20. In Figures 4 to 6 the embodiment illustrated, the size is set in units of tiles 50, but as described above with respect to Figure 1 it may be slightly different when using another concept of providing a spatially varying quality scene 30 at the server 20.

[0063] According to an example, the information 70 may (for example) include the probability of a given movement speed of the viewport of the viewing section 28. As already indicated above, the information 74 may cause a media presentation description to be available for the client device 40 (which may be, for example, a DASH client), or some in-band mechanism may be used to convey the information 74, such as an event block, that is, an EMSG or a SAND message in the case of DASH. The information 74 may also be included in any container format, such as the ISO file format or a transport format beyond MPEG-DASH (such as MPEG-2TS). The information may also be conveyed in the video bitstream (such as, in an SEI message as described later). In other words, the information 74 may indicate a predetermined value for the measure of the spatial speed for the viewing section 28. In this way, the information 74 indicates the size of the portion 64, either in the form of a scaling relative to the size of the viewing section 28 or in the form of an increment relative to the size of the viewing section 28. That is, the information 74 starts from the "basic size" of the portion 64 necessary to cover the size of the section 28 and appropriately (such as incrementally or proportionally) increases this "basic size". For example, the aforementioned movement speed of the viewing section 28 may be used to correspondingly scale the perimeter of the current position of the viewing section 28 in order to determine (for example) the farthest position of the perimeter of the viewing section 28 in any spatial direction feasible after this time interval, for example, to determine the latency when adjusting the spatial position of the portion 64, such as the duration of the time segment 54 corresponding to the time length of the media segment 58. The speed multiplied by this duration plus the perimeter of the current position of the omnidirectional viewport 28 may thus result in this worst-case perimeter and may be used to determine the magnification of the portion 64 relative to a certain minimum expansion of the portion 64 assuming a non-moving viewport 28.

[0064] The information 74 can even be about the evaluation of statistical data on user behavior. Subsequently, embodiments suitable for feeding such an evaluation program are described. For example, the information 74 can indicate the maximum speed regarding a certain percentage of users. For example, the information 74 can indicate that 90% of the users move at a speed lower than 0.2 radians per second and 98% move at a speed lower than 0.5 radians per second. The information 74 or the message carrying said information can be defined such that a probability-speed pair is defined or the message can be defined to signal the maximum speed of a fixed percentage of users (e.g., always 99% of the users). The movement speed signaling 74 can additionally include direction information, i.e., an angle in 2D, or depth in 2D plus 3D (also referred to as light field applications). The information 74 can indicate different probability-speed pairs for different movement directions.

[0065] In other words, the information 74 can be applied to a given time span, such as the time length of a media segment. The information can consist of a trajectory-based (x percentage, average user path) or speed-based pair (x percentage, speed) or distance-based pair (x percentage, pore / diameter / preferred) or area-based pair (x percentage, recommended preferred area) or a single maximum boundary value of a path, speed, distance, or preferably area. Instead of associating the information with a percentage, a simple frequency grading can be made according to the fact that most users move at a particular speed, the second most users move at another speed, and so on. Additionally or alternatively, the information 74 is not limited to indicating the speed of the observation segment 28, but can equally indicate the preferred regions to be observed separately, in order to guide the attempt to track parts 62 and / or 64 of the observation segment 28, with or without an indication of the statistical significance of the indication (such as the percentage of users who have complied with that indication or an indication of whether the indication is consistent with the most frequently recorded user observation speed / observation segment), and with or without an indication of the time duration of the indication. The information 74 can indicate another measure of the speed of the observation segment 28, such as a measure of the travel distance of the observation segment 28 over a particular time period (such as over the time length of a media clip, or more specifically over the time length of the time segment 54). Alternatively, the information 74 can be notified in a way that differentiates between the particular movement directions in which the observation segment 28 can travel. This applies both to indicating the rate or speed of the observation segment 28 in a particular direction and to indicating the travel distance of the observation segment 28 with respect to a particular movement direction. Additionally, the extension of part 64 can be signaled directly by the information 74 omnidirectionally or in a way that differentiates between different movement directions. Additionally, all of the examples outlined just above can be modified, where the information 74 indicates these values as well as the percentage of users for whom these values are sufficient to explain the statistical behavior when moving the observation segment 28. In this regard, it should be noted that the observation speed (i.e., the speed of the observation segment 28) can be quite large and is not limited to (for example) the speed value of the user's head. In fact, the observation segment 28 can move depending on (for example) the user's eye movements, in which case the observation speed can be significantly larger. The observation segment 28 can also move according to the movement of another input device (such as according to the movement of a tablet computer, etc.). Since all of these "input possibilities" that enable the user to move the segment 28 result in different expected speeds of the observation segment 28, the information 74 can even be designed such that the information differentiates between different concepts for controlling the movement of the observation segment 28. That is, the information 74 can indicate the size of part 64 in a way that indicates different sizes for different methods of controlling the movement of the observation segment 28, and the device 40 can use the size indicated by the information 74 for correct observation segment control.That is, the device 40 obtains knowledge of the way the viewing section 28 is controlled by the user, that is, checks whether the viewing section 28 is controlled by head movement, eye movement, or tablet computer movement or the like, and sets the size according to a part of the information 74 corresponding to such viewing section control.

[0066] Generally, the movement speed can be signaled according to content, time period, representation, segment, according to the SRD position, according to pixels, according to tiles (e.g., at any temporal or spatial granularity, etc.). The movement speed can also distinguish head movement and / or eye movement, as outlined just now. Additionally, the information 74 regarding the user movement probability can be conveyed as a recommendation regarding high-resolution prefetching (i.e., the video area outside the user's viewport, or the sphere coverage).

[0067] Figures 7a to 7c Briefly outline some of the options as explained regarding the information 74 in terms of its use by the device 40 to respectively modify the size and / or the position of the part 64 or the part 62. According to Figure 7a In the option shown in, the device 40 magnifies the perimeter of the section 28 by a distance corresponding to the product of the signaled speed v and the duration Δt, and the duration can correspond to a time period corresponding to the time length of the time segment 54 encoded in the individual media segment 50a. Additionally and / or alternatively, the greater the speed, the further the position of the part 62 and / or 64 can be placed away from the current position of the section 28, or the current position of the part 62 and / or 64 can be in the direction of the signaled speed or movement, as signaled by the information 74. The speed and direction can be derived from the newly developed or changed measurement or extrapolation of the recommended preferred area indicated by the information 74. Instead of applying v×Δt omnidirectionally, the speed can be signaled differently by the information 74 for different spatial directions. Figure 7b The alternative example depicted in shows that the information 74 can directly indicate the distance to magnify the perimeter of the viewing section 28, which is indicated by the parameter s in Figure 7b Again, magnification with a change in the direction of the section can be applied. Figure 7c Shows that the magnification of the perimeter of the section 28 can be indicated by an area increase by the information 74, such as in the form of the ratio of the area of the magnified section to the original area of the section 28. In any case, the perimeter of the area 28 after magnification (indicated by 76 in Figures 7a to 7c ) can be used by the selector 56 to size or set the size of the part 64 such that the part 64 covers at least the entire area within the magnified section 76 by a predetermined amount. Obviously, the larger the section 76, for example, the larger the number of tiles within the part 64. According to another alternative, the section 74 can directly indicate the size of the part 64, such as in the form of the number of tiles constituting the part 64.

[0068] In Figure 5A further possibility of signaling the size of the signaling part 64 is depicted. Figure 5 The embodiments of Figure 4 can be modified in a manner similar to the way the embodiments of Figure 6 are modified, i.e., the entire area of the section 62 can be retrieved from the server 20 via the segments 58 with the quality of the part 64.

[0069] In any case, at the end of Figure 5 , the information 74 differentiates between different sizes of the viewing section 28, i.e., different fields of view seen by the viewing section 28. The information 74 simply indicates the size of the part 64 depending on the size of the viewing section 28 at which the device 40 is currently aimed. This enables the service of the server 20 to be used by devices having different fields of view or different sizes of the viewing section 28 without a device such as the device 40 having to cope with calculating or otherwise guessing the size of the part 64 such that the part 64 is sufficient to cover the viewing section 28 regardless of any movement of the section 28 (as discussed with respect to Figure 4 , Figure 6 and FIG. 7). As may become clear from the description of Figure 1 , it is easy to evaluate (e.g.) which constant number of tiles may be sufficient to completely cover a particular size of the viewing section 28 (i.e., a particular field of view) regardless of the orientation of the viewing section 28 for spatial positioning 30. Here, the information 74 alleviates this situation, and the device 40 can simply look up the value of the size of the part 64 in the information 74 for the size of the viewing section 28 of the device 40. That is, according to the embodiments of Figure 5 , a media presentation description (such as an event chunk or a SAND message) available for a DASH client or some intended agency may include the information 74 regarding the sphere coverage or the field of view of a set of representations or a set of tiles, respectively. An example could be providing M representations of tiles as depicted in Figure 1 . The information 74 may indicate a recommended number n < M tiles (referred to as representations) to be downloaded for covering a given terminal device field of view. For example, in a cube representation tiled into 6×4 tiles as depicted in Figure 1 , 12 tiles are considered sufficient to cover a 90°×90° field of view. Due to the fact that the terminal device field of view may not always align perfectly with the tile boundaries, this recommendation cannot be trivially generated by the device 40 itself. The device 40 may use the information 74 by downloading (e.g.) at least N tiles, i.e., the media segment 58 is related to N tiles. Another way of using the information could be to focus on the quality of the N tiles closest to the current viewing center of the terminal device within the section 62, i.e., using the N tiles to constitute the part 64 of the section 62.

[0070] Regarding Figure 8a , embodiments regarding another aspect of the present application are described. Here,Figure 8a Show the client device 10 and the server 20, which communicate with each other according to any one of the possibilities described above with respect to Figure 1 Figures 3 to 7. That is, the device 10 can be embodied according to any one of the embodiments described with respect to Figure 2 Figures 3 to 7, or can simply operate in the manner described above without these details. However, advantageously, the device 10 is embodied according to any one of the embodiments described above with respect to Figure 1 Figures 3 to 7 or any combination thereof, and additionally inherits the operating mode now described with respect to Figure 2 Figures 3 to 7. In particular, the device 10 is understood internally as described above with respect to Figure 8a Figures 3 to 8, that is, the device 40 includes a selector 56, an extractor 60 and optionally includes a derivator 66. The selector 56 performs a selection for targeted unequal streaming, that is, selects media segments in such a way that the media content is encoded into the selected and extracted media segments in a way that the quality varies spatially and / or there are uncoded portions. However, in addition to this, the device 40 also includes a log message transmitter 80 that issues log messages recorded in (for example) the following to the server 20 or the evaluation device 82: Figure 2 Figures 3 to 8, that is, the device 40 includes a selector 56, an extractor 60 and optionally includes a derivator 66. The selector 56 performs a selection for targeted unequal streaming, that is, selects media segments in such a way that the media content is encoded into the selected and extracted media segments in a way that the quality varies spatially and / or there are uncoded portions. However, in addition to this, the device 40 also includes a log message transmitter 80 that issues log messages recorded in (for example) the following to the server 20 or the evaluation device 82:

[0071] Instantaneous measurement results or statistical values of the spatial position and / or movement of the first part 64,

[0072] Instantaneous measurement results or statistical values of the quality of the spatial scene measured until it is encoded into the selected media segment and until it is visible in the observation section 28, and / or

[0073] Instantaneous measurement results or statistical values of the quality of the first part or the quality of the spatial scene 30 measured until it is encoded into the selected media segment and until it is visible in the observation section 28.

[0074] The motivation is as follows.

[0075] In order to be able to derive statistical data, such as the areas of greatest interest or speed-probability pairs, as previously described, a reporting mechanism from the user is required. Additional DASH metrics for the statistical data defined in Appendix D of ISO / IEC 23009-1 are necessary.

[0076] One metric can be the field of view of the client as a DASH metric, where the DASH client sends the characteristics of the terminal device regarding the field of view back to the metric server (which can be the same as the DASH server or another server).

[0077] Keywords Type Description EndDeviceFoVH Integer Horizontal field of view of the end device, in degrees EndDeviceFoVV Integer Vertical field of view of the end device, in degrees

[0078] A metric can be a ViewportList, where the DASH client sends back to the metric server (which can be the same as the DASH server or another server) the viewports seen by each client in a timely manner. The instantiation of this message can be as follows.

[0079]

[0080] For viewport (region of interest) messages, the DASH client can be required to report when a viewport change occurs, possibly with a given granularity (to avoid or not avoid reporting very small movements) or a given periodicity. This message can be included in the MPD as an attribute @reportViewPortPeriodicity or as an element or descriptor. This message can also be indicated out-of-band, such as using a SAND message or any other means.

[0081] Viewports can also be signaled with respect to tile granularity.

[0082] Additionally or alternatively, log messages can report on other current scene-related parameters that change in response to user input, such as any of the parameters discussed below with respect to Figure 10 the current user distance from the scene center and / or the current viewing depth.

[0083] Another metric can be a ViewportSpeedList, where the DASH client indicates the speed of movement of a given viewport when a movement occurs.

[0084]

[0085] This message can be sent only when the client performs a viewport movement. However, as in the previous case, the server can indicate that the message should be sent only when the movement is significant. This configuration can be somewhat similar to @minViewportDifferenceForReporting, for signaling the size in pixels or degrees or any other quantity that needs to be changed for the message being sent.

[0086] Another important aspect of the VR-DASH service (where asymmetric quality as described above is provided) is to evaluate how quickly a user switches from an asymmetric representation or a set of unequal quality / resolution representations of a viewport to a more adequate another representation or set of representations of another viewport. Using this metric, the server can derive statistics that help it understand the relevant factors affecting QoE. This metric can look as follows.

[0087]

[0088]

[0089] Alternatively, the durations described previously can be given as an average value.

[0090]

[0091] For other DASH metrics, all such metrics may additionally have a time measured in which they are performed.

[0092] t Real-time The time at which the parameter is measured.

[0093] In some cases, it may occur that if unequal quality content is downloaded and poor quality (or a mixture of good and poor quality) is presented for a sufficient length of time (which may be only a few seconds), the user will be unhappy and leave the session. Under the condition of leaving the session, the user may send a message with the quality presented in the most recent x time interval.

[0094]

[0095] Alternatively, the maximum quality difference can be reported or the maximum and minimum quality of the viewport can be reported.

[0096] As becomes clear from the above description regarding Figure 8a it is advantageous for a tile-based DASH streaming service operator to be able to derive statistics that exemplify the client reporting mechanism described above in order to set up and optimize their service in a meaningful way (e.g., regarding resolution ratio, bitrate, and segment duration). Additional DASH metrics are stated below in addition to the metrics defined above and in addition to Appendix D of the document "ISO / IEC 23009-1:2014, Information technology--Dynamic adaptive streaming over HTTP (DASH)--Part 1: Media presentation description and segment formats".

[0097] It is envisioned that a tile-based streaming service uses a video with a cube projection as depicted in Figure 1 The reconstruction on the client side is in Figure 8bIt is described in [description], where the small circles 198 indicate the projection of the observation directions that are horizontally and vertically equiangularly distributed within the client's viewport 28 onto the two-dimensional distribution on the image regions covered by the individual tiles 50. The tiles marked with hatching indicate high-resolution tiles, thus forming the high-resolution portion 64, and the tiles 50 shown without hatching represent low-resolution tiles, thus forming the low-resolution portion 66. It can be seen that as the viewport 28 changes, the user is partially presented with low-resolution tiles because the resolution of each tile on the cube determined by the most recent update of the segment selection and download, and the projection plane or pixel array of the tiles 50 encoded into the downloadable segment 58 falls on the cube.

[0098] Although the above description actually generally (especially) indicates feedback or log messages that indicate the quality of the video presented to the user in the viewport, hereinafter, more specific and advantageous metrics applicable in this regard will be outlined. The metrics now described can be reported back from the client side and are called the effective viewport resolution. It is speculated that the metric indicates to the service operator the effective resolution in the client's viewport. In the case where the reported effective viewport resolution indicates a resolution at which the user is only presented with the resolution towards the low-resolution tiles, the service operator can accordingly change the tiling configuration, resolution ratio, or segment length to achieve a higher effective viewport resolution.

[0099] One embodiment can be the average pixel count in the viewport 28 measured in the projection plane, and the pixel array of the tiles 50 encoded into the segment 58 falls in the projection plane. The measurement can be differentiated with respect to the covered field of view (FoV) of the viewport 28 or specific to the horizontal direction 204 and the vertical direction 206. The following table shows possible examples of the appropriate syntax and semantics that can be included in the log message to signal the outlined viewport quality metrics.

[0100]

[0101] The decomposition in the horizontal and vertical directions can be stopped by alternatively using a scalar value of the average pixel count. Together with an indication of the aperture or size of the viewport 28, the average count can also be reported to the recipient of the log message (i.e., the evaluator 82), and the average count indicates the pixel density within the viewport.

[0102] It can be advantageous to reduce the field of view considered for the metric to be less than the field of view of the viewport actually presented to the user, thus excluding regions that are only for peripheral vision towards the boundaries of the viewport and therefore have no impact on the subjective quality perception. This alternative is illustrated by the dashed line 202, which encloses the pixels in this central section of the viewport 28. The report of the considered field of view 202 for the reported metric with respect to the total field of view of the viewport 28 can also be communicated to the log message recipient 82. The following table shows the corresponding extension of the previous example.

[0103]

[0104] According to another embodiment, instead of measuring the average pixel density by spatially averaging the mass in a uniform manner in the projection plane (as was the case in the examples actually described so far for the examples containing EffectiveFoVResolutionH / V), it is measured in a manner that non-uniformly weights this averaging for the pixels (i.e., the projection plane). The averaging can be performed in a spherically uniform manner. As an example, the averaging can be performed uniformly with respect to sample points distributed as in circle 198. In other words, the averaging can be performed by weighting the regional density with weights that decrease quadratically with increasing local projection plane distance and increase according to the sine of the local tilt of the projection relative to the line connecting to the viewport. The message can include an optional (flag-controlled) step size to adjust for the inherent oversampling in some available projections (such as the equirectangular projection), for example by using a uniform spherical sampling grid. Some projections do not have a large oversampling problem, and forcing a calculation to remove the oversampling can create unnecessary complexity issues. This must not be limited to the equirectangular projection. The report does not need to distinguish between horizontal and vertical resolutions, but can combine them. An example is given below.

[0105]

[0106]

[0107] In Figure 8e the application of equiangular uniformity in the averaging is illustrated by showing how points 302 (within the viewport 28) that are equiangularly horizontally and vertically distributed on a sphere 304 centered on the viewport 306 project onto the projection plane 308 of the tile (here a cube) in order to perform the averaging of the pixel density 308 of the pixels arranged in an array (by rows and columns) in the projection area, in order to set the local weights for the pixel density according to the local density of the projections 198 of the points 302 onto the projection plane. Figure 8f A very similar method is depicted in Figure 8f . Here, the points 302 are equally spaced in the viewport plane perpendicular to the viewing direction 312, i.e., horizontally and vertically uniformly distributed in rows and columns, and the projections onto the projection plane 308 define points 198, the local density of which controls the weights, and the local density pixel density 308 that varies due to high and low resolution tiles within the viewport 28 contributes to the averaging with these weights. In examples such as the above table of the latest, Figure 8f an alternative to Figure 8e the example depicted in

[0108] In the following, embodiments of another classification of log messages are described, which relate to a DASH client 10 having a plurality of media buffers 300, as Figure 8a illustrated illustratively in which the DASH client 10 forwards the downloaded segments 58 to subsequent decoding by one or more decoders 42 (compare Figure 1 ). The distribution of the segments 58 onto the buffers can be done in different ways. For example, the distribution can be made such that specific regions of the 360 video are downloaded separately from each other, or cached after being downloaded into separate buffers. The following examples illustrate different distributions by indicating about the following: which Figure 1 tiles T indexed #1 to #24 as shown in which the enantiomers have a total of 25 are encoded to which individual downloadable representations R#1 to #P in which quality Q out of qualities #1 to #M (1 being the best and M being the worst), and how these P representations R can be grouped into adaptive sets A indexed #1 to #S in the MPD (optional), and how the segments 58 of the P representations R can be distributed onto the buffers of buffers B indexed #1 to #N.

[0109]

[0110] Here, the representations can be provided at the server and announced in the MPD for download, each of these representations relating to one tile 50, i.e., a section of the scene. Representations that are related to one tile 50 but encode this tile 50 in different qualities can be outlined in optionally grouped adaptive sets, but precisely, this grouping is for association to the buffers. Thus, according to this example, for each tile 50, or in other words, for each viewport (observation section) encoded, there will be one buffer.

[0111] Another set of representations and distribution can be:

[0112]

[0113] According to this example, each representation can cover the entire area, but the high-quality area will be focused on one hemisphere, while the lower quality is used for the other hemisphere. Representations that only differ in the exact quality used in this way (i.e., evenly in the position of the higher-quality hemisphere) can be collected in one adaptive set and distributed onto the buffers (here illustratively, six) according to this characteristic.

[0114] Therefore, the following description assumes that this distribution to the buffers according to different viewport encodings (video sub-regions, such as tiles) associated with an adaptive set or the like is applied. Figure 8cDescribe the buffer fullness levels over time of two separate buffers (e.g., tile 1 and tile 2) in a tile-based streaming scenario, which has been described in the last but not least table. Enabling the client to report the fullness levels of all its buffers allows the service operator to correlate the data with other streaming parameters to understand the quality of experience (QoE) impact of its service settings.

[0115] Advantageously, the buffer fullness of multiple media buffers on the client side can be reported using metrics and identified and associated with the buffer type. For example, the association types are as follows:

[0116] ● Tile

[0117] ● Viewport

[0118] ● Region

[0119] ● Adaptation set

[0120] ● Representation

[0121] ● Low-quality version of the full content

[0122] An embodiment of the present invention is given in Table 1, which defines the metrics for reporting buffer level status events for each identified and associated buffer.

[0123] Table 1: List of buffer levels

[0124]

[0125]

[0126] Another embodiment using viewport-dependent encoding is as follows.

[0127] In a viewport-dependent streaming scenario, the DASH client downloads and pre-buffers several media segments related to a specific viewing orientation (viewport). If the amount of pre-buffered content is too high and the client changes its viewing orientation, the part of the pre-buffered content that will be played after the viewport change is not presented and the corresponding media buffer is cleared. This scenario is depicted in Figure 8d .

[0128] Another embodiment can be related to a traditional video streaming scenario with multiple representations (quality / bitrate) of the same content and the quality used for encoding the video content can be spatially uniform.

[0129] The distribution can thus look as follows:

[0130]

[0131] That is, here, each representation can cover, for example, a complete scene that may not be a panoramic 360 scene with different qualities (i.e., spatially uniform qualities), and these representations can be individually distributed to the buffer. All examples stated in the last three tables should be considered as not restricting the way of distributing the segments 58 of the representations provided at the server to the buffer. There are different methods, and the rules can be based on the membership of the segment 58 to the representation, the membership of the segment 58 to the adaptive set, the direction of the locally increased quality encoding the spatially non-uniform scene into the representation to which the corresponding segment belongs, the quality used to encode the scene into the corresponding segment to which it belongs, etc.

[0132] The client can maintain a buffer for each representation and, after experiencing an increase in the available throughput, decide to clear the remaining low-quality / bitrate media buffer before playback and download high-quality media segments with a duration into the existing low-quality / bitrate buffer. Similar embodiments can be constructed for tile-based streaming and viewport-dependent encoding.

[0133] The service operator may not be interested in understanding how much and what kind of data is downloaded without presenting, because this introduces a non-beneficial cost on the server side and reduces the quality on the client side. Therefore, the present invention should provide a reporting metric that correlates the two events "media download" and "media presentation" for easy interpretation. The present invention avoids analyzing the information about the download and playback status of each media segment reported at length and only allows efficient reporting of the clearing event. The present invention also includes the identification of the buffer as described above and the association to the type. Embodiments of the present invention are given in Table 2.

[0134] Table 2: List of clearing events

[0135]

[0136]

[0137] Figure 9 Another embodiment showing how the device 40 can be advantageously implemented Figure 9 The device 40 can correspond to any of the examples stated above with respect to Figure 1 to FIG. 8. That is, the device may include a log transmitter as described above with respect to Figure 8a but not necessarily, and can use the information 68 as described above with respect to Figure 2 and Figure 3 or the information 74 as described above with respect to Figures 5 to 7c but not necessarily. However, different from the description of Figure 2 to FIG. 8, with respect to Figure 9, assume that the tile - based streaming method is truly applied. That is, the scene content 30 is provided at the server 20 in a tile - based manner, and the tile - based manner is discussed as the option regarding Figure 2 to FIG. 8.

[0138] Although the internal structure of the device 40 may be different from Figure 9 the internal structure depicted in Figure 2 to FIG. 8, the device 40 is illustratively shown as including the selector 56 and the extractor 60 that have been discussed above regarding Figure 9 to FIG. 8, and optionally includes the inferencer 66. However, additionally, the device 40 includes a media presentation description analyzer 90 and a matcher 92. The MPD analyzer 90 is used to derive from the media presentation description obtained from the server 20: at least one version, for tile - based streaming, in which the spatio - temporal scene 30 is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for benefiting from the spatio - temporal scene in the corresponding version of the tile - based streaming. The meaning of "version" will become clear from the following description. In particular, the matcher 92 matches the beneficial requirements so obtained with the device capabilities of the device 40 or another device that interacts with the device 40 (such as, the decoding capabilities of one or more decoders 42, the number of decoders 42, or the like).The background or idea implicit in the concept is as follows. Imagine, assume a specific size of the viewing section 28. Based on the tile-based approach, a specific number of tiles need to be included by the section 62. Additionally, it can be assumed that the media segments belonging to a specific tile form a media stream or video stream, and the stream should be decoded by a separate decoding instance, separately from decoding the media segments belonging to another tile. Thus, the movement aggregation of a specific number of tiles within the section 62 where the corresponding media segments are selected by the selector 56 requires a specific decoding capability, such as the existence of corresponding decoding resources in the form of, for example, a corresponding number of decoding instances (i.e., a corresponding number of decoders 42). If this number of decoders does not exist, the service provided by the server 20 is not available to the client. Therefore, the MPD provided by the server 20 can indicate the "useful requirement", that is, the number of decoders required to use the provided service. However, the server 20 can provide MPDs for different versions. That is, different MPDs for different versions can be obtained by the server 20, or the MPD provided by the server 20 can be structured internally to distinguish the different versions that can be served and used. For example, the versions can differ in the field of view (i.e., the size of the viewing section 28). Different sizes of the field of view manifest themselves as different numbers of tiles within the section 62, and thus can differ in useful requirements because (for example) these versions may require different numbers of decoders. Other examples can also be imagined. For example, although versions with different fields of view may involve the same number of media segments 46, according to another example, the differences between different versions of the scenario 30 provided for tile-based streaming at the server 20 can even lie in the number of 46 media segments involved according to the corresponding version. For example, the tile segmentation according to one version is coarser compared to the tile-hyphenation segmentation of the scenario according to another version, thus requiring (for example) a smaller number of decoders.

[0139] The matcher 92 matches the useful requirement and thus selects the corresponding version or completely rejects all versions.

[0140] However, the useful requirement can additionally focus on the profile / level that one or more decoders 42 must be able to handle. For example, the DASH MPD includes multiple locations that allow indicating the profile. A typical profile describes the attributes, elements that can exist in the MPD, and the video or audio profile for each representation of the provided media stream.

[0141] Other examples of beneficial requirements concern, for example, the ability on the client side to move the viewport 28 across scenes. The beneficial requirement may indicate a required viewport speed that should be available to the user to move the viewport so as to be able to truly enjoy the provided scene content. The matcher may check, for example, whether this requirement is met, e.g., inserted in a user input device such as the HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in a directional sense, the set of "sufficient types of input devices" may be indicated by the beneficial requirement.

[0142] In the tile streaming service of spherical video, there are an excessive number of configuration parameters that can be set dynamically, such as the number of qualities, the number of tiles. In the case where the tiles are independent bitstreams that need to be decoded by separate decoders, if the number of tiles is too high, a hardware device with several decoders will not be able to decode all the bitstreams simultaneously. The possibility is to keep this as a degree of freedom, and the DASH device parses all possible representations and counts how many decoders are needed to decode all the representations or the given number of the field of view of the covering device, and thus determines whether the DASH client is likely to consume the content. However, a smarter solution for interoperability and capability negotiation is to use signaling in the MPD mapped to a profile, which is used as a commitment to the client: if the profile is supported, the provided VR content can be consumed. This signaling should be in the form of a URN (such as urn::dash-mpeg::vr::2016) that can be encapsulated at the MPD level or at the adaptation set. This parsing will mean that N decoders at profile X are sufficient to consume the content. Depending on the profile, the DASH client may ignore or accept the MPD or part of the MPD (adaptation set). Additionally, there are several mechanisms that do not include all the information, such as Xlink or MPD links, where little signaling for selection is available. In this case, the DASH client will not be able to determine whether it can consume the content. It is necessary to expose the decoding capabilities regarding the number of decoders and the profile / level of each decoder through this urn (or something similar) so that the DASH client can now make sense of whether to perform an Xlink or MPD link or a similar mechanism. The signaling may also mean different operating points, such as N decoders with profile X / level or Z decoders with profile Y / level.

[0143] Figure 10 Further illustration, regarding the Figures 1 to 9 Any of the above-described embodiments and descriptions presented by the client, device 40, server, etc. can be extended to the following scope: the provided service is extended to a scope where the spatio-temporal scene that changes over time not only changes over time but also depends on another parameter. For example, Figure 10 Illustration Figure 1variant, where a plurality of available media segments are obtained on a server to describe scene content 30 for different positions of the viewing center 100. In Figure 10 In the schematic diagram shown in, the scene center is depicted as varying only along one direction X, but obviously, the viewing center can vary along more than one spatial direction (such as two-dimensionally or three-dimensionally). For example, this corresponds to a change in the user's location in a specific virtual environment. Depending on the user's location in the virtual environment, the available view changes, and correspondingly, the scene 30 changes. Thus, in addition to describing that the scene 30 is subdivided into tiles and time segments and media segments of different qualities, other media segments describe different content of the scene 30 for different positions of the scene center 100. The device 40 or the selector 56 calculates the addresses of the media segments to be extracted within the selection process from among the plurality of 46 media segments respectively depending on the viewing section position and at least one parameter (such as parameter X), and can then use the calculated addresses to extract these media segments from the server. For this purpose, the media presentation description can describe a function depending on the tile index, the quality index, the scene center position, and time t, and generate the corresponding addresses of the corresponding media segments. Thus, according to Figure 10 the embodiment of, the media presentation description can include this calculation rule, in addition to the parameters described above with respect to Figures 1 to 9 which, the calculation rule also depends on one or more additional parameters. Parameter X can be quantized to any level in a hierarchy, for which level the corresponding scene representation is encoded by the corresponding media segment within the plurality of 46 media segments in the server 20.

[0144] As an alternative, X can be a parameter defining the viewing depth (i.e., the distance radially from the scene center 100). Although providing scenes in different versions with different X values in the viewing center part allows the user to "walk" through the scene, providing scenes in different versions with different viewing depths can allow the user to "radially zoom" back and forth in the scene.

[0145] For a plurality of non-concentric viewports, the MPD can thus be signaled using another signal of the position of the current viewport. The signaling can be done at the segment, representation, or period level or the like.

[0146] Non-concentric spheres: To enable user movement, the spatial relationship of different spheres should be signaled in the MPD. This signaling can be done through coordinates (x, y, z) in any unit relative to the sphere diameter. Additionally, the diameter of each sphere should be indicated. The sphere can be "good enough" for a user at its center and the additional space (for which the content will be good). If the user can move beyond the signaled diameter, another sphere should be used for presenting the content.

[0147] Exemplary signaling of the viewport can be relative to a predefined center point in space. Each viewport will be signaled relative to that center point. In MPEG-DASH, this can be signaled (for example) in the AdaptationSet element.

[0148]

[0149] Finally, Figure 11 information that describes information such as or similar to that described above with respect to reference numeral 74 can reside in video bitstream 110, into which video 112 is encoded. Decoder 114 that decodes this video 110 can use information 74 to determine the size of the focused region 116 within video 112, to which the decoding capabilities for decoding video 110 should be concentrated. For example, information 74 can be conveyed within the SEI information of video bitstream 110. For example, the focused region can be decoded specifically, or decoder 114 can be configured to start decoding each image of the video at the focused region, rather than (for example) at the upper left image corner, and / or decoder 114 can stop decoding each image of the video after the focused region 116 has been decoded. Additionally or alternatively, information 74 can be present in the data stream for use only in forwarding to subsequent renderers or viewport controls or streaming media devices of the client or segment selector, for deciding which segments to download or stream in order to cover a spatial section completely or with increased or predetermined quality. For example, as outlined above, information 74 indicates a recommended preferred region as a recommendation for placing viewing segment 62 or segment 66 to coincide with or cover or track this region. Information 74 can be used by the segment selector of the client. Just as is true with respect to the description of Figures 4 to 7c information 74 can set the size of region 116 absolutely (such as in terms of the number of tiles), or can set, for example, the speed of region 116 used for region movement according to user input, etc., in order to follow the content of interest in the video spatiotemporally, thereby scaling region 116 to increase as the indication of speed increases.

[0150] A first aspect of the present application provides an apparatus for streaming media content regarding a spatially varying scene (30) over time.

[0151] (A1) The apparatus is configured to:

[0152] select media segments from among a plurality of media segments available on a server,

[0153] the selected media segments being from the server, wherein the apparatus is configured to:

[0154] Perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene with the time variation that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and

[0155] derive the predetermined relationship from information included in the selected media segment and / or from a signal obtained from the server.

[0156] (A2), The apparatus according to (A1), wherein each of the plurality of media segments has an associated spatio-temporal part of the spatially varying scene with the time variation encoded therein at an associated quality level within a set of quality levels.

[0157] (A3), The apparatus according to (A1), wherein each of the spatio-temporal parts of the spatially varying scene (30) encoded into the plurality of media segments is a time segment of the spatially varying scene at a corresponding one of tiles (50) into which the spatially varying scene is spatially subdivided.

[0158] (A4), The apparatus according to any one of (A1)-(A3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.

[0159] (A5), The apparatus according to (A4), wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the viewing section.

[0160] (A6), The apparatus according to (A4) or (A5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the viewing section by a list of pairs of the respective distance from the viewing section and the corresponding tolerable value for the measure of the difference beyond the respective distance.

[0161] (A7), The apparatus according to any one of (A4)-(A6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.

[0162] (A8), The apparatus according to any one of (A4)-(A7), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval during which the second part can be together with the first part within the observation section.

[0163] (A9), The apparatus according to any one of (A8), wherein the information indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval during which the second part can be together with the first part within the observation section.

[0164] (A10), The apparatus according to any one of (A4)-(A9), wherein the information is time-varying and / or spatially varying.

[0165] (A11), The apparatus according to any one of (A1)-(A10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.

[0166] (A12), The apparatus according to any one of (A1)-(A11), wherein the apparatus is configured to perform the selection such that the first part follows a time-varying observation section of the time-varying spatial scene.

[0167] (A13), The apparatus according to (A12), wherein the apparatus is configured to cause the spatial position of the time-varying observation section to change according to a user input.

[0168] (A14), The apparatus according to any one of (A1)-(A13), wherein the apparatus is configured to determine the first part to correspond to a region of interest.

[0169] (A15), The apparatus according to any one of (A14), wherein the apparatus is configured to extract information about the region of interest from the server.

[0170] The present invention also provides a streaming server for media content of a time-varying spatial scene.

[0171] (B1) The streaming server is configured to:

[0172] Make multiple media segments available for extraction by a device, enabling the device to select a media segment for extraction, where the media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and

[0173] Signal information about a predetermined relationship in the media segment and / or by signaling to the device, where the another quality satisfies the predetermined relationship with respect to the predetermined quality.

[0174] (B2), The streaming media server according to (B1), wherein each of the multiple media segments has an associated spatio-temporal part of the spatially varying scene with the time variation encoded therein at an associated quality level in a set of quality levels.

[0175] (B3), The streaming media server according to (B2), wherein each of the spatio-temporal parts of the spatially varying scene encoded into the multiple media segments is a time segment of the spatially varying scene at a corresponding one of the tiles into which the spatially varying scene is spatially subdivided.

[0176] (B4), The streaming media server according to any one of (B1)-(B3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.

[0177] (B5), The streaming media server according to (B4), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from an observation section.

[0178] (B6), The streaming media server according to (B4) or (B5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section by a list of pairs of a corresponding distance from the observation section and a corresponding tolerable value for the measure of the difference beyond the corresponding distance.

[0179] (B7), The streaming media server according to any one of (B4)-(B6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.

[0180] (B8), The streaming media server according to any one of (B4)-(B7), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second part can be together with the first part within the observation section.

[0181] (B9), The streaming media server according to (B8), wherein the information indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval within which the second part can be together with the first part within the observation section.

[0182] (B10), The streaming media server according to any one of (B4)-(B9), wherein the information is time-varying and / or spatially varying.

[0183] (B11), The streaming media server according to any one of (B4)-(B10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.

[0184] (B12), The streaming media server according to any one of (B4)-(B11), wherein the server is configured to send information about the region of interest to the device.

[0185] This application also provides a media presentation description.

[0186] (C1), The media presentation description includes:

[0187] Information about calculating addresses of a plurality of media segments such that a device can use the information to select and extract a media segment from the plurality of media segments, wherein the media segment has at least a spatial section of the spatially varying temporal scene encoded therein, and is encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying temporal scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and

[0188] Information about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.

[0189] (C2), The media presentation description according to (C1), wherein each media segment of the plurality of media segments has an associated spatio-temporal part of the spatially varying temporal scene encoded therein at an associated quality level within a set of quality levels.

[0190] (C3), the media presentation description according to (C2), wherein each of the spatio-temporal parts of the spatially varying scene changing over time encoded into the plurality of media segments is a temporal segment of the spatially varying scene changing over time at a respective one of the tiles into which the spatially varying scene changing over time is spatially subdivided.

[0191] (C4), the media presentation description according to any one of (C1)-(C3), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.

[0192] (C5), the media presentation description according to (C4), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section.

[0193] (C6), the media presentation description according to (C4) or (C5), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section by a list of pairs of the respective distances from the viewing section and the corresponding tolerable values for the measure of the difference beyond the respective distances.

[0194] (C7), the media presentation description according to any one of (C4)-(C6), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.

[0195] (C8), the media presentation description according to any one of (C4)-(C7), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval within which the second part can be together with the first part within the viewing section.

[0196] (C9), the media presentation description according to (C8), wherein the information indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of the maximum allowable time interval within which the second part can be together with the first part within the viewing section.

[0197] (C10), the media presentation description according to any one of (C4)-(C9), wherein the information is time-varying and / or spatially varying.

[0198] (C11), a media presentation description according to any one of (C4)-(C10), wherein the information indicates an allowed concurrent setting pair for the another quality and the predetermined quality.

[0199] (C12), a media presentation description according to any one of (C4)-(C11), wherein the media presentation description includes information about a region of interest.

[0200] This application also provides an apparatus for streaming media content of a spatial scene varying over time.

[0201] (D1) The apparatus is configured to:

[0202] select media segments from a plurality of media segments available on a server,

[0203] wherein the apparatus is configured to:

[0204] perform the selection such that the selected media segments have at least a spatial section of the spatial scene varying over time encoded therein, encoded in a manner that:

[0205] a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatial scene varying over time that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and

[0206] such that the first part follows an observation section varying over time of the spatial scene varying over time, and

[0207] set the size and / or position of the first part depending on information included in the selected media segment and / or a signal received from the server.

[0208] (D2), the apparatus according to (D1), wherein the information indicates the size in the form of an increment relative to the size of the observation section varying over time or a scaling of the size of the observation section varying over time.

[0209] (D3), the apparatus according to (D1) or (D2), wherein the information indicates a predetermined value of a measure of the spatial speed for the observation section.

[0210] (D4), the apparatus according to (D3), wherein the information indicates the predetermined value of the measure of the spatial speed for the observation section for:

[0211] a default percentage of users whose measured spatial speed does not exceed the predetermined value, and / or

[0212] A percentage value indicating the percentage of users whose spatial speed of the viewing section does not exceed the predetermined value, and / or

[0213] A percentage value indicating the percentage of users for whom the viewing section is in a predetermined region, and / or

[0214] A hint of one or more types of user input for controlling the movement of the viewing section to which the predetermined value is applicable.

[0215] (D5), the apparatus according to (D3) or (D4), the apparatus being configured to perform the setting such that

[0216] The higher the predetermined value d of the measure of the spatial speed for the viewing section, the larger the size.

[0217] (D6), the apparatus according to any one of (D1) or (D5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement for the viewing section.

[0218] (D7), the apparatus according to (D6), the apparatus being configured to perform the setting such that

[0219] The higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.

[0220] (D8), the apparatus according to any one of (D1) or (D7), wherein the information indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.

[0221] (D9), the apparatus according to any one of (D1) or (D8), the apparatus being configured such that the first part coincides with the spatial section.

[0222] (D10), the apparatus according to any one of (D1) or (D9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.

[0223] (D11), the apparatus according to (D10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level within a set of quality levels.

[0224] (D12), The device according to (D1), wherein the device is configured to set the size in a manner independent of the size of the observation section that varies with time depending on the information.

[0225] (D13), The device according to (D1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.

[0226] (D14), The device according to (D12) or (D13), wherein the information indicates the size in terms of the number of tiles.

[0227] This application also provides a streaming media server for media content of a spatial scene that varies over time.

[0228] (E1), The streaming media server is configured to:

[0229] Make a plurality of media segments available for extraction by a device, so that the device can select media segments for extraction, and the media segments at least have a spatial section of the spatial scene that varies over time encoded therein, and the encoding is such that: a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatial scene that varies over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment with another quality reduced relative to the predetermined quality, and the first part follows the time-varying observation section of the spatial scene that varies over time, and

[0230] Signal information in the media segment and / or by signaling to the device regarding how to set the size and / or position of the first part.

[0231] This application also provides a signal defining a media presentation description.

[0232] (F1) The signal includes:

[0233] Information regarding calculating addresses of multiple media segments such that a device can use the information to select and extract a media segment from the multiple media segments, the media segment having at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment with another quality that is reduced relative to the predetermined quality, and such that the first part follows an observation section of the time variation of the spatially varying scene, and

[0234] Information in the media segment and / or via signaling to the device regarding how to set the size and / or position of the first part.

[0235] (F2), according to the signal of (F1), wherein the information indicates the size in the form of an increment relative to the size of the observation section of the time variation or a scaling of the size of the observation section of the time variation.

[0236] (F3), according to the signal of (F1) or (F2), wherein the information indicates a predetermined value of a measure of the spatial speed for the observation section.

[0237] (F4), according to the signal of (F3), wherein the information indicates the predetermined value of the measure of the spatial speed for the observation section for:

[0238] A default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or

[0239] A percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or

[0240] A percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or

[0241] A hint of one or more types of user input for controlling the movement of the observation section for which the predetermined value is applicable.

[0242] (F5), according to the signal of (F3) or (F4), the signal being configured to perform the setting such that

[0243] The higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.

[0244] (F6), according to the signal of any one of (F1)-(F5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement for the observation section.

[0245] (F7), according to the signal of (F6), the signal is configured to perform the setting such that

[0246] the higher the probability of the corresponding moving direction, the more the first part extends into the corresponding direction.

[0247] (F8), according to the signals of (F1)-(F7), wherein the information indicates the predetermined spatial velocity in a time-varying and / or spatially varying and / or directionally varying manner.

[0248] (F9), according to the signal of any one of (F1) or (F8), is configured such that the first part coincides with the spatial section.

[0249] (F10), according to the signal of any one of (F1) or (F9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.

[0250] (F11), according to the signal of (F10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level among a set of quality levels.

[0251] (F12), according to the signal of (F11), wherein the device is configured to set the size in a manner independent of the size of the time-varying observation section depending on the information.

[0252] (F13), according to the signal of (F1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.

[0253] (F14), according to the signal of (F11) or (F12), wherein the information indicates the size in terms of the number of tiles.

[0254] This application also provides a video bitstream.

[0255] (G1) having video encoded therein, the video bitstream includes a signal function for one or more of the size of the focused area into which the decoding ability for decoding the video should be concentrated within the video and the recommended preferred observation section area of the video.

[0256] The present application also provides a decoder for decoding video from a video bitstream.

[0257] (H1) The decoder is configured to:

[0258] derive a signal effect of the size of a focused area within the video from the video bitstream, and

[0259] concentrate the decoding capability for decoding the video to the focused area.

[0260] (H2), the decoder according to (H1), the decoder is configured to specifically decode the focused area.

[0261] (H3), the decoder according to (H1), the decoder is configured to start decoding each image of the video at the focused area.

[0262] (H4), the decoder according to (H1), the decoder is configured to stop decoding each image of the video after decoding the focused area.

[0263] (H5), the decoder according to any one of (H1)-(H4), wherein the signal effect absolutely indicates the size, or the decoder is configured to scale the size of the focused area with a parameter included in the signal effect.

[0264] The present application also provides a device for streaming media content of a spatial scene changing over time.

[0265] (I1) The device is configured to:

[0266] derive from a media presentation description:

[0267] at least one version, for tile-based streaming, the spatial scene changing over time is provided in the at least one version,

[0268] for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming of the corresponding version of the spatial scene changing over time,

[0269] match the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming.

[0270] (I2), the device according to (I1), wherein the beneficial requirements and the device capabilities relate to decoding capabilities.

[0271] (I3), the apparatus according to (I1) or (I2), wherein the beneficial requirements and the apparatus capabilities are related to the number of available decoders.

[0272] (I4), the apparatus according to any one of (I1)-(I3), wherein the beneficial requirements and the apparatus capabilities are related to the layer and / or profile descriptor.

[0273] (I5), the apparatus according to any one of (I1)-(I4), wherein the beneficial requirements and the apparatus capabilities are related to the type of input device for moving the viewing section across the spatial scene that varies over time, or to the speed of moving the viewing section across the spatial scene that varies over time using the input device.

[0274] (I6), the apparatus according to any one of (I1)-(I5), wherein the apparatus is configured to:

[0275] select a media segment from a plurality of media segments available on the server, the selection being made by using a calculation rule included in the media presentation description to calculate the address of the selected media segment,

[0276] extract the selected media segment from the server using the calculated address,

[0277] wherein the apparatus is configured to perform the selection such that the selected media segment has at least a spatial section of the spatial scene that varies over time encoded therein, encoded in such a way that:

[0278] a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatial scene that varies over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and

[0279] such that the first part follows the viewing section that varies over time of the spatial scene that varies over time.

[0280] This application also provides a streaming media server for streaming media content regarding a spatial scene that varies over time.

[0281] (J1) The streaming media server is configured to:

[0282] provide a media presentation description from which

[0283] at least one version can be obtained, and for tile-based streaming, the spatial scene that varies over time is provided in the at least one version,

[0284] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene that benefits from the tile-based streaming with respect to the time variation,

[0285] such that the apparatus for streaming the media content from the streaming server is able to

[0286] match the beneficial requirements of the at least one version with the apparatus capabilities of the apparatus or another apparatus that interacts with the apparatus with respect to the tile-based streaming.

[0287] This application also provides a media presentation description.

[0288] (K1) The media presentation description includes:

[0289] information about at least one version, for tile-based streaming, the spatially varying scene with time variation is provided in the at least one version,

[0290] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene that benefits from the tile-based streaming with respect to the time variation.

[0291] (K2), the media presentation description according to (K1), wherein the beneficial requirements and the apparatus capabilities relate to decoding capabilities.

[0292] (K3), the media presentation description according to (K1) or (K2), wherein the beneficial requirements and the apparatus capabilities relate to the number of available decoders.

[0293] (K4), the media presentation description according to any one of (K1)-(K3), wherein the beneficial requirements and the apparatus capabilities relate to layer and / or profile descriptors.

[0294] (K5), the media presentation description according to any one of (K1)-(K4), wherein the beneficial requirements and the apparatus capabilities relate to the type of input device for moving the viewing segment across the spatially varying scene with time variation, or to the speed of moving the viewing segment across the spatially varying scene with time variation using the input device.

[0295] (K6), the media presentation description according to any one of (K1)-(K5), further includes a calculation rule, using which the apparatus is able to

[0296] select a media segment from a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rule.

[0297] The present application also provides an apparatus for streaming media content of a spatial scene varying over time.

[0298] (L1) The apparatus is configured to:

[0299] calculate an address of a media segment depending on a spatial viewport position and at least one parameter, the media segment describing a spatial scene varying in time and in the at least one parameter,

[0300] use the calculated address to extract the media segment.

[0301] (L2) The apparatus according to claim (L1), wherein the at least one parameter includes one or more coordinates of a center of view and / or a viewing depth.

[0302] The present application also provides a media presentation description.

[0303] (M1) The media presentation description includes:

[0304] a calculation rule for calculating an address of a media segment depending on a spatial viewport position and at least one parameter, so as to use the calculated address to extract the media segment, the media segment describing a spatial scene varying in time and in the at least one parameter.

[0305] The present application also provides a streaming server for allowing an apparatus to stream media content of a spatial scene varying over time from a server.

[0306] (N1) The streaming server is configured to provide the media presentation description according to (M1).

[0307] The present application also provides an apparatus for streaming media content of a spatial scene varying over time.

[0308] (O1) The apparatus is configured to:

[0309] select a media segment from a plurality of media segments available on a server,

[0310] wherein the apparatus is configured to:

[0311] encode a first part of the spatial scene varying over time into the selected media segment with increased quality compared to a spatial neighborhood of the first part or in such a way that the spatial neighborhood of the first part is not encoded into the selected media segment,

[0312] emit a log message that records:

[0313] an instantaneous measurement result of measuring a spatial position and / or movement of the first part; and / or

[0314] Measure statistical values of the spatial position and / or movement of the first part, such as time averages; and / or

[0315] Measure instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or

[0316] Indications of the set of buffers of the device involved in buffering the selected media segment, descriptions of the distribution rules applied in distributing the selected media segment to the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or

[0317] Measurement results of the amount of the selected media segment that has not been output from the buffer of the device for undergoing decoding; and / or

[0318] Measure statistical values of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time averages; and / or

[0319] Measure instantaneous measurement results of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or

[0320] Measure statistical values of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time averages; and / or

[0321] The field of view covered by the observation section; and / or

[0322] Measure instantaneous measurement results of the user position or the observation depth relative to the scene center; and / or

[0323] Measure statistical values of the user position or the observation depth relative to the scene center, such as time averages.

[0324] (O2), The device according to (O1), wherein the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.

[0325] (O3), The device according to (O1) or (O2), configured to perform the selection such that the first part of the spatial scene of the time variation pins the observation section.

[0326] (O4) The apparatus according to any one of (O1)-(O3), configured to emit a log message that records the instantaneous measurement of the quality of the spatial scene that measures the time variation up to being encoded into the selected media segment and up to being visible in the observation section as one of the following:

[0327] The measurement result of the average density of the pixels falling within the observation section, where the spatial scene with the time variation is encoded into the selected media segment at the average density.

[0328] (O5) The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.

[0329] (O6) The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.

[0330] (O7) The apparatus according to (O4), configured to make the emitted log message indicate whether the measurement result measures the average density of the pixels by

[0331] averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment, or

[0332] averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.

[0333] (O8) The apparatus according to (O6) or (O7), wherein averaging the pixel density in a spatially non-uniform manner corresponds to

[0334] averaging in a spherically uniform manner, or

[0335] averaging uniformly in the viewport plane space, where the viewport plane is perpendicular to the central viewing direction of the observation section.

[0336] (O9) The apparatus according to any one of (O4)-(O8), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a way that limits the averaging to the central sub-section of the observation section, or applies a higher averaging weight to the central sub-section compared to the edge part of the observation section around the central sub-section.

[0337] (O10) The apparatus according to any one of (O4)-(O9) is configured such that the measurement results measure the average density of the pixels in a manner that is separate along the horizontal observation section axis and the vertical observation section axis, respectively.

[0338] (O11) The apparatus according to any one of (O4)-(O10) is configured to intermittently emit log messages.

[0339] (O12) The apparatus according to any one of (O4)-(O11) is configured to emit log messages at a rate controlled by a manifest file, and the apparatus performs the selection of the media segments for download based on the manifest file.

[0340] (O13) The apparatus according to any one of (O4)-(O12), wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, and the representations differ in one or more of the following:

[0341] The scene segment of the time-varying spatial scene encoded into the media segment,

[0342] The quality used to encode the time-varying spatial scene into the media segment,

[0343] The spatial quality variation used to encode the time-varying spatial scene into the media segment,

[0344] wherein the apparatus is configured to emit log messages, and the log messages record, in the form of an association of each buffer with one or a combination of two or more of the following, the description of the distribution rule applied in distributing the selected media segments to the set of buffers:

[0345] The scene segment,

[0346] The quality,

[0347] The spatial quality distribution,

[0348] Representation.

[0349] (O14) The apparatus according to any one of (O4)-(O13), wherein the representations are grouped into adaptive sets according to one or more of the following:

[0350] The scene segment of the time-varying spatial scene encoded into the media segment,

[0351] The spatial quality variation used to encode the time-varying spatial scene into the media segment,

[0352] wherein the device is configured to emit a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, in the form of an association of each buffer with one of the adaptive sets, or in the form of an association of each buffer with one of the representations.

[0353] (O15) The device according to any one of (O4)-(O14), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in the form of a time measurement result.

[0354] (O16) The device according to (O15), wherein the device is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in the form of a time measurement result in a time unit less than the time length of the media segment and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.

[0355] (O17) The device according to any one of (O4)-(O16), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in a format classified by one or more of the following:

[0356] the buffer of the decoder in which the corresponding media segment has been buffered,

[0357] the scene segment encoded into the corresponding media segment,

[0358] the quality used to encode the spatially varying scene over time into the corresponding media segment,

[0359] the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.

[0360] The present application also provides a method for streaming media content regarding a spatially varying scene over time.

[0361] (P1) The method includes:

[0362] selecting a media segment from a plurality of media segments available on a server,

[0363] Extract the selected media segment from the server, perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and

[0364] The method further includes deriving the predetermined relationship from information included in the selected media segment and / or from a signal obtained from the server.

[0365] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.

[0366] (Q1) The method includes:

[0367] Make a plurality of media segments available for extraction by a device, so that the device can select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and

[0368] Signal information regarding the predetermined relationship in the media segment and / or by signal action to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.

[0369] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.

[0370] (R1) The method includes:

[0371] Select a media segment from a plurality of media segments available on the server,

[0372] Extract the selected media segment from the server, wherein the selection is performed such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that

[0373] A first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced with respect to the predetermined quality, and

[0374] An observation section that causes the first part to follow the temporal change of the spatial scene that changes over time, and

[0375] The method further includes setting the size of the first part depending on information included in the selected media segment and / or a signal action obtained from the server.

[0376] This application also provides a method for streaming media content regarding a spatial scene that changes over time.

[0377] (S1) The method includes:

[0378] Making a plurality of media segments available for extraction by a device, so that the device can select a media segment for extraction, where the selected media segment has at least a spatial section of the spatial scene that changes over time encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatial scene that changes over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment with another quality that is reduced relative to the predetermined quality, and an observation section that causes the first part to follow the temporal change of the spatial scene that changes over time, and

[0379] Signaling information in the media segment and / or via a signal action to the device regarding how to set the size of the first part.

[0380] This application also provides a method for decoding video from a video bitstream.

[0381] (T1) The method includes:

[0382] Deriving a signal action of the size of a focus area within the video from the video bitstream, and

[0383] Concentrating the decoding ability for decoding the video on the focus area.

[0384] This application also provides a method for streaming media content regarding a spatial scene that changes over time, performed by a device.

[0385] (U1) The method includes:

[0386] Deriving from a media presentation description:

[0387] At least one version, for tile - based streaming, the spatial scene that changes over time is provided with the at least one version,

[0388] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatial scene that benefits from the time variation for tile-based streaming

[0389] Match the beneficial requirements of the at least one version with the device capabilities of the device or another device that interacts with the device regarding tile-based streaming.

[0390] This application also provides a method for streaming media content of a spatial scene that varies over time.

[0391] (V1) The method includes:

[0392] Select media segments from a plurality of media segments available on a server

[0393] Extract the selected media segments from the server, where the selection is performed such that the selected media segments have an increased quality compared to the spatial neighborhood of the first part or are encoded therein in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segments, for the first part of the spatial scene that varies over time

[0394] The method further includes emitting a log message that records:

[0395] Instantaneous measurement results of the spatial position and / or movement of the first part; and / or

[0396] Statistical values of the spatial position and / or movement of the first part, such as time average; and / or

[0397] Instantaneous measurement results of the quality of the spatial scene that varies over time until it is encoded into the selected media segments and until it is visible in the observation section; and / or

[0398] Indications of the set of buffers of the device involved in buffering the selected media segments, a description of the distribution rules applied in distributing the selected media segments to the set of buffers, and the instantaneous buffer fullness of each of the set of buffers; and / or

[0399] Measurement results of the amount of the selected media segments yet to be output from the device's buffer for undergoing decoding; and / or

[0400] Statistical values of the quality of the spatial scene that varies over time until it is encoded into the selected media segments and until it is visible in the observation section, such as time average; and / or

[0401] An instantaneous measurement of the quality of the first part or of the spatial scene up to the time variation encoded into the selected media segment and up to the quality visible in the viewing section; and / or

[0402] A statistical value of the quality of the first part or of the spatial scene up to the time variation encoded into the selected media segment and up to the quality visible in the viewing section, such as a time average; and / or

[0403] The field of view covered by the viewing section; and / or

[0404] An instantaneous measurement of the user position or the viewing depth relative to the scene center; and / or

[0405] A statistical value of the user position or the viewing depth relative to the scene center, such as a time average.

[0406] The present application also provides a method for streaming media content of a spatial scene with time variation.

[0407] (W1) The method includes:

[0408] Providing a media presentation description from which can be derived:

[0409] At least one version, for tile-based streaming, the spatial scene with time variation is provided in the at least one version,

[0410] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatial scene with time variation to benefit from the tile-based streaming,

[0411] Thereby enabling a device for streaming the media content from a streaming server to

[0412] Match the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming.

[0413] The present application also provides a method for allowing a device to stream media content of a spatial scene with time variation from a server.

[0414] (X1) The method includes providing the media presentation description as described in (M1).

[0415] The present application also provides a computer program having program code for performing the method according to any one of claims (P1)-(X1) when the program is executed on a computer.

[0416] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or objects or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0417] Signals as mentioned above (such as, a streamed signal, an MPD, or any other of the mentioned signals) may be stored on a digital storage medium or may be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet.

[0418] Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or in software. The implementation may be carried out using a digital storage medium storing an electronically readable control signal, such as a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, the electronically readable control signal cooperating (or being capable of cooperating) with a programmable computer system to cause the execution of a corresponding method. Thus, the digital storage medium may be computer-readable.

[0419] Some embodiments according to the invention include a data carrier having an electronically readable control signal, the electronically readable control signal being capable of cooperating with a programmable computer system to cause the execution of one of the methods described herein.

[0420] In general, embodiments of the invention may be implemented as a computer program product having program code, which is operable to execute one of the methods when the computer program product is executed on a computer. The program code may (for example) be stored on a machine-readable carrier.

[0421] Other embodiments include a computer program stored on a machine-readable carrier for executing one of the methods described herein.

[0422] In other words, thus, embodiments of the inventive method are computer programs having program code for executing one of the methods described herein when the computer program is executed on a computer.

[0423] Thus, another embodiment of the inventive method is a data carrier (or a digital storage medium, or a computer-readable medium) that includes a computer program recorded thereon for executing one of the methods described herein. The data carrier, the digital storage medium, or the recorded medium is generally tangible and / or non-transitory.

[0424] Accordingly, another embodiment of the method of the present invention is a data stream or signal sequence that represents a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured (e.g.) to be transmitted via a data communication connection (e.g., via the Internet).

[0425] Another embodiment includes a processing component, such as a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0426] Another embodiment includes a computer on which is installed a computer program for performing one of the methods described herein.

[0427] Another embodiment according to the present invention includes an apparatus or system configured to (e.g., electrically or optically) transmit a computer program for performing one of the methods described herein to a receiver. The receiver can be (e.g.) a computer, a mobile device, a storage device, etc. The apparatus or system can (e.g.) include a file server for transmitting the computer program to the receiver.

[0428] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.

[0429] The apparatuses described herein can be implemented using hardware devices or using a computer or using a combination of hardware devices and a computer.

[0430] The apparatuses described herein or any components of the apparatuses described herein can be implemented at least in part in hardware and / or in software.

[0431] The methods described herein can be performed using hardware devices or using a computer or using a combination of hardware devices and a computer.

[0432] The methods described herein or any components of the apparatuses described herein can be performed at least in part by hardware and / or by software.

[0433] The above embodiments merely illustrate the principles of the present invention. It should be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is only intended to be limited by the scope of the appended patent claims and not by the specific details presented by the description and explanation of the embodiments herein.

Claims

1. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select (56) media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the selected media segments are from the server (20), and wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) varying over time encoded therein, encoded such that a first part (64) of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part (66) of the spatially varying scene varying over time that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and derive the predetermined relationship from information (68) included in the selected media segment and / or from a signal obtained from the server (20).

2. The apparatus according to claim 1, wherein each of the plurality (46) of media segments has an associated spatio-temporal part of the spatially varying scene (30) varying over time encoded therein at an associated quality level within a set of quality levels.

3. The apparatus according to claim 2, wherein each spatio-temporal part of the spatially varying scene (30) encoded into the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of tiles (50) into which the spatially varying scene (30) is spatially subdivided.

4. The apparatus according to any one of claims 1 to 3, wherein the information (68) indicates a tolerable value for a measure of a difference between the another quality and the predetermined quality.

5. The apparatus according to claim 4, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on a distance from an observation section (28).

6. The apparatus according to claim 4 or 5, wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on a distance from the observation section (28) by a list of pairs of respective distances from the observation section (28) and corresponding tolerable values for the measure of the difference beyond the respective distances.

7. The apparatus according to any one of claims 4 to 6, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.

8. The apparatus according to any one of claims 4 to 7, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality, and an indication of a maximum allowable time interval during which the second part can be together with the first part within the observation section.

9. The apparatus according to claim 8, wherein the information (68) indicates another tolerable value for a measure of the difference between the another quality and the predetermined quality, and the second portion is capable of indicating another maximum allowable time interval within the observation section together with the first portion.

10. The apparatus according to any one of claims 4 to 9, wherein the information (68) is time-varying and / or spatially varying.

11. The apparatus according to any one of claims 1 to 10, wherein the information (68) indicates an allowable concurrent setting pair for the another quality and the predetermined quality.

12. The apparatus according to any one of claims 1 to 11, wherein the apparatus is configured to perform the selection such that the first portion (64) follows a time-varying observation section (28) of the time-varying spatial scene (30).

13. The apparatus according to claim 12, wherein the apparatus is configured such that a spatial position of the time-varying observation section (28) is changed according to a user input.

14. The apparatus according to any one of claims 1 to 13, wherein the apparatus is configured to determine the first portion (64) to correspond to a region of interest.

15. The apparatus according to claim 14, wherein the apparatus is configured to extract information about the region of interest from the server.

16. A streaming media server for media content of a time-varying spatial scene, configured to: make a plurality of media segments available for extraction by an apparatus, so that the apparatus can select media segments for extraction, the media segments having at least a spatial section of the time-varying spatial scene encoded therein, encoded in such a way that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the time-varying spatial scene spatially adjacent to the first portion is encoded into the selected media segment at another quality, and signal information about a predetermined relationship in the media segment and / or by a signal to the apparatus, the another quality satisfying the predetermined relationship with respect to the predetermined quality.

17. The streaming media server according to claim 16, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the time-varying spatial scene (30) encoded therein at an associated quality level within a set of quality levels.

18. The streaming media server according to claim 17, wherein each of the spatio-temporal portions of the time-varying spatial scene (30) encoded into the plurality (46) of media segments is a time segment (54) of the time-varying spatial scene (30) at a corresponding one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

19. The streaming media server according to any one of claims 16 to 18, wherein the information (68) indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.

20. The streaming media server according to claim 19, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).

21. The streaming media server according to claim 19 or 20, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by means of a list of pairs of corresponding distances from the viewing section (28) and corresponding tolerable values for the measure of the difference beyond the corresponding distances.

22. The streaming media server according to any one of claims 19 to 21, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.

23. The streaming media server according to any one of claims 19 to 22, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.

24. The streaming media server according to claim 23, wherein the information (68) indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.

25. The streaming media server according to any one of claims 19 to 24, wherein the information (68) is time - varying and / or spatially varying.

26. The streaming media server according to any one of claims 19 to 25, wherein the information (68) indicates allowed concurrent setting pairs for the other quality and the predetermined quality.

27. The streaming media server according to any one of claims 19 to 26, wherein the server is configured to send information about the region of interest to the device.

28. A media presentation description, comprising: information about calculating addresses of a plurality of media segments such that a device can use the information to select and extract media segments from the plurality of media segments, wherein at least a spatial section of the spatially varying scene over time is encoded into the media segments in such a way that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatially varying scene over time that is spatially adjacent to the first part is encoded into the selected media segment with another quality, and information (68) about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.

29. The media presentation description according to claim 28, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the spatially varying scene (30) with the time variation encoded therein at an associated quality level within a set of quality levels.

30. The media presentation description according to claim 29, wherein each of the spatio-temporal portions of the spatially varying scene (30) encoded in the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.

31. The media presentation description according to any one of claims 28 to 30, wherein the information (68) indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.

32. The media presentation description according to claim 31, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).

33. The media presentation description according to claim 31 or 32, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by a list of pairs of respective distances from the viewing section (28) and corresponding tolerable values for measures of differences beyond the respective distances.

34. The media presentation description according to any one of claims 31 to 33, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.

35. The media presentation description according to any one of claims 31 to 34, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.

36. The media presentation description according to claim 35, wherein the information (68) indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and another indication of a maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.

37. The media presentation description according to any one of claims 31 to 36, wherein the information (68) is time-varying and / or spatially varying.

38. The media presentation description according to any one of claims 31 to 37, wherein the information (68) indicates an allowed concurrent setting pair for the other quality and the predetermined quality.

39. The media presentation description according to any one of claims 31 to 38, wherein the media presentation description includes information about a region of interest.

40. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) varying over time encoded therein, encoded in such a way that: a first part (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second part (66; 72) of the spatially varying scene varying over time that is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced with respect to the predetermined quality, and such that the first part (64) follows an observation section (28) of the spatially varying scene (30) varying over time, and set the size and / or position of the first part (64) depending on information (74) included in the selected media segment and / or a signal received from the server.

41. The apparatus according to claim 40, wherein the information (74) indicates the size in the form of an increment with respect to the size of the observation section varying over time or a scaling of the size of the observation section varying over time.

42. The apparatus according to claim 40 or 41, wherein the information (74) indicates a predetermined value of a measure of the spatial speed for the observation section.

43. The apparatus according to claim 42, wherein the information (74) indicates the predetermined value of the measure of the spatial speed for the observation section for: a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or a hint indicating one or more types of user input for controlling the movement of the observation section to which the predetermined value is applicable.

44. The apparatus according to claim 42 or 43, the apparatus being configured to perform the setting such that the higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.

45. The apparatus according to any one of claims 40 or 44, wherein the information (74) indicates a predetermined value of a measure of the probability of the direction of movement of the observation section.

46. The apparatus according to claim 45, the apparatus being configured to perform the setting such that the higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.

47. The apparatus according to any one of claims 40 or 46, wherein the information (74) indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.

48. The apparatus according to any one of claims 40 or 47, configured such that the first portion (64) coincides with the spatial section (62).

49. The apparatus according to any one of claims 40 or 48, wherein each of the spatio-temporal portions of the spatially varying scene (30) encoding the temporal variation into the plurality of media segments is a temporal segment (54) of the spatially varying scene (30) at a respective one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.

50. The apparatus according to claim 49, wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying scene encoding the temporal variation therein at an associated quality level within a set of quality levels.

51. The apparatus according to claim 40, wherein the apparatus is configured to set the size in a manner independent of the size of the temporally varying observation section depending on the information (74).

52. The apparatus according to claim 40, wherein the information includes different values for the size for different size options of the time-varying observation section, and the apparatus uses the values included by the information for a size option suitable for the actual size of the time-varying observation section.

53. The apparatus according to claim 51 or 52, wherein the information indicates the size in terms of the number of tiles.

54. A streaming media server for media content of a spatially varying scene over time, configured to: make a plurality of media segments available for extraction by an apparatus, such that the apparatus can select media segments for extraction, the media segments having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced with respect to the predetermined quality, and such that the first portion follows a temporally varying observation section of the spatially varying scene over time, and signal in the media segments and / or by signaling to the apparatus information on how to set the size and / or position of the first portion.

55. A signal defining a media presentation description, comprising: Information regarding calculating addresses of multiple media segments such that a device, using the information, can select and extract a media segment from the multiple media segments, the media segment having at least a spatial section of the spatially varying scene varying in time encoded therein in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene varying in time and spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first part follows an observation section of the time variation of the spatially varying scene, and Information in the media segment and / or signaled to the device regarding how to set the size and / or position of the first part.

56. The signal according to claim 55, wherein the information (74) indicates the size in the form of an increment relative to the size of the observation section of the time variation or a scaling of the size of the observation section of the time variation.

57. The signal according to claim 55 or 56, wherein the information (74) indicates a predetermined value of a measure of the spatial speed for the observation section.

58. The signal according to claim 57, wherein the information (74) indicates the predetermined value of the measure of the spatial speed for the observation section for: A default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or A percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or A percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or A hint of one or more types of user input controlling the movement of the observation section for which the predetermined value is applicable.

59. The signal according to claim 57 or 58, the signal being configured to perform the setting such that The higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.

60. The signal according to any one of claims 55 or 59, wherein the information (74) indicates a predetermined value of a measure of the probability of the direction of movement of the observation section.

61. The signal according to claim 60, the signal being configured to perform the setting such that The higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.

62. The signal according to any one of claims 55 or 61, wherein the information (74) indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.

63. The signal according to any one of claims 55 or 62, configured such that the first part (64) coincides with the spatial section (62).

64. The signal according to any one of claims 55 or 63, wherein each of the spatio-temporal parts of the spatially varying spatio-temporal scene (30) encoded into the plurality of media segments is a temporal segment (54) of the spatially varying spatio-temporal scene (30) at a corresponding one of the tiles (50) into which the spatially varying spatio-temporal scene (30) is spatially subdivided.

65. The signal according to claim 64, wherein each of the plurality of media segments has an associated spatio-temporal part of the spatially varying spatio-temporal scene encoded therein at an associated quality level within a set of quality levels.

66. The signal according to claim 55, wherein the device is configured to set the size in a manner independent of the size of the temporally varying observation section depending on the information (74).

67. The signal according to claim 55, wherein the information includes different values for the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.

68. The signal according to claim 65 or 66, wherein the information indicates the size in terms of the number of tiles.

69. A video bitstream having video encoded therein, the video bitstream including a signal effect for one or more of the size of a focus area within the video to which the decoding ability for decoding the video should be concentrated and a recommended preferred observation section area of the video.

70. A decoder for decoding video from a video bitstream, configured to: derive a signal effect (74) of the size of a focus area within the video from the video bitstream, and concentrate the decoding ability for decoding the video to the focus area (116).

71. The decoder according to claim 70, the decoder being configured to specifically decode the focus area.

72. The decoder according to claim 70, the decoder being configured to decode each image of the video starting from where the focus area is decoded.

73. The decoder according to claim 70, the decoder being configured to stop decoding each image of the video after decoding the focus area.

74. The decoder according to any one of claims 70 to 73, wherein the signal effect absolutely indicates the size, or the decoder is configured to scale the size of the focus area with a parameter included in the signal effect.

75. A device for streaming media content regarding a spatially varying spatio-temporal scene (30), configured to: derive (90) from a media presentation description: at least one version, for tile-based streaming, the spatially varying spatio-temporal scene being provided in the at least one version, for each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatially varying spatio-temporal scene for benefiting from the tile-based streaming, Match the benefit requirements of the at least one version to the device capabilities of the device or another device interacting with the device regarding the tile-based streaming media (92).

76. The device according to claim 75, wherein the benefit requirements and the device capabilities relate to decoding capabilities.

77. The device according to claim 75 or 76, wherein the benefit requirements and the device capabilities relate to the number of available decoders.

78. The device according to any one of claims 75 to 77, wherein the benefit requirements and the device capabilities relate to layer and / or profile descriptors.

79. The device according to any one of claims 75 to 78, wherein the benefit requirements and the device capabilities relate to the type of input device for moving an observation section across the spatial scene (30) that changes over time, or to the speed of moving the observation section across the spatial scene (30) that changes over time using the input device.

80. The device according to any one of claims 75 to 79, wherein the device is configured to: Select media segments from a plurality (46) of media segments (58) available on the server (20), the selection calculating the addresses of the selected media segments by using calculation rules included in the media presentation description, Extract the selected media segments from the server (20) using the calculated addresses, wherein the device is configured to perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) that changes over time encoded therein, encoded in such a way that: A first part (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second part (66; 72) of the spatially varying scene that changes over time and is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first part (64) follows the observation section (28) of the spatially varying scene (30) that changes over time.

81. A streaming media server for streaming media content regarding a spatially varying scene (30) that changes over time, configured to: Provide (90) a media presentation description from which can be derived At least one version, for tile-based streaming, the spatially varying scene that changes over time is provided in the at least one version, For each of the at least one version, an indication of the benefit requirements for the corresponding version of the spatially varying scene that benefits from the tile-based streaming, Thereby enabling a device streaming the media content from the streaming media server to Match the benefit requirements of the at least one version to the device capabilities of the device or another device interacting with the device regarding the tile-based streaming (92).

82. A media presentation description, comprising: Information about at least one version, for tile - based streaming, a spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene over time to benefit from the tile - based streaming.

83. The media presentation description according to claim 82, wherein the beneficial requirements and the device capabilities relate to decoding capabilities.

84. The media presentation description according to claim 82 or 83, wherein the beneficial requirements and the device capabilities relate to the number of available decoders.

85. The media presentation description according to any one of claims 82 to 84, wherein the beneficial requirements and the device capabilities relate to layer and / or profile descriptors.

86. The media presentation description according to any one of claims 82 to 85, wherein the beneficial requirements and the device capabilities relate to the type of input device for moving an observation section across the spatially varying scene (30) over time, or to the speed of using the input device to move the observation section across the spatially varying scene (30) over time.

87. The media presentation description according to any one of claims 82 to 86, further comprising a calculation rule, using which the device can select a media segment from a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rule.

88. A device for streaming media content regarding a spatially varying scene (30) over time, configured to: calculate the address of a media segment depending on a spatial viewport position and at least one parameter, the media segment describing the spatially varying scene (30) varying in time and in the at least one parameter, use the calculated address to extract the media segment.

89. The device according to claim 88, wherein the at least one parameter includes one or more coordinates of an observation center and / or an observation depth.

90. A media presentation description, comprising: a calculation rule for calculating the address of a media segment depending on a spatial viewport position and at least one parameter, in order to use the calculated address to extract the media segment, the media segment describing the spatially varying scene (30) varying in time and in the at least one parameter.

91. A streaming server for allowing a device to stream media content regarding a spatially varying scene over time from a server, configured to provide the media presentation description according to claim 90.

92. A device for streaming media content regarding a spatially varying scene over time, configured to: select a media segment from a plurality of media segments available on the server, wherein the device is configured to: encode a first part of the spatially varying scene over time into the selected media segment with increased quality compared to the spatial neighborhood of the first part or in such a way that the spatial neighborhood of the first part is not encoded into the selected media segment, issue a log message that records: Instantaneous measurement results of the spatial position and / or movement of the first part; and / or Statistical values of the spatial position and / or movement of the first part, such as time average values; and / or Instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or Indications of the set of buffers (300) of the device involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment to the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or Measurement results of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding (42); and / or Statistical values of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or Instantaneous measurement results of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or Statistical values of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or The field of view covered by the observation section; and / or Instantaneous measurement results of the user position or the observation depth relative to the scene center (100); and / or Statistical values of the user position or the observation depth relative to the scene center (100), such as time average values.

93. The device according to claim 92, wherein the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.

94. The device according to claim 92 or 93, configured to perform the selection such that the first part (64) of the spatial scene of the time variation pins the observation section (28).

95. The device according to any one of claims 92 to 94, configured to issue a log message that records the instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section as one of the following: Measurement results of the average density of the pixels falling within the observation section (28), with the spatial scene of the time variation being encoded into the selected media segment at the average density.

96. The device according to claim 95, configured such that the measurement results measure the average density of the pixels by averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.

97. The apparatus according to claim 95, configured such that the measurement result measures the average density of pixels by averaging the pixel density in a spatially non-uniform manner with respect to a pixel grid of an image encoded in the selected media segment.

98. The apparatus according to claim 95, configured such that the emitted log message indicates whether the measurement result measures the average density of pixels by: averaging the pixel density in a spatially uniform manner with respect to a pixel grid of an image encoded in the selected media segment, or averaging the pixel density in a spatially non-uniform manner with respect to a pixel grid of an image encoded in the selected media segment.

99. The apparatus according to claim 97 or 98, wherein averaging the pixel density in a spatially non-uniform manner corresponds to averaging in a spherical uniform manner, or averaging spatially uniformly with respect to a viewport plane (310) that is perpendicular to the central viewing direction (312) of the viewing section (28).

100. The apparatus according to any one of claims 95 to 99, configured such that the measurement result measures the average density of pixels by: averaging the pixel density in a manner that limits the averaging to a central sub-section of the viewing section (28), or applying a higher averaging weight to the central sub-section (202) compared to an edge portion (204) of the viewing section surrounding the central sub-section.

101. The apparatus according to any one of claims 95 to 100, configured such that the measurement result measures the average density of pixels separately along a horizontal viewing section axis (204) and a vertical viewing section axis (206).

102. The apparatus according to any one of claims 95 to 101, configured to intermittently emit log messages.

103. The apparatus according to any one of claims 95 to 102, configured to emit log messages at a rate controlled by a manifest file, and the apparatus performs the selection of the media segment for download based on the manifest file.

104. The apparatus according to any one of claims 95 to 103, wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, and the representations differ in one or more of the following: a scene section (50) of the time-varying spatial scene encoded in the media segment, the quality used to encode the time-varying spatial scene in the media segment, the spatial quality variation used to encode the time-varying spatial scene in the media segment, wherein the apparatus is configured to emit a log message that records, in the form of an association of each buffer with a combination of one or two or more of the following, a description of the distribution rule applied in distributing the selected media segment to the set of buffers: the scene section, the quality, the spatial quality distribution, representation 105. The apparatus according to any one of claims 95 to 104, wherein the representations are grouped into adaptation sets according to one or more of the following: the scene segments of the spatially varying scene over time encoded into the media segment, the spatial quality variations used to encode the spatially varying scene over time into the media segment, wherein the apparatus is configured to emit a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, either in the form of an association of each buffer with one of the adaptation sets or in the form of an association of each buffer with one of the representations.

106. The apparatus according to any one of claims 95 to 105, wherein the apparatus is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42), in the form of a time measurement result.

107. The apparatus according to claim 106, wherein the apparatus is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42) in the form of a time measurement result in time units less than the time length of the media segment and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.

108. The apparatus according to any one of claims 95 to 107, wherein the apparatus is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42) in a format classified by one or more of the following: the buffer of the decoder in which the corresponding media segment has been buffered, the scene segment encoded into the corresponding media segment, the quality used to encode the spatially varying scene over time into the corresponding media segment, the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.

109. A method for streaming media content regarding a spatially varying scene (30) over time, comprising: selecting (56) media segments from a plurality (46) of media segments (58) available on a server (20), extracting (60) the selected media segments from the server (20), performing the selection such that the selected media segments have at least a spatial segment (62) of the spatially varying scene (30) encoded therein, encoded such that a first part (64) of the spatial segment is encoded into the selected media segment at a predetermined quality and a second part (66) of the spatially varying scene that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and The method further includes deriving the predetermined relationship from information (68) included in the selected media segment and / or from a signal received from the server (20).

110. A method for streaming media content of a spatially varying scene over time, comprising: making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and signaling information about a predetermined relationship in the media segment and / or by a signal to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.

111. A method for streaming media content of a spatially varying scene (30) over time, comprising: selecting a media segment from a plurality (46) of media segments (58) available on a server (20), extracting the selected media segment from the server (20), wherein the selection is performed such that the selected media segment has at least a spatial section (62) of the spatially varying scene (30) over time encoded therein, encoded such that a first portion (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality and a second portion (66; 72) of the spatially varying scene over time that is spatially adjacent to the first portion (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and causing the first portion (64) to follow an observation section (28) of the spatially varying scene (30) over time, and the method further includes setting a size of the first portion (64) depending on information (74) included in the selected media segment and / or a signal received from the server.

112. A method for streaming media content of a spatially varying scene over time, comprising: making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and causing the first portion to follow an observation section of the spatially varying scene over time, and Signal information on how to size the first part in the media segment and / or via signaling to the device.

113. A method for decoding video from a video bitstream, comprising: Signaling the size of a focus region within the video derived from the video bitstream, and Focusing decoding capabilities for decoding the video on the focus region.

114. A method for streaming media content regarding a spatially varying scene (30) over time, performed by a device, comprising: Deriving (90) from a media presentation description: At least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene over time to benefit from the tile-based streaming, Matching (92) the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming.

115. A method for streaming media content regarding a spatially varying scene over time, comprising: Selecting a media segment from a plurality of media segments available on a server, Extracting the selected media segment from the server, wherein the selection is performed such that the selected media segment has the first part of the spatially varying scene encoded therein with increased quality compared to the spatial neighborhood of the first part or in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segment, The method further comprises emitting a log message that records: Instantaneous measurements of the spatial position and / or movement of the first part; and / or Statistical values of the spatial position and / or movement of the first part, such as a time average; and / or Instantaneous measurements of the quality of the spatially varying scene up to being encoded into the selected media segment and up to being visible in the viewing segment; and / or An indication of the set of buffers of the device involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment into the set of buffers, and the instantaneous buffer fullness of each of the set of buffers; and / or A measurement of the amount of the selected media segment yet to be output from the buffers of the device for undergoing decoding (42); and / or Statistical values of the quality of the spatially varying scene up to being encoded into the selected media segment and up to being visible in the viewing segment, such as a time average; and / or Instantaneous measurements of the quality of the first part or of the spatially varying scene up to being encoded into the selected media segment and up to being visible in the viewing segment; and / or Statistical values of the quality of the first part or of the spatially varying scene up to being encoded into the selected media segment and up to being visible in the viewing segment, such as a time average; and / or The field of view covered by the viewing segment; and / or Instantaneous measurements of the user position or viewing depth relative to the scene center (100); and / or Statistical values of the user position or viewing depth relative to the scene center (100), such as time averages.

116. A method for streaming media content of a spatio-temporal scene (30) that varies over time, comprising: Providing (90) a media presentation description from which can be derived: At least one version, for tile-based streaming, the spatio-temporal scene that varies over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatio-temporal scene that varies over time for benefiting from the tile-based streaming, Thereby enabling a device for streaming the media content from a streaming server to Match (92) the beneficial requirements of the at least one version with the device capabilities of the device or another device that interacts with the device regarding the tile-based streaming.

117. A method for allowing a device to stream media content of a spatio-temporal scene (30) that varies over time from a server, comprising providing the media presentation description as claimed in claim 90.

118. A computer program having program code for performing the method according to any one of claims 109 to 117 when the program is executed on a computer.

Citation Information

Patent Citations

  • Spatially unequal streaming

    CN115037917A