Spatial inequality streaming

By adopting spatial uneven encoding method in VR streaming, media segments are selected and extracted, and the quality is dynamically adjusted according to user viewport direction and other parameters, the problem of user visible quality reduction and calculation complexity caused by spatial unevenness in the prior art is solved, and efficient bandwidth utilization and user experience improvement are achieved.

CN115037917BActive Publication Date: 2025-05-13FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210671217.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-08
Filing Date
2017-10-11
Publication Date
2025-05-13
Estimated Expiration
2037-10-11

AI Technical Summary

Technical Problem

In the virtual reality (VR) streaming, the prior art is difficult to effectively solve the problems of reduced user visible quality caused by spatial inequality and increased computing complexity and bandwidth requirements of streaming extraction sites.

Method used

By selecting and extracting media segments, and using information contained in the media segments and signals obtained from the server, the spatial uneven method is used to encode the spatial scene of time-changing, thereby realizing the streaming of videos. The specific method includes setting different qualities of encoding in a media segment with a predetermined relationship, calculating the address of the media segment according to the viewport direction and other parameters, and extracting the media segment from the server.

Benefits of technology

Improve the user's visible quality under relatively low bandwidth consumption and reduce the computational complexity of streaming receiving sites, avoiding the decline in user experience caused by quality degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115037917B_ABST
    Figure CN115037917B_ABST
Patent Text Reader

Abstract

Various concepts for media content streaming are described. Some concepts allow spatial scene content to be streamed in a spatially unequal manner so that the visible quality to the user is increased, or the processing complexity or required bandwidth at the streaming extraction site is reduced. Other concepts allow spatial scene content to be streamed in a manner that increases applicability to other application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application whose applicant is Fraunhofer Gesellschaft, whose application date is October 11, 2017, whose application number is 201780076761.7, and whose invention name is “Spatially Unequal Streaming”. Technical Field

[0002] This application is about spatially uneven streaming, such as occurs in virtual reality (VR) streaming. Background Art

[0003] VR streaming typically involves the transmission of very high-resolution video. The resolution capability of the human fovea is about 60 pixels per degree. If the transmission of a complete sphere of 360°×180° is considered, the transmission can end up sending a resolution of about 22k×11k pixels. Since sending this high resolution will generate extremely high bandwidth requirements, another solution is to send only the viewport displayed at the head-mounted display (HMD), which has a field of view of 90°×90°, resulting in a video of about 6k×6k pixels. The compromise between sending the full video at the highest resolution and sending only the viewport is to send the viewport at a high resolution and send some adjacent data (or the rest of the sphere video) at a lower resolution or lower quality.

[0004] In a DASH scenario, omnidirectional video (also called spherical video) can be provided in a manner as described previously where mixed resolution or mixed quality video is controlled by the DASH client. The DASH client only needs to know information describing how the content is provided.

[0005] One example could be to provide different representations with different projections, with asymmetric properties, such as different quality and distortion for different parts of the video. Each representation would correspond to a given viewport and would have that viewport encoded at a higher quality / resolution than the rest of the content. Knowing the directional information (the direction of the viewport where the content has been encoded at a higher quality / resolution), the DASH client can dynamically select one or the other representation to match the viewing direction of the user at any time.

[0006] A more flexible option for a DASH client to select this asymmetric characteristic for omni-directional video may be when the video is split into several spatial regions, each of which may be available at a different resolution or quality. One option may be to split the video into rectangular regions (also called tiles) based on a grid, but other options are foreseeable. In this case, the DASH client will need some signaling about the different qualities at which different regions are provided, and the DASH client may download different regions at different qualities so that the quality of the viewport shown to the user is better than other content that is not shown.

[0007] In either of the previous cases, when user interaction occurs and the viewport has changed, it takes some time for the DASH client to react to the user movement and download the content in a way that matches the new viewport. During the time between the user movement and the DASH client adapting its request to match the new viewport, the user will see some areas of high quality and low quality simultaneously in the viewport. While it is acceptable that the quality / resolution difference is content-dependent, the quality seen by the user is reduced in any case.

[0008] Therefore, the concept of having a relieved or more efficient presentation or even increased visible quality for a user regarding a portion of a spatial scene content streamed by adaptive streaming may be advantageous. Summary of the invention

[0009] Therefore, an object of the present invention is to provide a concept for streaming spatial scene content in a spatially unequal manner so that the visible quality to the user is increased, or the processing complexity or required bandwidth at the streaming extraction site is reduced, or to provide a concept for streaming spatial scene content in a manner that increases applicability to other application scenarios.

[0010] This object is achieved by the subject matter of the independent claims of the application.

[0011] A first aspect of the present application is based on the discovery that streaming media content (such as video) about a time-varying spatial scene in a spatially unequal manner can be improved in terms of visible quality and / or computational complexity at a streaming reception site at comparable bandwidth consumption if the selected and extracted media segments and / or the signaling obtained from the server provide the extraction device with hints about the predetermined relationship to which the qualities used to encode different parts of the time-varying spatial scene into the selected and extracted media segments comply. Otherwise, the extraction device may not know in advance how the juxtaposition of parts encoded at different qualities into the selected and extracted media segments will negatively affect the overall visible quality experienced by the user. Information contained in the media segments and / or the signaling obtained from the server (such as in a manifest file (media presentation description) or additional streaming-related control messages (such as SAND messages) from the server to the client) enables the extraction device to make an appropriate selection among the media segments provided at the server. In this way, virtual reality streaming or partial streaming of video content can become more robust against quality degradation that occurs due to insufficient distribution of available bandwidth over this spatial segment of the time-varying spatial scene presented to the user.

[0012] Another aspect of the invention is based on the finding that streaming media content (such as video) about a temporally varying spatial scene in a spatially unequal manner (such as using a first quality in a first portion and a lower second quality in a second portion or leaving the second portion unstreamed) can be improved in visible quality and / or become less complex with respect to bandwidth consumption and / or computational complexity at the streaming extraction side by determining the size and / or position of the first portion depending on information contained in the media fragment and / or a signaling effect obtained from a server. For example, for tile-based streaming, it is envisaged that the temporally varying spatial scene may be provided at the server in a tile-based manner, i.e. the media fragments may represent spectro-temporal portions of the temporally varying spatial scene, each of which may be a temporal segment of the spatial scene within a corresponding tile of the distribution of tiles into which the spatial scene is subdivided. In this case, the decision is made by the extraction device (client) as to how to distribute the available bandwidth and / or computational power in the spatial scene (i.e. at tile granularity). The extraction device may perform the selection of the media segments so that the first part of the spatial scene (which respectively follows the observation segment that tracks the temporal changes of the spatial scene) is encoded into the selected and extracted media segments with a predetermined quality, which may be (for example) the highest quality that can be implemented under the current bandwidth and / or computing power conditions. For example, the second part of the spatial scene that is spatially adjacent may not be encoded into the selected and extracted media segments, or may be encoded into the media segments with another quality that is reduced relative to the predetermined quality. In this case, it is computationally complex or even infeasible to calculate the number / count of neighboring tiles, the aggregation of which completely covers the temporal changing observation segment, regardless of the orientation of the observation segment. Depending on the projection selected in order to map the spatial scene onto the individual tiles, the angular scene coverage of each tile may vary in this scene, and the fact that the individual tiles may overlap each other even makes it more difficult to calculate the count of neighboring tiles that are sufficient to cover the observation segment in terms of space (regardless of the orientation of the observation segment). Thus, in this case, the aforementioned information may indicate the size of the first part as a count N of tiles or as the number of tiles, respectively. By this measure, the device will be able to track the time-varying observation segment by selecting those media segments having a co-located aggregation of N tiles encoded with a predetermined quality. The fact that the aggregation of these N tiles adequately covers the observation segment may be ensured by the information indicating N. Another example may be information contained in the media segment and / or a signaling obtained from a server, which indicates the size of the first part relative to the size of the observation segment itself. For example, this information may to some extent set a "safety zone" or pre-fetch zone around the actual observation segment in order to take into account the movement of the time-varying observation segment. The greater the speed at which the time-varying observation segment moves across the spatial scene, the larger the safety zone should be.Thus, the aforementioned information may indicate the size of the first portion in a manner relative to the size of the time-varying observation segment, such as in an incremental or scaled manner. An extraction device setting the size of the first portion according to this information will be able to avoid quality degradation, which may otherwise occur due to unextracted or low-quality parts of the spatial scene being visible in the observation segment. It is irrelevant here whether this scene is provided in a tile-based manner or in some other manner.

[0013] In connection with the just mentioned aspect of the application, a video bitstream encoded with a video can be decoded with increased quality, provided that the video bitstream is provided with a signaling of the size of a focus region within said video, to which the decoding capabilities for decoding the video should be focused. By this measure, a decoder decoding the video from a bitstream can focus or even limit its decoding capabilities for decoding the video to a portion having the size of the focus region signaled in said video bitstream, knowing from this that (for example) the portion so decoded is decodable with the available decoding capabilities and spatially covers a desired segment of the video. For example, the size of the focus region so signaled can be chosen to be large enough to cover the size of an observation segment and the movement of this observation segment, thereby taking into account the decoding delay when decoding the video. Or, in other words, the signaling of a recommended preferred observation segment area of ​​the video included in the video bitstream can allow the decoder to process this area in a better way, thereby allowing the decoder to focus its decoding capabilities accordingly. Regardless of whether region-specific decoding capability concentration is performed, the region signaling may be forwarded to the platform, which selects which media segments to download, i.e., where to place and how to size the increased quality portions.

[0014] The first and second aspects of the present application are closely related to the third aspect of the application, according to which the fact that a large number of extraction devices stream media content from a server is used in order to obtain information that can then be used to appropriately set the aforementioned type of information, thereby allowing the size of the first part or the size and / or position to be set, and / or the predetermined relationship between the first quality and the second quality to be appropriately set. Therefore, according to this aspect of the present application, the extraction device (client) issues a log message, which records one of the following: an instantaneous measurement result or statistic measuring the spatial position and / or movement of the first part; an instantaneous measurement result or statistic measuring the quality of the time-varying spatial scene until it is encoded into the selected media segment and until it is visible in the observation segment; and an instantaneous measurement result or statistic measuring the quality of the first part or the quality of the time-varying spatial scene until it is encoded into the selected media segment and until it is visible in the observation segment. The instantaneous measurement results and / or statistics can be provided with time information related to the time when the corresponding instantaneous measurement results or statistics have been obtained. The log message may be sent to a server providing the media segment, or to some other device evaluating incoming log messages so as to update current settings of the aforementioned information for setting the size of the first portion or the size and / or position based on the log message, and / or derive a predetermined relationship based on the log message.

[0015] According to another aspect of the present application, in particular, by providing a media presentation description including: at least one version, for tile-based streaming, the time-varying spatial scene is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements of the corresponding version of the time-varying spatial scene for benefiting from the tile-based streaming. By this measure, the extraction device is able to match the beneficial requirements of the at least one version with the device capabilities of the extraction device itself or another device interacting with the extraction device for tile-based streaming. For example, the benefit requirements may be related to decoding capability requirements. That is, if the decoding capabilities used to decode the streamed / extracted media content will be insufficient to decode all the media segments required to cover the observation segment of the time-varying spatial scene, then attempting to stream and present the media content will waste time, bandwidth and computing power, and therefore, it may be more efficient not to attempt to stream and present the media content in any case. For example, if, for example, the media segments related to a particular tile form a separate media stream (such as a video stream) from the media segments related to another tile, the decoding capability requirement may, for example, indicate the number of decoder instances required for the corresponding version. For example, the decoding capability requirement may also be related to other information, such as the specific portion of the decoder instances necessary to fit a predetermined decoding profile and / or tier, or may indicate a specific minimum capability of a user input device to move the viewport / segment in a fast enough manner for the user to view the scene. Depending on the scene content, a low movement capability may not be enough for the user to view the interesting portion of the scene.

[0016] Another aspect of the invention relates to the extension of streaming media content with respect to spatial scenes that vary in time. In particular, the idea according to this aspect is that the spatial scene may in fact vary not only in time, but also with respect to at least one other parameter that implies variation (e.g., view and position, viewing depth, or some other physical parameter). The extraction device may use adaptive streaming in this context by: calculating the address of a media segment depending on the viewport direction and at least one other parameter, the media segment describing the spatial scene that varies in time and in the at least one parameter; and extracting the media segment from a server using the calculated position.

[0017] The above-outlined aspects of the present application and the advantageous implementations that are the subject of the dependent claims can be combined individually or together. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The following describes a preferred embodiment of the present application with reference to the accompanying drawings, in which:

[0019] Figure 1 shows a schematic diagram illustrating a system of clients and servers for virtual reality applications as an example of a situation in which the embodiments described in the following figures may be advantageously used;

[0020] Figure 2 A block diagram showing a client device and a schematic illustration of a media segment selection procedure for describing a possible mode of operation of the client device according to an embodiment of the present application, wherein a server 10 provides the device with information about acceptable or tolerable quality variations within media content presented to a user;

[0021] Figure 3 exhibit Figure 2 The modification, the part of quality increase is not concerned with tracking the observation segment of the viewport, but with the region of interest of the media scene content signaled from the server to the client;

[0022] Figure 4 A block diagram showing a client device and a schematic illustration of a media segment selection procedure according to an embodiment, wherein a server provides information on how to set the size or size and / or position of the quality-increased portion, or the size or size and / or position of the actual extracted segment of the media scene;

[0023] Figure 5 exhibit Figure 5 A variation of , in which the information sent by the server directly indicates the size of portion 64, rather than scaling the portion depending on the intended movement of the viewport;

[0024] Figure 6 exhibit Figure 4 a variant according to which the extracted segment has a predetermined quality and its size is determined by information originating from a server;

[0025] Figures 7a to 7c Display description according to Figure 4 and Figure 6 A schematic diagram of the manner in which the size of a portion extracted with a predetermined quality is increased by a corresponding enlargement of the size of the viewport;

[0026] Figure 8a A schematic diagram illustrating an embodiment in which a client device sends log messages to a server or a specific evaluator for evaluating these log messages, for example, for Figures 2 to 7c The type of information discussed derives an appropriate setting;

[0027] Figure 8bSchematic diagram showing a tile-based cubic projection of a 360 scene onto tiles, and examples of how some of the tiles are overlaid with illustrative positions of the viewport. Small circles indicate positions in the viewport that are equiangularly distributed, and shaded tiles are encoded in the downloaded fragment at a higher resolution than non-shaded tiles;

[0028] Figure 8c and Figure 8d A schematic diagram showing a diagram showing how the buffer fullness (vertical axis) of different buffers of a client may develop along the time axis (horizontal), wherein Figure 8c Assume that the buffer will be used to buffer the representation of a particular tile, and Figure 8d Assuming that the buffer is to be used to buffer an omnidirectional representation of a scene having been encoded therein with non-uniform quality (i.e. increasing towards a certain direction specific to the respective buffer);

[0029] Figure 8e and Figure 8f a three-dimensional graph showing different pixel density measurements within the viewport 28, with the difference being uniformity in the sense of a sphere or viewing plane;

[0030] Fig. 9 A block diagram showing a client device and a schematic illustration of a media segment selection process when the device detects information from a server in order to evaluate whether a particular version of tile-based streaming provided by the server is acceptable to the client device;

[0031] Fig.10 A schematic diagram illustrating a plurality of media segments provided by a server according to an embodiment is shown, thereby allowing a media scene to depend not only on time but also on another non-time parameter (i.e., illustratively here on the scene center position);

[0032] Fig.11 Schematic diagrams are shown illustrating a video bitstream that includes information to manipulate or control the size of a focus region within the video encoded into the bitstream, and an example of a video decoder that is able to exploit this information. DETAILED DESCRIPTION

[0033] In order to easily understand the description of various aspects of the present application, Figure 1 Examples of environments in which the subsequently described embodiments of the present application may be applicable and advantageously used are presented. In particular, Figure 1A system consisting of a client 10 and a server 20 interacting via adaptive streaming is shown. For example, dynamic adaptive streaming over HTTP (DASH) may be used for communication 22 between the client 10 and the server 20. However, the embodiments outlined subsequently should not be interpreted as limited to the use of DASH, and likewise, terms such as media presentation description (MPD) should be understood broadly so as to also encompass manifest files whose definitions differ from those in DASH.

[0034] Figure 1 A system configured to implement a virtual reality application is described. That is, the system is configured to present to a user wearing a heads-up display 24 (i.e., via an internal display 26 of the heads-up display 24) an observation segment 28 of a spatial scene 30 that changes over time, the segment 28 corresponding to the orientation of the heads-up display 24 as illustratively measured by an internal orientation sensor 32 (such as an inertial sensor of the heads-up display 24). That is, the segment 28 presented to the user forms a segment of the spatial scene 30 whose spatial position corresponds to the orientation of the heads-up display 24. Figure 1 In the case of time-varying spatial scene 30, the time-varying spatial scene 30 is depicted as an omnidirectional video or a spherical video, but Figure 1 The description and subsequently explained embodiments may also be easily transferred to other examples, such as presenting a segment in a video, wherein the spatial position of the segment 28 is determined by the intersection of a face access or an eye access with a virtual or real projection wall or the like. Furthermore, the sensor 32 and the display 26 may be respectively, for example, comprised by different devices, such as a remote control and a corresponding television, or the sensor and the display may be part of a handheld device, such as a mobile device, such as a tablet or a mobile phone. Finally, it should be noted that some of the embodiments described later may also be applied to scenarios in which the area 28 presented to the user always covers the entire time-varying spatial scene 30, wherein the non-uniformity in presenting the time-varying spatial scene is related to, for example, an unequal distribution of mass in the spatial scene.

[0035] Additional details about the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are described in Figure 1 However, these details should not be considered as limiting the embodiments explained later, but should actually serve as examples of how to implement any one of the embodiments explained later.

[0036] In particular, if Figure 1As shown in FIG. 1 , the server 20 may include a memory 34 and a controller 36, such as a suitably programmed computer, an application specific integrated circuit, etc. The memory 34 has stored thereon media clips representing the time-varying spatial scene 30. Figure 1 The description of FIG. 1 outlines a specific example in more detail. Controller 36 answers a request sent by client 10 by resending the requested media segment to client 10, and the media presentation description may send information about itself to client 10. Details about this are also set out below. Controller 36 may retrieve the requested media segment from memory 34. Other information may also be stored in this memory, such as the media presentation description or portions thereof, which is sent from server 20 to client 10 in other signals.

[0037] like Figure 1 As shown in , the server 20 may optionally further include a stream modifier 38, which modifies the media segments sent from the server 20 to the client 10 in response to a request from the client 10, so as to generate a media data stream at the client 10, which forms a single media stream decodable by an associated decoder, but (for example) the media segments extracted in this way by the client 10 are actually aggregated from several media streams. However, the presence of this stream modifier 38 is optional.

[0038] Figure 1 The client 10 is illustratively depicted as comprising a client device or controller 40 or a plurality of decoders 42 and a re-projector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA or an application specific integrated circuit, etc. The client device 40 undertakes the responsibility of selecting the segments to be extracted from the server 20 from a plurality 46 of media segments provided at the server 20. For this purpose, the client device 40 first extracts a manifest or a media presentation description from the server 20. From the manifest or the media presentation description, the client device 40 obtains a calculation rule for calculating the addresses of the media segments of the plurality 46 of media segments that correspond to specific desired spatial portions of the spatial scene 30. The media segments thus selected are extracted from the server 20 by the client device 40 by sending corresponding requests to the server 20. These requests contain the calculated addresses.

[0039] The media segments thus extracted by the client device 40 are forwarded by the client device 40 to one or more decoders 42 for decoding. Figure 1In the example of , the media segment thus extracted and decoded represents for each temporal unit only a spatial segment 48 of the temporally varying spatial scene 30, but as already indicated above, this may be different depending on other aspects of the scene, for example, whether the observation segment 28 to be presented always covers the entire scene. The reprojector 44 may optionally reproject the observation segment 28 to be displayed to the user and cut it out from the extracted and decoded scene content of the selected, extracted and decoded media segment. For this purpose, as Figure 1 , client device 40 may, for example, continuously track the spatial position of observation segment 28 and update the spatial position in response to user orientation data from sensor 32, and notify reprojector 44, for example, of this current spatial position of scene segment 28 and a reprojection map to be applied to the extracted and decoded media content so as to be mapped to the area forming observation segment 28. Reprojector 44 may accordingly apply mapping and interpolation to, for example, a regular grid of pixels to be displayed on display 26.

[0040] Figure 1 The case where a spatial scene 30 has been mapped onto a tile 50 using a cube mapping is illustrated. The tile is thus depicted as a rectangular sub-region of a cube onto which the scene 30 in the form of a sphere has been projected. The reprojector 44 reverses this projection. However, other examples may also be applied. For example, instead of a cube projection, a projection onto a truncated cone or a cone without truncation may be used. Furthermore, although Figure 1 The tiles are depicted as non-overlapping with respect to the overlay spatial scene 30, but the subdivision into tiles may involve overlapping of tiles with each other. And as will be outlined in more detail below, the spatial subdivision of the scene 30 into tiles 50 (each tile forming a representation as will be further explained below) is also not mandatory.

[0041] Therefore, if Figure 1 As depicted in FIG. 5 , the entire spatial scene 30 is spatially subdivided into tiles 50. Figure 1 In the example of , each of the six faces of the cube is subdivided into four tiles. For illustration purposes, the tiles are listed. For each tile 50, the server 20 provides a video 52, such as Figure 1 . For greater precision, the server 20 even provides more than one video 52 per tile 50, the quality Q# of these videos being different. Even further, the video 52 is temporally subdivided into time segments 54. The time segments 54 of all videos 52 of all tiles T# are respectively formed or encoded into one of the media segments of the plurality 46 of media segments stored in the memory 34 of the server 20.

[0042] Even reiterated, Figure 1The example of tile-based streaming described in is only an example from which many deviations are possible. Figure 1 It seems that the media fragments representing the higher quality of the scene 30 correspond to tiles that are consistent with the tiles to which the media fragments belong, in which the scene 30 is encoded with quality Q1, but this consistency is not necessary and tiles of different qualities may even correspond to tiles of different projections of the scene 30. In addition, although it has not been discussed so far, it is possible that Figure 1 The media segments corresponding to different quality levels depicted in may differ in spatial resolution and / or signal-to-noise ratio and / or temporal resolution, among other things.

[0043] Finally, unlike the tile-based streaming concept according to which the media segments that may be individually extracted by the device 40 from the server 20 are spatially subdivided with respect to the tiles 50 into which the scene 30 is subdivided, the media segments provided at the server 20 may alternatively, for example, each encode the scene 30 therein in a spatially complete manner with a spatially varying sampling resolution, which, however, is maximized at different spatial locations in the scene 30. This may be achieved, for example, by providing at the server 20 a sequence of segments 54 relating to projections of the scene 30 onto truncated cones, the frustums of which may be oriented in mutually different directions, resulting in oriented differently resolution peaks.

[0044] Furthermore, with regard to the optionally present flow modifier 38, it should be noted that the flow modifier may alternatively be part of the client 10, or the flow modifier may even be located between the client 10 and the server 20, within a network device, via which the signals described herein are exchanged.

[0045] After the systems of the server 20 and the client 10 have been explained more generally, the functionality of the client device 40 will be described in more detail with respect to an embodiment according to the first aspect of the present application. For this purpose, see Figure 2 , which shows the device 40 in more detail. As already explained above, the device 40 is used to stream media content about a temporally changing spatial scene 30. Figure 1As explained, the device 40 may be configured so that the media content being streamed is spatially continuous about the entire scene, or only about a segment 28 of the scene. In any case, the device 40 includes: a selector 56 for selecting an appropriate media segment 58 from a plurality 46 of media segments available on the server 20; and an extractor 60 for extracting the selected media segment from the server 20 by a corresponding request, such as an HTTP request. As described above, the selector 56 may use the media presentation description in order to calculate the addresses of the selected media segments, which the extractor 60 uses when extracting the selected media segments 58. For example, the calculation rule indicated in the media presentation description for calculating the address may depend on the quality parameter Q, the tile T and a certain time segment t. For example, the address may be a URL.

[0046] As also discussed above, the selector 56 is configured to perform the selection such that the selected media segment has at least a spatial segment of a temporally varying spatial scene encoded therein. The spatial segment may cover the entire scene spatially contiguously. Figure 2 An exemplary case in which the device 40 adapts a spatial segment 62 of the scene 30 to overlap and surround the observation segment 28 is illustrated at 61. However, as already mentioned above, this is not necessarily the case and the spatial segment may cover the entire scene 30 continuously.

[0047] Furthermore, the selector 56 performs the selection such that the selected media segment has segments 62 encoded therein with spatially unequal quality. More precisely, the first part 64 of the spatial segment 62 (at Figure 2) is encoded into the selected media segment at a predetermined quality. This quality may be, for example, the highest quality provided by the server 20 or may be a "good" quality. For example, the device 42 moves or adapts the first portion 64 in a manner that spatially follows the temporally changing observation segment 28. For example, the selector 56 selects the current time segment 54 of those tiles that inherit the current position of the observation segment 28. Having so selected, the selector 56 may optionally keep the number of tiles that make up the first portion 64 constant, as explained below with respect to other embodiments. In any case, the second portion 66 of the segment 62 is encoded into the selected media segment 58 at another quality, such as a lower quality. For example, the selector 56 selects a media segment corresponding to a current time segment of tiles that are spatially adjacent to the tiles of the portion 64 and that belong to tiles of lower quality. For example, to address a possible moment when observation segment 28 moves so quickly that it leaves portion 64 and overlaps portion 66 before the end of the time interval corresponding to the current time segment, and selector 56 will be able to spatially reconfigure portion 64, selector 56 primarily selects the media segment corresponding to portion 66. In this case, the portion of segment 28 that protrudes into portion 66 may still be presented to the user (i.e., at reduced quality).

[0048] It is not possible for device 40 to assess to the user what kind of negative quality degradation may be caused by pre-presenting to the user scene content of reduced quality and scene content within portion 64 having a higher quality. In particular, a transition between these two qualities is produced that is clearly visible to the user. At least, this transition may be visible depending on the current scene content within segment 28. The severity of the negative impact of this transition within the user's field of view is a characteristic of the scene content provided by server 20 and may not be predicted by device 40.

[0049] Therefore, according to Figure 2In an embodiment of the present invention, the device 40 includes a deriver 66 that derives a predetermined relationship that will be satisfied between the quality of the portion 64 and the quality of the portion 66. The deriver 66 derives this predetermined relationship from information that may be contained in the media segment (such as within a transport block within the media segment 58) and / or contained in a signaling obtained from the server 20 (such as within a media presentation description or a dedicated signal (such as within a SAND message) sent from the server 20, etc.). An example of how the information 68 looks is presented below. The predetermined relationship 70 derived by the deriver 66 based on the information 68 is used by the selector 56 in order to perform the selection appropriately. For example, the restriction in selecting the quality of the portions 64 and 66 affects the distribution of the available bandwidth for extracting the media content for the segment 62 onto the portions 64 and 66 compared to a completely independent selection of the quality of the portions 64 and 66. In any case, the selector 56 selects the media segment so that the quality used to encode the portions 64 and 66 into the final extracted media segment satisfies the predetermined relationship. Examples of how a predetermined relationship may look are also set forth below.

[0050] The media segments selected and finally extracted by the extractor 60 are finally forwarded to one or more decoders 42 for decoding.

[0051] For example, according to a first example, the signaling mechanism embodied by information 68 involves information 68 indicating to device 40 (which may be a DASH client) which quality combinations are acceptable for the provided video content. For example, information 68 may be a list of quality pairs indicating to the user or device 40 that different regions 64 and 66 may be mixed with the maximum quality (or resolution) difference. Device 40 may be configured to inevitably use a specific quality level (such as the highest quality level provided at server 10) for portion 64, and derive the quality level used to encode portion 66 into the selected media segment from information 68, wherein the information is included in the form of a list of quality levels for portion 68, for example.

[0052] Information 68 may indicate a tolerable value for a measure of the difference between the quality of portion 68 and the quality of portion 64. As a "measure" of the quality difference, a quality index of media segment 58 may be used, by which the media segments may be distinguished in the media presentation description, and the addresses of the media segments are calculated by the quality index using the calculation rules described in the media presentation description. In MPEG-DASH, the corresponding attribute indicating the quality may be, for example, @qualityRanking. Device 40 may take into account the restrictions on the selectable pairs of quality levels with which portions 64 and 66 may be encoded into the selected media segment when performing the selection.

[0053] However, instead of this difference metric, the quality difference may alternatively be measured, for example, as a bitrate difference (i.e., a tolerable difference in the bitrates used to encode portions 64 and 66 into corresponding media segments, respectively), assuming that bitrate generally increases monotonically with increasing quality. Information 68 may indicate the allowed pairings of options for the qualities used to encode portions 64 and 66 into the selected media segment. Alternatively, information 68 simply indicates the allowed quality for encoding portion 66, thereby indirectly indicating an allowed or tolerable quality difference, assuming that main portion 64 is encoded using some default quality (such as the highest quality possible or achievable). For example, information 68 may be a list of acceptable representation IDs or may indicate a minimum bitrate level for the media segment associated with portion 66.

[0054] However, alternatively, a more gradual quality difference may be desired, wherein, instead of a quality pair, a quality group (more than two qualities) may be indicated, wherein, depending on the distance from the segment 28 (viewport), the quality difference may increase. That is, the information 68 may indicate tolerable values ​​for a measure of the difference between the qualities of the portions 64 and 66 in a manner that depends on the distance from the observation segment 28. This may be done by a list of pairs of respective distances from the observation segment and corresponding tolerable values ​​for a measure of the quality difference exceeding the respective distance. Below the respective distance, the quality difference must be lower. That is, each pair may indicate for a respective distance that a portion of the portion 66 that is further from the segment 28 than the respective distance may have a quality difference with the quality of the portion 64 that exceeds the respective tolerable value of this list entry.

[0055] The tolerable value may increase with increasing distance from the observation segment 28. The acceptability of the quality differences just discussed often depends on the time at which these different qualities are displayed to the user. For example, content with a high quality difference may be acceptable if the content is only displayed for 200 microseconds, while content with a lower quality difference may be acceptable if the content is displayed for 500 microseconds. Therefore, according to another example, in addition to the aforementioned quality combinations, or in addition to the allowed quality differences, the information 68 may also include time intervals in which the combination / quality difference may be acceptable. In other words, the information 68 may indicate the tolerable or maximum allowed difference between the quality of the portions 66 and 64, as well as an indication of the maximum allowed time interval in which the portion 66 may be displayed simultaneously with the portion 64 in the observation segment 28.

[0056] As already mentioned previously, the acceptability of quality differences depends on the content itself. For example, the spatial position of different tiles 50 has an impact on the acceptability. Quality differences in uniform background areas with low-frequency signals are expected to be more acceptable than quality differences in foreground objects. In addition, the temporal position also has an impact on the acceptability due to changing content. Therefore, according to another example, the signal forming the information 68 is sent to the device 40 intermittently (such as, per representation or period in DASH). That is, the predetermined relationship indicated by the information 68 can be updated intermittently. In addition and / or alternatively, the signaling mechanism implemented by the information 68 can be varied in space. That is, the information 68 can be made spatially dependent, such as, through the SRD parameter in DASH. That is, different predetermined relationships can be indicated by the information 68 for different spatial areas of the scene 30.

[0057] As about Figure 2 As described, an embodiment of device 40 is concerned with the fact that device 40 wishes to keep the quality degradation caused by pre-fetched portion 66 within extracted segment 62 of video content 30 that is briefly visible in segment 28 as low as possible before being able to change the position of segment 62 and portion 64 in order to adapt the segment and portion to the position change caused by segment 28. That is, Figure 2 , portions 64 and 66 are different portions of segment 62 whose quality is limited until their possible combination is of interest to information 68, while the transition between the two portions 64 and 66 is continuously shifted or adapted so as to track or exceed the moving observation segment 28. Figure 3 In the alternative embodiment shown in FIG. 4 , the device 40 uses the information 68 to control the possible combination of the qualities of the portions 64 and 66 , however, according to Figure 3 In an embodiment of the present invention, portions 64 and 66 are defined as being different or distinct portions from one another in a manner defined, for example, in a media presentation description (i.e., in a manner independent of the position of viewing segment 28). The positions of portions 64 and 66 and the transitions therebetween may be constant or vary in time. If varying in time, the variations are due to changes in the content of scene 30. For example, portion 64 may correspond to a region of interest that merits consuming a higher quality, while portion 66 is a portion whose quality reduction, for example due to low bandwidth conditions, should be considered before considering the quality reduction of portion 64.

[0058] In the following, another example of an advantageous implementation of the device 40 is described. In particular, Figure 4 Display device 40, which is structurally similar to Figure 2 and 3 Corresponding, but the operation mode is changed to correspond to the second aspect of the application.

[0059] That is, device 40 includes a selector 56, an extractor 60, and a deriver 66. Selector 56 makes a selection from a plurality 46 of media segments 58 provided by server 20, and extractor 60 extracts the selected media segment from the server. Figure 4 Assume that the device 40 is as described above. Figure 2 and Figure 3 The depicted and described operation is that the selector 56 performs the selection so that the selected media segment 58 encodes therein a spatial segment 62 of the scene 30 in such a way that this spatial segment follows the observation segment 28 whose spatial position changes in time. Figure 5 A variant corresponding to the same aspect of the present application is described, in which, for each time instant t, the selected and extracted media segment 58 has the entire scene or a constant spatial segment 62 encoded therein.

[0060] In any case, similar to Figure 2 and 3 , selector 56 selects media segments 58 such that a first portion 64 within segment 62 is encoded into the selected and extracted media segment at a predetermined quality, while a second portion 66 of segment 62 (which is spatially adjacent to first portion 64) is encoded into the selected media segment at a reduced quality relative to the predetermined quality of portion 64. Figure 6 , in which selector 56 restricts the selection and extraction of media segments for a moving template to tracks the position of viewport 28, and in which the media segments have segments 62 fully encoded therein at a predetermined quality such that first portion 64 completely covers segment 62 while being surrounded by unencoded portion 72. In any case, selector 56 performs the selection such that first portion 64 follows a viewing segment 28 whose spatial position changes in time.

[0061] In this case, it is also not easy for the client 40 to predict how large the segment 62 or portion 64 should be. Depending on the scene content, most users may make similar movements when moving the observation segment 28 across the scene 30, and therefore, the same movement applies to the interval of the observation segment 28, and the observation segment 28 is likely to move across the scene 30 at this speed. Therefore, according to Figures 4 to 6 In an embodiment, the information 74 is provided by the server 20 to the device 40 in order to help the device 40 to set the size or size and / or position of the first part 64 or the size or size and / or position of the section 62, respectively, depending on the information 74. As for the possibility of transmitting the information 74 from the server 20 to the device 40, as described above with respect to Figure 2 and Figure 3That is, the information may be contained within the media segment 58, such as within an event block of the media segment, or for this purpose, a media presentation description or a transmission within a dedicated message (such as a SAND message) sent from the server to the device 40 may be used.

[0062] That is, according to Figures 4 to 6 In an embodiment of the present invention, the selector 56 is configured to set the size of the first portion 64 depending on the information 74 originating from the server 20. Figures 4 to 6 In the embodiment described in FIG. , the size is set in units of tiles 50, but as described above with respect to Figure 1 As has been described, the situation may be slightly different when using the further concept of providing a scene 30 of spatially varying quality at the server 20 .

[0063] According to an example, the information 70 may, for example, include the probability of a given movement speed of the viewport observing the segment 28. As already indicated above, the information 74 may result in the media presentation description being available to the client device 40, which may be, for example, a DASH client, or some in-band mechanism may be used to deliver the information 74, such as an event block, i.e., an EMSG or SAND message in the case of DASH. The information 74 may also be included in any container format, such as an ISO file format or a transport format beyond MPEG-DASH, such as MPEG-2 TS. The information may also be delivered in the video bitstream, such as, for example, in an SEI message as described later. In other words, the information 74 may indicate a predetermined value of a measure of the spatial speed for the observation segment 28. In this way, the information 74 indicates the size of the portion 64, either in the form of a scaling relative to the size of the observation segment 28 or in the form of an increment relative to the size of the observation segment 28. That is, information 74 starts with a "base size" of portion 64 necessary to cover the size of segment 28, and increases this "base size" appropriately (such as incrementally or proportionally). For example, the aforementioned speed of movement of observation segment 28 may be used to scale the perimeter of the current position of observation segment 28 accordingly, so as to determine, for example, the farthest position of the perimeter of observation segment 28 along any spatial direction feasible after this time interval, e.g., to determine a delay in adjusting the spatial position of portion 64, such as the duration of time segment 54 corresponding to the temporal length of media segment 58. The speed multiplied by this duration plus the perimeter of the current position of the omnidirectional viewport 28 may therefore result in this worst-case perimeter and may be used to determine an enlargement of portion 64 relative to some minimum extension of portion 64 assuming a non-moving viewport 28.

[0064] The information 74 may even be about an evaluation of the statistics of the user's behavior. Subsequently, an embodiment suitable for feeding such an evaluation program is described. For example, the information 74 may indicate the maximum speed for a certain percentage of users. For example, the information 74 may indicate that 90% of the users move at a speed lower than 0.2 radians / second and 98% of the users move at a speed lower than 0.5 radians / second. The information 74 or the message carrying the information may be defined so that a probability-speed pair is defined or the message may be defined to signal the maximum speed of a fixed percentage of users (e.g., always 99% of the users). The movement speed signaling 74 may additionally include direction information, i.e., angle in 2D, or depth in 2D plus 3D (also known as light field application). The information 74 may indicate different probability-speed pairs for different movement directions.

[0065] In other words, the information 74 may apply to a given time span, such as the time length of a media segment. The information may consist of a track-based (x percent, average user path) or speed-based pairing (x percent, speed) or distance-based pairing (x percent, pore / diameter / preferred) or area-based pairing (x percent, recommended preferred area) or a single maximum boundary value for path, speed, distance, or preferred area. Instead of associating the information with percentages, a simple frequency ranking may be performed based on most users moving at a certain speed, the second most users moving at another speed, etc. Additionally or alternatively, the information 74 is not limited to indicating the speed of the observation segment 28, but may also indicate preferred areas to be observed separately to guide efforts to track portions 62 and / or 64 of the observation segment 28, with or without an indication of the statistical significance of the indication (such as the percentage of users who have complied with that indication or whether the indication is consistent with the most frequently recorded user observation speed / observation segment), and with or without the time duration of the indication. The information 74 may indicate another measure of the speed of the observation segment 28, such as a measure of the travel distance of the observation segment 28 within a specific time period (such as within the time length of the media segment, or more specifically within the time length of the time segment 54). Alternatively, the information 74 may be notified in a manner that distinguishes between specific movement directions in which the observation segment 28 may travel. This is in relation to both indicating the speed or velocity of the observation segment 28 in a specific direction and indicating the travel distance of the observation segment 28 with respect to a specific movement direction. In addition, the extension of the portion 64 may be directly signaled by the information 74 omnidirectionally or in a manner that distinguishes between different movement directions. In addition, all of the examples just outlined may be modified, wherein the information 74 indicates these values ​​and percentages of users, which are sufficient for the user to explain the statistical behavior when moving the observation segment 28. In this regard, it should be noted that the observation speed (that is, the speed of the observation segment 28) may be quite large and is not limited to, for example, the speed value of the user's head. In practice, the observation segment 28 may move depending on, for example, the user's eye movement, in which case the observation speed may be significantly larger. The viewing section 28 may also be moved according to another input device movement, such as according to the movement of a tablet computer, etc. Since all of these "input possibilities" that enable the user to move the section 28 result in different expected speeds of the viewing section 28, the information 74 may even be designed so that it distinguishes different concepts for controlling the movement of the viewing section 28. That is, the information 74 may indicate the size of the portion 64 in a manner that indicates different sizes of different methods for controlling the movement of the viewing section 28, and the device 40 may use the size indicated by the information 74 for the correct viewing section control.That is, the device 40 obtains knowledge about the way in which the viewing segment 28 is controlled by the user, i.e. checks whether the viewing segment 28 is controlled by head movement, eye movement or tablet movement or similar movement, and sets the size according to the part of the information 74 corresponding to such viewing segment control.

[0066] In general, movement speed may be signaled per content, per period, per representation, per segment, per SRD position, per pixel, per tile (e.g., at any temporal or spatial granularity, etc.). Movement speed may also distinguish between head movement and / or eye movement, as just outlined. In addition, information 74 on the probability of user movement may be conveyed as a recommendation for high-resolution prefetching (i.e., video area outside the user's viewport, or sphere coverage).

[0067] Figures 7a to 7c Some of the options explained with respect to information 74 in terms of its use by device 40 to modify the size of portion 64 or portion 62, respectively, and / or the position of said portion are briefly summarized. Figure 7a , device 40 enlarges the perimeter of segment 28 by a distance corresponding to the product of the signaled velocity v and a duration Δt, which may correspond to a time period corresponding to the time length of time segment 54 encoded in individual media segment 50a. Additionally and / or alternatively, the greater the velocity, the further away from the current position of segment 28 the position of portion 62 and / or 64 may be placed, or in the direction of the signaled velocity or movement, as signaled by information 74. The velocity and direction may be derived from measuring or extrapolating recent developments or changes in the recommended preferred area indicated by information 74. Instead of applying v×Δt omnidirectionally, the velocity may be signaled differently by information 74 for different spatial directions. Figure 7b The alternative example depicted in FIG. 7 shows that the information 74 can directly indicate the distance of the perimeter of the magnified observation section 28, which is determined by Figure 7b The parameter s in indicates that . Again, an amplification of the directional changes of the segments can be applied. Figure 7c 28, the enlargement of the perimeter of the section 28 may be indicated by the information 74 through an increase in area, such as in the form of a ratio of the area of ​​the enlarged section compared to the original area of ​​the section 28. In any case, the area 28 is enlarged in the perimeter (in Figures 7a to 7c 76) may be used by selector 56 to size or set the size of portion 64 so that portion 64 covers at least the entire area within the enlarged section 76 by a predetermined amount. Obviously, the larger the section 76, for example, the greater the number of tiles within portion 64. According to another alternative, section 74 may directly indicate the size of portion 64, such as in the form of the number of tiles that make up portion 64.

[0068] exist Figure 5A further possibility of signaling the size of the signaling part 64 is depicted. Figure 5 The embodiment of Figure 4 can be modified in a manner similar to the modification of the embodiment of Figure 6 ; that is, the entire area of the section 62 can be retrieved from the server 20 by the segment 58 with the quality of the part 64.

[0069] In any case, at the end of Figure 5 , the information 74 differentiates the different sizes of the viewing section 28, that is, the different fields of view seen by the viewing section 28. The information 74 simply indicates the size of the part 64 depending on the size of the viewing section 28 currently aimed at by the device 40. This enables the service of the server 20 to be used by devices having different fields of view or different sizes of the viewing section 28 without a device such as the device 40 having to deal with calculating or otherwise guessing the size of the part 64 such that the part 64 is sufficient to cover the viewing section 28 regardless of any movement of the section 28 (as discussed with respect to Figure 4 , Figure 6 and FIG. 7). As will become clear from the description of Figure 1 , it is easy to evaluate, for example, which constant number of tiles may be sufficient to completely cover a particular size of the viewing section 28 (i.e., a particular field of view) regardless of the orientation of the viewing section 28 for spatial positioning 30. Here, the information 74 alleviates this situation, and the device 40 can simply look up the value of the size of the part 64 in the information 74 for the size of the viewing section 28 applied to the device 40. That is, according to the embodiment of Figure 5 , a media presentation description (such as an event chunk or a SAND message) available for use by a DASH client or some interested institution may include the information 74 regarding the sphere coverage or the field of view of a set of representations or a set of tiles, respectively. An example may be to provide M representations of tiles, as depicted in Figure 1 . The information 74 may indicate a recommended number n < M tiles (referred to as representations) to be downloaded for covering the field of view of a given terminal device. For example, in a cube representation of tiles partitioned into 6×4 tiles as depicted in Figure 1 , it is considered that 12 tiles are sufficient to cover a 90°×90° field of view. Due to the fact that the field of view of the terminal device may not always align perfectly with the tile boundaries, this recommendation cannot be trivially generated by the device 40 itself. The device 40 may use the information 74 by downloading, for example, at least N tiles, that is, the media segment 58 is related to N tiles. Another way of using the information may be to focus on the quality of the N tiles closest to the current viewing center of the terminal device within the section 62, that is, using the N tiles to form the part 64 of the section 62.

[0070] Regarding Figure 8a , an embodiment regarding another aspect of the present application is described. Here, Figure 8a The client device 10 and the server 20 are shown, and the two are based on the above Figure 1 7 to communicate with each other. That is, the devices 10 may communicate with each other according to the Figure 2 7, or may be implemented without these details as described above with respect to Figure 1 However, the device 10 is advantageously configured according to the above description. Figure 2 7 or any combination thereof, and further inherits the present Figure 8a In particular, the device 10 is understood internally as described above with respect to Figure 2 8, that is, the device 40 includes a selector 56, an extractor 60 and optionally a deriver 66. The selector 56 performs the selection for targeting unequal streaming, that is, selecting the media segments in a way that the media content is encoded into the selected and extracted media segments in a way that the quality varies spatially and / or there are unencoded parts. However, in addition to this, the device 40 also includes a log message sender 80 that sends log messages recorded in (for example) the following to the server 20 or the evaluation device 82:

[0071] measuring instantaneous measurements or statistics of the spatial position and / or movement of the first portion 64,

[0072] measuring instantaneous measurements or statistics of the quality of the time-varying spatial scene up to encoding into the selected media segment and up to being visible in the observation section 28, and / or

[0073] An instantaneous measurement or statistical value of the quality of the first portion or of the temporally varying spatial scene 30 up to the encoding into the selected media segment and up to the visible observation section 28 is measured.

[0074] The motivation is as follows.

[0075] In order to be able to derive statistics, such as the most interesting regions or speed-probability pairs, a reporting mechanism from the user is required as described previously. Additional DASH metrics to the statistics defined in Annex D of ISO / IEC 23009-1 are required.

[0076] One metric may be the client's field of view which is a DASH metric, where the DASH client sends characteristics of the terminal device regarding the field of view back to a metric server (which may be the same as the DASH server or another server).

[0077] Keywords type describe EndDeviceFoVH Integer Horizontal field of view of the terminal device, in degrees EndDeviceFoVV Integer Vertical field of view of the terminal device, in degrees

[0078] One metric may be a ViewportList, where the DASH client sends the viewports seen by each client back to the metric server (which may be the same as the DASH server or another server) in a timely manner. An instantiation of this message may be as follows.

[0079]

[0080] For viewport (region of interest) messages, a DASH client may be required to report when viewport changes occur, possibly with a given granularity (to avoid or not avoid reporting very small movements) or with a given periodicity. This message may be included in the MPD as an attribute @reportViewPortPeriodicity or element or descriptor. This message may also be indicated out-of-band, such as using a SAND message or any other means.

[0081] The viewport can also be signaled regarding tile granularity.

[0082] Additionally or alternatively, the log message may report on other current scene related parameters that change in response to user input, such as described below with respect to Fig.10 Any of the parameters discussed, such as the current user distance from the center of the scene and / or the current viewing depth.

[0083] Another metric may be ViewportSpeedList, where the DASH client indicates the movement speed of a given viewport when movement occurs.

[0084]

[0085]

[0086] This message may only be sent when the client performs a viewport move. However, as with the previous case, the server may indicate that the message should only be sent when the movement is significant. This configuration may be somewhat similar to @minViewportDifferenceForReporting, used to signal the size in pixels or degrees or any other quantity that needs to change for the message to be sent.

[0087] Another important thing for VR-DASH services (where asymmetric quality as described above is provided) is to assess how quickly a user switches from an asymmetric representation or set of unequal quality / resolution representations of a viewport to another representation or set of representations that is more adequate for another viewport. Using this metric, the server can derive statistics that help it understand the relevant factors that affect QoE. This metric may look like the following.

[0088]

[0089] Alternatively, the previously described durations may be given as average values.

[0090]

[0091]

[0092] As with other DASH metrics, all such metrics may additionally have a time over which the measurement has been executed.

[0093] t real time The time at which the parameter was measured.

[0094] In some cases, it may happen that if unequal quality content is downloaded and bad quality (or a mix of good and bad quality) is shown for long enough (it may be only a few seconds), the user gets unhappy and leaves the session. Under the condition of leaving the session, the user may send a message with the quality shown in the last x time intervals.

[0095]

[0096] Alternatively, the maximum quality difference may be reported or the maximum and minimum quality of a viewport may be reported.

[0097] As from about Figure 8a As becomes clear from the above description, in order for a tile-based DASH streaming service operator to configure and optimize its service in a meaningful way (e.g., with respect to resolution ratio, bitrate, and segment duration), it is advantageous for the service operator to be able to derive statistics that require the client reporting mechanism described above. Additional DASH metrics in addition to the metrics defined above and in addition to Appendix D of [A1] are set out below.

[0098] Imagine a tile-based streaming service using Figure 1 The client-side reconstruction is shown in the video of the cube projection depicted in Figure 8b 6, wherein small circles 198 indicate the projection of the two-dimensional distribution of viewing directions within the client's viewport 28, distributed equiangularly horizontally and vertically, onto the image area covered by the individual tiles 50. Tiles marked with hatching indicate high resolution tiles, thus forming the high resolution portion 64, while tiles 50 shown unhatched represent low resolution tiles, thus forming the low resolution portion 66. It can be seen that as the viewport 28 changes, the user is partially presented with low resolution tiles, since the most recent update to the fragment selection and download determines the resolution of each tile on the cube upon which the projection plane or pixel array of the tiles 50 encoded into the downloadable fragment 58 falls.

[0099] While the above description actually indicates generally (among other things) feedback or log messages indicating the quality of the video presentation to the user in the viewport, in the following, more specific and advantageous metrics applicable in this regard will be outlined. The metric now described may be reported back from the client side and is referred to as the effective viewport resolution. Presumably the metric indicates to the service operator the effective resolution in the client's viewport. In the event that the reported effective viewport resolution indicates that the user is only presented with a resolution towards the resolution of the low-resolution tiles, the service operator may change the tile configuration, resolution ratio or segment length accordingly to achieve a higher effective viewport resolution.

[0100] One embodiment may be the average pixel count in the viewport 28 measured in the projection plane in which the pixel array of the tile 50 encoded into the fragment 58 falls. The measurement may differentiate or be specific to the horizontal direction 204 and the vertical direction 206 relative to the covered field of view (FoV) of the viewport 28. The following table shows possible examples of suitable syntax and semantics that may be included in the log message in order to signal the summarized viewport quality metric.

[0101]

[0102] The decomposition in the horizontal and vertical directions may be stopped by using a scalar value of the average pixel count instead.Also reported to the recipient of the log message (i.e., evaluator 82) along with the indication of the aperture or size of the viewport 28 is the average count, which indicates the pixel density within the viewport.

[0103] It may be advantageous to reduce the field of view considered for the metric to be smaller than the field of view of the viewport actually presented to the user, thereby excluding areas towards the border of the viewport that are used only for peripheral vision and therefore have no impact on subjective quality perception. This alternative is illustrated by the dashed line 202, which encloses pixels that are in this central segment of the viewport 28. Reports of the considered field of view 202 for reported metrics about the total field of view of the viewport 28 may also be signaled to the log message recipient 82. The following table shows the corresponding extensions of the previous example.

[0104]

[0105]

[0106] According to another embodiment, the average pixel density is not measured by averaging the quality in a spatially uniform manner in the projection plane (as is actually the case in the examples described so far with respect to the examples containing EffectiveFoVResolutionH / V), but rather in a way that this averaging is weighted in a non-uniform manner for the pixels (i.e., the projection plane). The averaging can be performed in a spherical uniform manner. As an example, the averaging can be performed uniformly with respect to sample points distributed as circle 198. In other words, the averaging can be performed by weighting the regional density with a weight that decreases in a quadratic manner with increasing local projection plane distance and increases according to the sine of the local tilt of the projection relative to the line connected to the viewport. The message may include an optional (flag-controlled) step size to adjust for the inherent oversampling of some of the available projections (such as equirectangular projections), for example by using a uniform spherical sampling grid. Some projections do not have a large oversampling problem, and forcing the calculation to remove the oversampling may create unnecessary complexity problems. This must not be limited to equirectangular projections. The report does not need to distinguish between horizontal and vertical resolutions, but can combine them. An embodiment is given below.

[0107]

[0108]

[0109] exist Figure 8e , the application of conformal uniformity in averaging is illustrated by showing how conformally horizontally and vertically distributed points 302 (within the viewport 28) on a sphere 304 centered on the viewport 306 are projected onto a projection plane 308 of the tile (here a cube) in order to perform an averaging of the pixel densities 308 of the pixels arranged in an array (in rows and columns) in the projection area in order to set local weights for the pixel densities based on the local density of the projection 198 of the point 302 onto the projection plane. Figure 8f A very similar approach is depicted in . Here, points 302 are equally spaced in a viewport plane 310 perpendicular to the viewing direction 312, i.e., evenly spaced horizontally and vertically in rows and columns, and the projection onto a projection plane 308 defines point 198, with the local density of the point controlling the weights, with which the local density pixel density 308, which varies due to high and low resolution tiles within the viewport 28, contributes to the averaging. In the above example such as the latest table, Figure 8f An alternative example can be replaced Figure 8e The examples depicted in are used.

[0110] In the following, another embodiment of classified log messages is described, which is related to a DASH client 10 having multiple media buffers 300, such as Figure 8aAs exemplarily depicted in FIG. 1 , the DASH client 10 forwards the downloaded segments 58 to subsequent decoding by one or more decoders 42 (compare Figure 1 ). The distribution of the segments 58 onto the buffers can be done in different ways. For example, the distribution can be made so that specific areas of the 360 ​​video are downloaded separately from each other, or cached after being downloaded into separate buffers. The following examples illustrate different distributions by indicating which are the following: Figure 1 The tiles T indexed #1 to #24 shown in the MPD (the enantiomer has a total of 25) are encoded into which individual downloadable representations R#1 to #P with which quality Q from #1 to #M (1 being the best and M being the worst), and how these P representations R can be grouped into adaptation sets A indexed #1 to #S in the MPD (optional), and how the fragments 58 of the P representations R can be distributed to the buffers of the buffer B indexed #1 to #N.

[0111]

[0112] Here, representations may be provided at the server and announced in the MPD for download, each of these representations relating to one tile 50, i.e. one segment of the scene. The representations relating to one tile 50, but encoding this tile 50 at different qualities, may be summarized in an adaptation set grouped as optional, but precisely this grouping is used to associate to the buffer. Thus, according to this example, there will be one buffer per tile 50, or in other words per viewport (observation segment) encoding.

[0113] Another representation of the set and distribution can be:

[0114]

[0115] According to this example, each representation may cover the entire region, but the high quality region will be focused on one hemisphere, while lower quality is used for the other hemisphere. Representations that differ only in the exact quality used in this way (i.e., equal in the position of the higher quality hemisphere) may be collected in an adaptive set and distributed to buffers (here illustratively, six) according to this characteristic.

[0116] Therefore, the following description assumes that this distribution to buffers according to different viewport encodings (video sub-regions, such as tiles) associated with adaptation sets or the like is applied. Figure 8c Buffer fullness levels over time for two separate buffers (e.g., Tile 1 and Tile 2) in a tile-based streaming scenario are illustrated, this scenario being illustrated in the last but not least table. Enabling clients to report fullness levels for all their buffers allows service operators to correlate the data with other streaming parameters to understand the Quality of Experience (QoE) impact of their service setup.

[0117] The advantage is that the buffer fullness of multiple media buffers on the client side can be reported using metrics and identified and associated to buffer types. For example, the associated types are as follows:

[0118] Tiles

[0119] Viewport

[0120] ·district

[0121] Adaptive Set

[0122] ·express

[0123] Low-quality versions of the full content

[0124] One embodiment of the present invention is presented in Table 1, which defines the metrics used to report buffer level status events for each buffer identified and associated.

[0125] Table 1: List of buffer levels

[0126]

[0127] Another embodiment using viewport-dependent encoding is as follows.

[0128] In a viewport-dependent streaming scenario, a DASH client downloads and pre-buffers several media segments that are relevant to a specific viewing orientation (viewport). If the amount of pre-buffered content is too high, and the client changes its viewing orientation, the portion of the pre-buffered content that will be played after the viewport change is not presented and the corresponding media buffer is cleared. This scenario is depicted in Figure 8d middle.

[0129] Another embodiment may pertain to a traditional video streaming scenario with multiple representations (quality / bitrate) of the same content and the quality used to encode the video content may be spatially uniform.

[0130] The distribution thus looks like this:

[0131]

[0132] That is, here, each representation may cover, for example, a complete scene that may not be a panoramic 360 scene, with a different quality (i.e., spatially uniform quality), and these representations may be distributed to the buffer individually. All examples stated in the last three tables should be considered as not limiting the way in which the fragments 58 of the representation provided at the server are distributed to the buffer. There are different methods, and the rules may be based on the affiliation of the fragments 58 to the representation, the affiliation of the fragments 58 to the adaptation set, the direction of locally increasing quality of encoding the space of the scene unevenly into the representation to which the corresponding fragment belongs, the quality used to encode the scene into the corresponding fragment, etc.

[0133] The client may maintain a buffer for each representation and, after experiencing an increase in available throughput, decide to clear the remaining low quality / bitrate media buffer before playback and download high quality media segments with duration into the existing low quality / bitrate buffer. Similar embodiments may be constructed for tile-based streaming and viewport-dependent encoding.

[0134] The service operator may be interested in understanding what amount and what kind of data is downloaded but not presented, because this introduces a cost without gain on the server side and reduces the quality on the client side. Therefore, the present invention should provide reporting metrics that relate the two events "media download" and "media presentation" to be easily interpreted. The present invention avoids the information about each media segment download and playback status being analyzed tediously, and only allows the clearing events to be reported efficiently. The present invention also includes the identification of buffers as described above and the association to types. An embodiment of the present invention is given in Table 2.

[0135] Table 2: List of clearing events

[0136]

[0137]

[0138] Fig. 9 Another embodiment showing how the device 40 can be advantageously implemented, Fig. 9 The device 40 may correspond to the above Figure 1 to any of the examples set forth in FIG. 8. That is, the device may include Figure 8a The log sender discussed above, but not necessarily includes, and can use the above-mentioned Figure 2 and Figure 3 The information 68 discussed above or Figures 5 to 7c The information discussed 74 is not necessarily used. However, unlike Figure 2 To the description of FIG. 8, about Fig. 9, assuming that a tile-based streaming method is actually applied. That is, the scene content 30 is provided at the server 20 in a tile-based manner, which is discussed above with respect to Figure 2 To the options of Figure 8.

[0139] Although the internal structure of the device 40 may be different from Fig. 9 , but the device 40 is exemplarily shown as including the internal structure described above with respect to Figure 2 The selector 56 and extractor 60 discussed to Figure 8, and optionally include a deriver 66. However, in addition, the device 40 includes a media presentation description analyzer 90 and a matcher 92. The MPD analyzer 90 is used to derive from the media presentation description obtained from the server 20: at least one version, in which the time-varying spatial scene 30 is provided for tile-based streaming; and for each of the at least one version, an indication of the beneficial requirements of the corresponding version of the time-varying spatial scene for benefiting from the tile-based streaming. The meaning of "version" will become clear from the following description. In particular, the matcher 92 matches the beneficial requirements thus obtained with the device capabilities of the device 40 or another device interacting with the device 40 (such as, the decoding capabilities of one or more decoders 42, the number of decoders 42, or the like). Fig. 9The background or idea behind the concept of is as follows. It is envisioned that, assuming a specific size of the observed segment 28, the tile-based approach requires a specific number of tiles to be contained by the segment 62. In addition, it can be assumed that the media segments belonging to a specific tile form a media stream or video stream, which should be decoded by a separate decoding instance, separate from the decoding of the media segments belonging to another tile. Therefore, the mobile aggregation of a specific number of tiles within the segment 62 selected by the selector 56 for the corresponding media segment requires a specific decoding capability, such as the presence of corresponding decoding resources in the form of, for example, a corresponding number of decoding instances (i.e., a corresponding number of decoders 42). If this number of decoders does not exist, the service provided by the server 20 is not available to the client. Therefore, the MPD provided by the server 20 may indicate a "beneficial requirement", that is, the number of decoders required to use the provided service. However, the server 20 may provide MPDs for different versions. That is, different MPDs for different versions may be obtained by the server 20, or the MPD provided by the server 20 may be structured internally so as to distinguish different versions that can be used by the service. For example, the versions may differ in the field of view (i.e., the size of the observation segment 28). The different sizes of the field of view manifest themselves in different numbers of tiles within the segment 62, and may therefore differ in the beneficial requirements, because, for example, these versions may require different numbers of decoders. Other examples are also conceivable. For example, while the versions with different fields of view may involve the same plurality of media segments 46, according to another example, the different versions of the scene 30 provided for tile streaming at the server 20 may even differ in the plurality 46 of media segments involved according to the corresponding versions. For example, the tile segmentation according to one version is coarser than the tile-hyphen segmentation of the scene according to another version, thereby requiring, for example, a smaller number of decoders.

[0140] The matcher 92 matches beneficial requirements and accordingly selects the corresponding version or completely rejects all versions.

[0141] However, the beneficial requirements may additionally focus on the profiles / levels that one or more decoders 42 must be able to address. For example, the DASH MPD includes multiple locations that allow profiles to be indicated. A typical profile describes the attributes, elements that may be present at the MPD, and the video or audio profiles for each media stream provided for each representation.

[0142] Other examples of beneficial requirements concern, for example, the client-side ability to move the viewport 28 across the scene. A beneficial requirement may indicate a necessary viewport speed that should be available to the user to move the viewport so that the provided scene content can be truly enjoyed. The matcher may check, for example, whether this requirement is met, e.g., inserted in a user input device such as an HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in a directional sense, a set of "sufficient types of input devices" may be indicated by a beneficial requirement.

[0143] In a tiled streaming service of spherical video, there are too many configuration parameters that can be set dynamically, such as the number of qualities, the number of tiles. In the case of tiles being independent bitstreams that need to be decoded by separate decoders, if the number of tiles is too high, a hardware device with several decoders will not be able to decode all bitstreams simultaneously. A possibility is to keep this as a degree of freedom, and the DASH device parses all possible representations and counts how many decoders are needed to decode all representations or a given number of fields of view of the coverage device, and thus derives whether the DASH client can consume the content. However, a smarter solution for interoperability and capability negotiation is to use signaling in the MPD mapped to a profile, which is used as a promise to the client that if the profile is supported, the provided VR content can be consumed. This signaling should be in the form of a URN (such as urn::dash-mpeg::vr::2016) that can be encapsulated at the MPD level or at an adaptation set. This parsing will mean that N decoders at the X profile are sufficient for consuming the content. Depending on the profile, the DASH client can ignore or accept the MPD or parts of the MPD (adaptation set). In addition, there are several mechanisms that do not include all the information, such as Xlink or MPD link, where little signaling for selection is available. In this case, the DASH client will not be able to figure out whether it can consume the content. The decoding capabilities need to be exposed through this urn (or something similar) about the number of decoders and the profile / level of each decoder so that the DASH client can now make sense of whether to execute Xlink or MPD link or similar mechanisms. Signaling can also mean different operating points, such as N decoders with X profile / level or Z decoders with Y profile / level.

[0144] Fig.10 Further description, the client, device 40, server, etc. Figures 1 to 9 Any of the above embodiments and descriptions presented can be extended to the extent that the service provided is extended to the extent that the time-varying spatial scene not only varies in time, but also varies depending on another parameter. For example, Fig.10 illustrate Figure 1A variation of , in which multiple available media segments are obtained on the server to describe the scene content 30 for different locations of the observation center 100. Fig.10 In the schematic diagram shown in , the scene center is depicted as changing only along one direction X, but it is obvious that the observation center may change along more than one spatial direction, such as two-dimensionally or three-dimensionally. For example, this corresponds to a change in the user position of a user in a specific virtual environment. Depending on the user position in the virtual environment, the view available changes and, accordingly, the scene 30 changes. Therefore, in addition to describing the scene 30 as being subdivided into tiles and time segments and media segments of different qualities, other media segments describe different contents of the scene 30 for different positions of the scene center 100. The device 40 or the selector 56 calculates the addresses of the media segments to be extracted within the selection procedure from the plurality of 46 media segments, respectively, depending on the observation segment position and at least one parameter, such as the parameter X, and may then use the calculated addresses to extract these media segments from the server. For this purpose, the media presentation description may describe a function that depends on the tile index, the quality index, the scene center position and the time t, and generates corresponding addresses for the corresponding media segments. Therefore, according to Fig.10 In an embodiment of the present invention, the media presentation description may include this calculation rule, in addition to the above Figures 1 to 9 In addition to the parameters described, the calculation rule also depends on one or more additional parameters. The parameter X may be quantized to any one of the levels for which the corresponding scene representation is encoded by a corresponding media segment within the plurality 46 of media segments in the server 20 .

[0145] Alternatively, X may be a parameter defining the viewing depth, i.e., the distance in the radial direction from the scene center 100. While providing the scene in different versions with different viewing center portions X allows the user to "walk" through the scene, providing the scene in different versions with different viewing depths may allow the user to "radially zoom" back and forth in the scene.

[0146] For multiple non-concentric viewports, the MPD may therefore use another signaling of the position of the current viewport. The signaling may be done at the segment, representation or period level or the like.

[0147] Non-concentric spheres: To allow the user to move, the spatial relationship of the different spheres should be signaled in the MPD. This signaling can be done with coordinates (x,y,z) in arbitrary units relative to the sphere diameter. In addition, the diameter of the sphere should be indicated for each sphere. A sphere may be "good enough" for a user in its center and extra space for which the content will be good. If the user can move beyond the signaled diameter, another sphere should be used for showing the content.

[0148] An exemplary signaling of viewports may be done relative to a predefined center point in space. Each viewport will be signaled relative to that center point. In MPEG-DASH, this may be signaled, for example, in an AdaptationSet element.

[0149]

[0150]

[0151]

[0152] at last, Fig.11 Information describing information such as or similar to the information described above with respect to reference numeral 74 may reside in a video bitstream 110 into which a video 112 is encoded. A decoder 114 decoding this video 110 may use the information 74 to determine the size of a focus region 116 within the video 112 to which decoding capabilities for decoding the video 110 should be focused. For example, the information 74 may be conveyed within SEI information of the video bitstream 110. For example, the focus region may be decoded exclusively, or the decoder 114 may be configured to start decoding each picture of the video at the focus region, rather than (for example) at the top left picture corner, and / or the decoder 114 may stop decoding each picture of the video after the focus region 116 has been decoded. Additionally or alternatively, the information 74 may be present in the data stream for forwarding only to a subsequent renderer or viewport control or streaming device of a client or segment selector for use in deciding which segments to download or stream in order to cover the spatial segment completely or at an increased or predetermined quality. For example, as outlined above, information 74 indicates a preferred area of ​​recommendation as a recommendation to place observation segment 62 or segment 66 to coincide with or cover or track this area. Information 74 can be used by a segment selector of the client. Figures 4 to 7c If the description is true, information 74 may set the size of region 116 absolutely (such as in a number of tiles), or may set the speed of region 116 at which the region is moved, for example based on user input, so as to follow the interesting content of the video in time and space, thereby scaling region 116 so that it increases as the speed indicated increases.

[0153] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, where blocks or apparatuses correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or objects or features of corresponding apparatuses. Some or all of the method steps may be performed by (or using) a hardware device (e.g., a microprocessor, a programmable computer, or an electronic circuit). In some embodiments, one or more of the most important method steps may be performed by this device.

[0154] The signals appearing above (such as the streamed signal, MPD or any other of the mentioned signals) may be stored on a digital storage medium or may be transmitted on a transmission medium (such as a wireless transmission medium or a wired transmission medium such as the Internet).

[0155] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or in software. The implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored, which cooperate (or are capable of cooperating) with a programmable computer system to cause the execution of the corresponding method. Thus, the digital storage medium may be computer readable.

[0156] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0157] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product is executed on a computer. The program code may, for example, be stored on a machine readable carrier.

[0158] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0159] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0160] Therefore, another embodiment of the inventive method is a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.

[0161] Therefore, another embodiment of the inventive method is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transmitted via a data communication connection, for example, via the Internet.

[0162] A further embodiment comprises processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0163] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0164] Another embodiment according to the invention comprises an apparatus or system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0165] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.

[0166] The devices described herein may be implemented using hardware devices or using computers or using a combination of hardware devices and computers.

[0167] The devices described herein or any component of a device described herein may be implemented at least partially in hardware and / or in software.

[0168] The methods described herein may be performed using a hardware device or using a computer or using a combination of a hardware device and a computer.

[0169] Any component of a method described herein or an apparatus described herein may be performed at least in part by hardware and / or by software.

[0170] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the appended patent claims, rather than by the specific details presented by describing and explaining the embodiments herein.

[0171] References

[0172] [A1]ISO / IEC 23009-1:2014,Information technology--Dynamic adaptive streaming over HTTP(DASH)--Part 1:Media presentation description and segmentformats.

Claims

1. An apparatus for streaming media content about a temporally changing spatial scene, configured to: Select a media clip from among multiple media clips available on the server, The device is configured as follows: encoding a first portion of the time-varying spatial scene into the selected media segment with increased quality compared to a spatial neighborhood of the first portion or in such a way that the spatial neighborhood of the first portion is not encoded into the selected media segment and that the first portion (64) of the time-varying spatial scene is pinned to the viewing segment (28), In response to the movement of the observed section, a log message is issued, the log message recording: An instantaneous measurement of the quality of the time-varying spatial scene visible in a viewing segment up to encoding into the selected media segment and during the movement of the viewing segment is measured.

2. An apparatus as claimed in claim 1, wherein the quality of the first portion or the quality of the temporally varying spatial scene until encoded into the selected media segment and until visible in an observation segment is measured as the duration for which a lower quality portion together with a higher quality portion is visible in the observation segment.

3. The apparatus of claim 1 , configured to emit a log message recording the instantaneous measurement of the quality of the time-varying spatial scene measured up to encoding into the selected media segment and up to being visible in the observation segment as one of: A measure of the average density of pixels falling within the observation segment (28) is measured, with which average density the time-varying spatial scene is encoded into the selected media segment.

4. An apparatus as claimed in claim 3, configured such that the measurement result measures the average density of pixels by averaging the pixel density in a spatially uniform manner relative to a pixel grid encoded into the image in the selected media segment.

5. An apparatus as claimed in claim 3, configured so that the measurement result measures the average density of pixels by averaging the pixel density in a spatially non-uniform manner relative to a pixel grid encoded into the image in the selected media segment.

6. The apparatus of claim 3, configured so that the issued log message indicates whether the measurement result measures the average density of pixels by: averaging the pixel density in a spatially uniform manner relative to the pixel grid of the image encoded into the selected media segment, or Pixel densities are averaged in a spatially non-uniform manner relative to a pixel grid encoded into an image in the selected media segment.

7. The apparatus of claim 5, wherein averaging the pixel density in a spatially non-uniform manner corresponds to Averaging with a spherical uniformity, or The averaging is performed spatially uniformly with respect to a viewport plane (310), which is perpendicular to a central viewing direction (312) of the viewing section (28).

8. The apparatus of claim 3, wherein the measurement result measures the average density of pixels by averaging the pixel density in a manner that limits the averaging to a central sub-segment of the observation segment (28), or applying a higher averaging weight to the central sub-segment (202) than to an edge portion of the observation segment surrounding the central sub-segment.

9. The apparatus of claim 3, configured such that the measurements measure the average density of pixels separately along a horizontal viewing segment axis and a vertical viewing segment axis (206).

10. The apparatus of claim 3, configured to emit log messages intermittently.

11. The apparatus of claim 3, configured to emit log messages at a rate controlled by a manifest file based on which the apparatus performs the selection of the media segments for downloading.

12. An apparatus as claimed in claim 3, wherein the apparatus is configured to emit a log message recording, in the form of a time measurement, the measurement of an amount of the selected media segment that has not yet been output from a buffer of the apparatus for undergoing decoding.

13. An apparatus as claimed in claim 12, wherein the apparatus is configured such that the measurement result of the amount of the selected media segment that has not yet been output from the buffer of the apparatus for undergoing decoding is presented in the form of a time measurement result measured in time units that are less than the time length of the media segment (58) and / or in a form defined independently of the time length of the media segment and / or in milliseconds.

14. The apparatus of claim 3, wherein the apparatus is configured to emit a log message recording the measurement of the amount of the selected media segment that has not yet been output from a buffer of the apparatus for undergoing decoding in a format categorized by one or more of: the decoder's buffer in which the corresponding media segment is already buffered, The scene segments encoded into the corresponding media fragments, encoding the temporally varying spatial scene to a quality used in a corresponding media segment, A spatial quality distribution is used to encode the temporally varying spatial scene into a corresponding media segment.

15. A method for streaming media content about a temporally varying spatial scene, comprising: Select a media clip from among multiple media clips available on the server, extracting the selected media segment from the server, wherein the selection is performed such that the selected media segment has a first portion of the time-varying spatial scene encoded therein with increased quality compared to a spatial neighborhood of the first portion or such that the spatial neighborhood of the first portion is not encoded into the selected media segment and such that the first portion (64) of the time-varying spatial scene is pinned to an observation segment (28), The method further comprises, in response to the movement of the observation section, issuing a log message, the log message recording: An instantaneous measurement of the quality of the time-varying spatial scene visible in a viewing segment up to encoding into the selected media segment and during the movement of the viewing segment is measured.

16. A storage medium storing a computer program having a program code for executing the method according to claim 15 when the program is executed on a computer.

Citation Information

Patent Citations

  • A video client and video server for panoramic video consumption

    EP2824883A1

  • Video Quality of Experience Based on Video Quality Estimation

    US20160088322A1

  • Reduced bit rate immersive video

    WO2016050283A1