Spatially unequal streaming

By streaming media content with spatially non-uniform quality distribution and adaptive signaling, the challenges of maintaining visual quality and reducing complexity in VR streaming are addressed, enhancing user experience and resource efficiency.

JP2025131610APending Publication Date: 2025-09-09FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025081752
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-07-08
Filing Date
2025-05-15
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing VR streaming technologies face challenges in maintaining visual quality and reducing processing complexity due to high bandwidth requirements and uneven distribution of quality in viewport changes, leading to quality degradation and inefficiencies in bandwidth and computational resources.

Method used

Streaming media content in a spatially non-uniform manner by providing hints and signaling to the client about predetermined relationships between different parts of the spatial scene, allowing for dynamic selection of media segments based on quality and computational power, and using adaptive streaming to manage viewport changes.

Benefits of technology

Improves visual quality and reduces computational complexity while optimizing bandwidth consumption by ensuring higher quality content is prioritized in the viewport, minimizing quality degradation during viewport changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131610000001_ABST
    Figure 2025131610000001_ABST
Patent Text Reader

Abstract

To provide an apparatus for providing an omni-directional video (also known as spherical video) so that mixed resolution or mixed quality video is controlled by DASH client, a streaming server, and a media presentation description.SOLUTION: A system for implementing a virtual reality application that uses dynamic adaptive streaming over HTTP (DASH) for communication 22 between a client 10 and a server 20 presents, to a user wearing a head up display 24, a view section 28 out of a temporally-varying spatial scene 30, the section corresponding to an orientation of the head up display measured by an internal orientation sensor 32 such as an inertial sensor of head up display, via an internal display 26 of the head up display. That is, the section presented to the user forms a section of a spatial scene the spatial position of which corresponds to the orientation of head up display.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to spatially non-uniform streaming, such as occurs in virtual reality (VR) streaming. [Background technology]

[0002] VR streaming typically involves transmitting very high-resolution video. The resolution of the human fovea is approximately 60 pixels per degree. Considering a full 360° x 180° spherical transmission, we would transmit a resolution of approximately 22k x 11k pixels. Because transmitting such high resolution results in very high bandwidth requirements, another solution is to transmit only the viewport shown on a head-mounted display (HMD) with a 90° x 90° FoV. In this case, the video would be approximately 6k x 6k pixels. The trade-off between transmitting the entire video at the highest resolution and transmitting only the viewport is to transmit the viewport at high resolution and some adjacent data (or the rest of the spherical video) at lower resolution or quality.

[0003] In a DASH scenario, omnidirectional video (also known as spherical video) can be provided such that the aforementioned mixed resolution or mixed quality video is controlled by the DASH client, which only needs to know information that describes how the content is being provided.

[0004] One example is being able to provide different representations using different projections with asymmetric properties such as different quality and distortion for different parts of the video. Each representation corresponds to a given viewport, where that viewport is encoded at a higher quality / resolution than other content. Knowing the directional information (the direction of the viewport where the content is encoded at a higher quality / resolution), a DASH client can dynamically select one or the other representation to match the user's gaze direction at any given time.

[0005] A more flexible option for a DASH client to select such asymmetric properties for omnidirectional video would be when the video is divided into several spatial regions, each available at a different resolution or quality. One option is to divide it into rectangular regions (a.k.a. tiles) based on a grid, but other options are possible. In such a case, the DASH client needs some signaling about the different qualities at which different regions are offered, and can download different regions at different qualities so that the viewport displayed to the user is of higher quality than other non-displayed content.

[0006] In either case, when a user action occurs and the viewport changes, the DASH client will need some time to react to the user's movement and download content in a way that matches the new viewport. Between the time the user moves and the time the DASH client adjusts its request to match the new viewport, the user will see some high-quality and some low-quality areas of the viewport simultaneously. The acceptable difference in quality / resolution depends on the content, but the user will see a degradation in quality in either case.

[0007] It is therefore preferable to have at hand a concept that provides relief, more efficient rendering, or even improved visual quality for the user regarding the partial presentation of streamed spatial scene content by adaptive streaming. Summary of the Invention [Problem to be solved by the invention]

[0008] It is therefore an object of the present invention to provide a concept for streaming spatial scene content in a spatially unequal manner to improve the visual quality for the user, or to reduce the processing complexity or required bandwidth in a streaming search site, or to provide a concept for streaming spatial scene content in a way that expands its applicability to further application scenarios.

[0009] This object is achieved by the subject matter of the pending independent claims. [Means for solving the problem]

[0010] A first aspect of the present invention is based on the discovery that streaming media content related to time-varying spatial scenes, such as videos, in a spatially unequal manner can be improved in terms of visual quality and / or computational complexity at a streaming receiving site for comparable bandwidth consumption if the selected and retrieved media segments and / or signaling obtained from the server provide the searching device with hints about predetermined relationships that different parts of the time-varying spatial scene should observe by the quality with which they are encoded in the selected and retrieved media segments. Otherwise, the searching device may not know in advance about the adverse effect that the juxtaposition of parts encoded with different qualities in the selected and retrieved media segments will have on the overall visual quality experienced by the user. Information contained in the media segments and / or signaling obtained from the server, for example in a manifest file (media presentation description), or additional streaming-related control messages from the server to the client, such as SAND messages, allows the searching device to make an appropriate selection from among the media segments provided by the server. In this way, virtual reality streaming or partial streaming of video content can be made more robust to quality degradation, as may otherwise occur due to an inappropriate distribution of available bandwidth over this spatial section of the time-varying spatial scene presented to the user.

[0011] A further aspect of the present invention is based on the discovery that streaming media content relating to a time-varying spatial scene, such as a video, in a spatially non-uniform manner, such as using a first quality for a first portion and a second, lower quality for a second portion, or leaving the second portion non-streaming, may improve visual quality and / or reduce complexity in terms of bandwidth consumption and / or computational complexity on the streaming search side by determining the size and / or position of the first portion according to information contained in the media segment and / or signaling obtained from the server. For example, assume that the time-varying spatial scene is provided by a server in a tile-based manner for tile-based streaming. That is, assume that media segments represent spectral temporal portions of the time-varying spatial scene, each of which is a time segment of the spatial scene within a corresponding tile of a distribution of tiles into which the spatial scene is subdivided. In such a case, it is up to the search device (client) to decide how to distribute available bandwidth and / or computational power across the spatial scene, i.e., at the granularity of a tile. The search device performs the selection of media segments to the extent that a first portion of a subsequent spatial scene, each tracking a time-varying view section of the spatial scene, is encoded into the selected retrieved media segment at a predetermined quality, which may be, for example, the highest quality achievable under current bandwidth and / or computational power conditions. A spatially adjacent second portion of the spatial scene may, for example, not be encoded into the selected retrieved media segment, or may be encoded therein at a further quality that is reduced relative to the predetermined quality. In such a situation, calculating the number / sum of adjacent tiles whose set completely covers the time-varying view section, regardless of the orientation of the view section, is a computationally complex problem or is infeasible.Depending on the projection selected to map the spatial scene to individual tiles, angular scene coverage per tile may vary due to the fact that the scene and individual tiles may overlap each other, making it even more difficult to calculate the number of adjacent tiles sufficient to spatially cover the view section, regardless of the view section's orientation. Therefore, in such situations, the aforementioned information may indicate the size of the first portion as the number of tiles N or the number of tiles, respectively. By this means, the device can track the time-varying view section by selecting those media segments that have a co-located collection of N tiles encoded therein with a predetermined quality. The fact that a collection of these N tiles sufficiently covers the view section can be guaranteed by the information indicating N. Another example is information included in the media segment and / or signaling obtained from the server that indicates the size of the first portion relative to the size of the view section itself. For example, this information may somehow establish a "safety zone" or prefetch zone around the actual view section to account for the movement of the time-varying view section. The faster the time-varying view section moves across the spatial scene, the larger the safety zone. The information may thus indicate the size of the first portion in a manner relative to the size of the view section that changes over time, such as incrementally or scaling. A search device that sets the size of the first portion according to such information may avoid quality degradation that would otherwise occur due to unsearched or low-quality parts of the spatial scene appearing in the view section. Here, it is irrelevant whether the scene is provided in a tile-based manner or in some other manner.

[0012] In relation to the immediately preceding aspect of the present application, a video bitstream having video encoded therein may be decodable with higher quality if the video bitstream is provided with signaling of the size of a focus region within the video where decoding power for decoding the video should be concentrated. By this means, a decoder decoding video from the bitstream can concentrate or even limit its decoding power to decoding a portion of the video having the size of the focus region signaled in the video bitstream, knowing, for example, that the portion thus decoded is decodable by the available decoding power and spatially covers a required portion of the video. For example, the size of the focus region signaled in this way may be selected to be large enough to cover the size of a view section and the motion of this view section, taking into account the decoding latency when decoding the video. Or, in other words, signaling a recommended preferred view section region of the video included in the video bitstream allows the decoder to treat this region in a preferred manner, thereby concentrating its decoding power accordingly. Independently of performing area-specific concentration of decoding capabilities, area signaling can be transferred to the stage of selecting on which media segments to download, i.e., where to place them, and how to size the quality-enhanced portions.

[0013] The first and second aspects of the present application are closely related to the third aspect of the present application and take advantage of the fact that a huge number of search devices stream media content from a server, and subsequently obtain information that can be used to appropriately set the size, or size and / or location, of the first portion and / or the type of information described above that enables appropriately setting a predetermined relationship between the first and second qualities. Thus, according to this aspect of the present application, the search device (client) sends a log message collecting one of momentary measurements or statistics measuring the spatial location and / or movement of the first portion, momentary measurements or statistics measuring the quality of the time-varying spatial scene as encoded in the selected media segment and visible in the view section, and momentary measurements or statistics measuring the quality of the first portion or the quality of the time-varying spatial scene as encoded in the selected media segment and visible in the view section. The momentary measurements and / or statistics can be provided with time information regarding the time at which each momentary measurement or statistic was obtained. The log message may be sent to a server where the media segment is provided or to another device which evaluates the received log message and may based thereon update the size of the first portion or the current settings of the information used to set the size and / or position thereof and / or derive a predetermined relationship based thereon.

[0014] According to a further aspect of the present application, streaming media content related to time-varying spatial scenes such as videos, particularly streaming media content in a tile-based manner, can be made more effective in terms of avoiding useless streaming attempts by providing a media presentation description that includes at least one version in which the time-varying spatial scene is provided for tile-based streaming and, for each of the at least one version, an indication of the benefit requirements for each version of the time-varying spatial scene to benefit from tile-based streaming. By this means, a search device can match the benefit requirements of the at least one version with the device capabilities of the search device itself or other devices that interact with the search device regarding tile-based streaming. For example, the benefit requirement can be related to a decoding capability requirement. That is, if the decoding capability for decoding the streamed / searched media content is not sufficient to decode all media segments necessary to cover a view section of the time-varying spatial scene, attempting to stream and present the media content would be wasteful in terms of time, bandwidth, and computational power, and accordingly, it would be more effective not to attempt it anyway. For example, if media segments for one tile form a media stream, such as a video stream, separately from media segments for other tiles, the decoding capability requirement may indicate, for example, several decoder instantiations required for each version. The decoding capability requirement may also relate to further information, such as, for example, the specific parts of decoder instantiations required to conform to a given decoding profile and / or level, or may indicate a specific minimum capability of a user input device to move a viewport / section fast enough for a user to view a scene through it. Depending on the content of the scene, low movement capability may not be sufficient for a user to view interesting parts of the scene.

[0015] A further aspect of the present invention relates to an extension of streaming of media content with respect to time-varying spatial scenes. In particular, the idea according to this aspect is that the spatial scene may not only actually vary in time, but may also vary with respect to at least one further parameter that indicates, for example, view and position, view depth, or some other physical parameter. The search device can use adaptive streaming in this context by calculating addresses of media segments depending on the viewport direction and the at least one further parameter, the media segments describing the time-varying spatial scene and the at least one parameter, and retrieving the media segments using the calculated addresses from the server.

[0016] The above-outlined aspects of the present application and their advantageous embodiments, which are the subject of the dependent claims, can be combined individually or all together.

[0017] Preferred embodiments of the present application are described below with reference to the drawings. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 shows a schematic diagram illustrating a client and server system for a virtual reality application, as an example of where the embodiments shown in the following figures can be used to advantage. [Figure 2] FIG. 2 shows a block diagram of a client device together with a schematic diagram of a media segment selection process to illustrate a possible mode of operation of the client device according to one embodiment of the present application in which the server 10 provides the device with information about acceptable or tolerable quality variations in the media content presented to the user. [Figure 3] Figure 3 shows a variation of Figure 2, where the quality improvement is not related to tracking view sections of the viewport, but rather to regions of interest in the media scene content that are signaled from the server to the client. [Figure 4] FIG. 4 shows a block diagram of a client device together with a schematic diagram of a media segment selection process according to an embodiment in which the server provides information on how to set the size or size and / or location of the quality-enhanced portion or the size or size and / or location of the actually retrieved section of the media scene. [Figure 5] FIG. 5 illustrates a variation of FIG. 5 in that the information sent by the server indicates the size of the portion 64 directly, rather than scaling according to the expected movement of the viewport. [Figure 6] FIG. 6 shows a variation of FIG. 4 in which the retrieved section has a predetermined quality and its size is determined by information originating from the server. [Figure 7a] FIG. 7a is a schematic diagram showing how the information 74 according to FIGS. 4 and 6 is transformed to increase the size of the retrieved portion at a given quality with a corresponding increase in the size of the viewport. [Figure 7b] FIG. 7b is a schematic diagram showing how the information 74 according to FIGS. 4 and 6 is transformed to increase the size of the retrieved portion at a given quality with a corresponding increase in the size of the viewport. [Figure 7c] FIG. 7c is a schematic diagram showing how the information 74 according to FIGS. 4 and 6 is transformed to increase the size of the retrieved portion at a given quality with a corresponding increase in the size of the viewport. [Figure 8a] FIG. 8a is a schematic diagram illustrating an embodiment in which a client device sends log messages to a server or a specific evaluator to evaluate these log messages and derive appropriate settings for the types of information, e.g., as described with respect to FIGS. 2 to 7c. [Figure 8b]Figure 8b is a schematic diagram showing an example of a tile-based stereoscopic projection of a 360° scene onto tiles, and how some tiles are covered by exemplary positions in the viewport. The small circles indicate the positions of the equiangular distribution in the viewport, and the hatched tiles are encoded at a higher resolution than the non-hatched tiles in the downloaded segment. [Figure 8c] Fig. 8c shows a schematic diagram of a diagram showing along the time axis (horizontal) how the buffer occupancy (vertical axis) of different buffers in the client may occur, where Fig. 8c assumes a buffer used to buffer a representation encoding a particular tile. [Figure 8d] Fig. 8d shows a schematic diagram of a diagram showing along the time axis (horizontal) how the buffer occupancy (vertical axis) of different buffers of a client may occur, where Fig. 8d assumes buffers used to buffer an omnidirectional representation with an encoded scene of non-uniform quality, i.e., increased towards a certain direction specific to each buffer. [Figure 8e] FIG. 8e shows a three-dimensional view of the measurement of different pixel densities in different viewports 28 with respect to uniformity in a spherical or view plane sense. [Figure 8f] FIG. 8f shows a three-dimensional view of the magnitude of different pixel densities in different viewports 28 with respect to uniformity in a spherical or view plane sense. [Figure 9] FIG. 9 shows a block diagram of a client device and a schematic diagram of the media segment selection process when the device examines information originating from the server to evaluate whether a particular version in which the server provides tile-based streaming is acceptable to the client device. [Figure 10]FIG. 10 shows a schematic diagram illustrating multiple media segments provided by a server according to an embodiment that allows not only temporal dependency of a media scene but also dependency of another non-temporal parameter, namely here exemplarily scene center position. [Figure 11] FIG. 11 shows a schematic diagram illustrating a video bitstream containing information that creates or controls the size of focus regions within video encoded in the bitstream, along with an example of a video decoder that can utilize this information. DETAILED DESCRIPTION OF THE INVENTION

[0019] To facilitate understanding of the description of embodiments of the present application with respect to various aspects thereof, Figure 1 illustrates an example of an environment in which the below-described embodiments of the present application may be applied and advantageously used. In particular, Figure 1 illustrates a system consisting of a client 10 and a server 20 interacting via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) may be used for communication 22 between the client 10 and the server 20. However, the embodiments outlined below should not be construed as being limited to the use of DASH, and similarly, terms such as Media Presentation Description (MPD) should be understood to be broad enough to cover manifest files that are defined differently from DASH.

[0020] FIG. 1 illustrates a system configured to implement a virtual reality application. That is, the system is configured to present to a user wearing a head-up display 24, i.e., via an internal display 26 of the head-up display 24, view sections 28 from a time-varying spatial scene 30, where the sections 28 correspond to the orientation of the head-up display 24, illustratively measured by an internal orientation sensor 32, such as an inertial sensor, of the head-up display 24. That is, the sections 28 presented to the user form sections of the spatial scene 30 whose spatial positions correspond to the orientation of the head-up display 24. In the case of FIG. 1, the time-varying spatial scene 30 is depicted as an omnidirectional or spherical video, but the description of FIG. 1 and the embodiments described thereafter are equally easily transferable to other examples, such as presenting sections from a video whose spatial positions are determined by the intersection of face access or eye access with a virtual or real projector wall, etc. Furthermore, the sensor 32 and the display 26 may be constituted by different devices, such as, for example, a remote control and a corresponding television, respectively, or may be part of a handheld device, such as a mobile device, such as a tablet or mobile phone. Finally, it should be noted that some of the embodiments described below may also be applied to scenarios involving unevenness in the presentation of the time-varying spatial scene 30, such as uneven distribution of quality across the spatial scene, where the region 28 presented to the user always covers the entire time-varying spatial scene 30.

[0021] Further details regarding server 20, client 10, and the manner in which spatial content 30 is provided on server 20 are shown in Figure 1 and described below. However, these details should also not be treated as limiting the embodiments described later, but rather should serve as examples of how to implement any of the embodiments described later.

[0022] In particular, as shown in Figure 1, the server 20 may include a storage device 34 and a controller 36, such as a suitably programmed computer, application specific integrated circuit, or the like. The storage device 34 stores media segments representing a time-varying spatial scene 30. A particular example is outlined in more detail below with respect to the diagram of Figure 1. The controller 36 responds to requests sent by the client 10 by resending the requested media segments, media presentation descriptions, to the client 10, and may itself send further information to the client 10. More details on this are also provided below. The controller 36 may retrieve the requested media segments from the storage device 34. Other information, such as a media presentation description or portions thereof, may also be stored in this storage device in other signals sent from the server 20 to the client 10.

[0023] 1, the server 20 may optionally further comprise a stream modifier 38 which modifies the media segments sent by the server 20 to the client 10 in response to a request from the latter so as to form a media data stream at the client 10, e.g., the media segments thus retrieved by the client 10 are in fact aggregated from several media streams, but one single media stream decodable by one associated decoder. However, the presence of such a stream modifier 38 is optional.

[0024] The client 10 of FIG. 1 is illustratively depicted as including a client device or controller 40, or more preferably a decoder 42 and a reprojector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA, or an application-specific integrated circuit. The client device 40 undertakes to select a media segment to be retrieved from the server 20 from a plurality of 46 media segments provided by the server 20. For this purpose, the client device 40 first retrieves a manifest or media presentation description from the server 20. Similarly, the client device 40 obtains computation rules for calculating the address of a media segment from the plurality 46 that corresponds to a required spatial portion of the spatial scene 30. The media segments thus selected are retrieved from the server 20 by the client device 40 by sending respective requests to the server 20. These requests include the computed addresses.

[0025] The media segments retrieved by the client device 40 in this manner are forwarded by the latter to one or more decoders 42 for decoding. In the example of FIG. 1 , the retrieved and decoded media segments thus retrieved represent merely a spatial section 48 of the time-varying spatial scene 30 for each temporal unit of time, but as already mentioned above, this may differ according to other aspects, e.g., the presented view section 28 always covers the entire scene. The reprojector 44 may optionally reproject and crop the view section 28 as displayed to the user from the retrieved and decoded scene content of the selected retrieved and decoded media segment. To this end, as shown in FIG. 1 , the client device 40 continuously tracks and updates the spatial position of the view section 28, e.g., in response to user orientation data from the sensor 32, and informs the reprojector 44 of this current spatial position of the scene section 28 as well as the reprojection mapping to be applied to the retrieved and decoded media content as it is mapped to the region-forming view section 28. Reprojector 44 may apply mapping and interpolation accordingly, for example, to the regular grid of pixels displayed on display 26 .

[0026] FIG. 1 illustrates the use of cubic mapping to map the spatial scene 30 onto tiles 50. Thus, tiles are depicted as rectangular subregions of a cube onto which the spherical scene 30 is projected. The reprojector 44 inverts this projection. However, other examples may be used as well. For example, instead of a stereographic projection, a projection onto a truncated or untruncated pyramid may be used. Furthermore, although the tiles in FIG. 1 are depicted as non-overlapping in terms of coverage of the spatial scene 30, the subdivision into tiles may include mutual tile overlap. And, as outlined in more detail below, it is also not required that the scene 30 be spatially subdivided into tiles 50, each forming a single representation, as further described below.

[0027] Thus, as shown in FIG. 1, the entire spatial scene 30 is spatially subdivided into tiles 50. In the example of FIG. 1, each of the six faces of a cube is subdivided into four tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52 as shown in FIG. 1. More precisely, the server 20 may even provide multiple videos 52 per tile 50, with these videos having different qualities Q#. Furthermore, the videos 52 are temporally subdivided into temporal segments 54. The temporal segments 54 of all the videos 52 of all the tiles T# form or are each encoded into one of a plurality of 46 media segments stored in the storage device 34 of the server 20.

[0028] It is emphasized again that even the example of tile-based streaming shown in Figure 1 merely forms an example from which many variations are possible. For example, Figure 1 seems to suggest that media segments relating to a representation of scene 30 at a higher quality are associated with tiles that match the tiles to which the media segments in which scene 30 was encoded at quality Q1 belong. This correspondence is not necessary; tiles of different qualities may even correspond to tiles of different projections of scene 30. Furthermore, although not discussed so far, media segments corresponding to the different quality levels shown in Figure 1 may differ in spatial resolution and / or signal-to-noise ratio and / or temporal resolution, etc.

[0029] Finally, unlike the tile-based streaming concept in which media segments that can be individually retrieved from server 20 by device 40 relate to tiles 50 into which scene 30 is spatially subdivided, the media segments provided at server 20 may alternatively, for example, each encode scene 30 in a spatially complete manner with spatially varying sampling resolutions, but with a maximum sampling resolution at a different spatial location within scene 30. For example, this may be achieved by providing server 20 with a series of segments 54 that relate to the projection of scene 30 onto a truncated pyramid whose truncated tips would be oriented in different directions from one another, thereby resulting in resolution peaks of different orientations.

[0030] Furthermore, it should be noted that with respect to presenting the stream modifier 38 as needed, the same may be part of the client 10 or may be located between the client 10 and the server 20 in the network equipment through which they exchange the signals described herein.

[0031] After describing the server 20 and the client 10 systems in a fairly general manner, the functionality of the client device 40 for an embodiment according to a first aspect of the present application will be described in more detail. For this purpose, reference is made to FIG. 2, which shows the device 40 in more detail. As already mentioned above, the device 40 is for streaming media content relating to a time-varying spatial scene 30. As explained with reference to FIG. 1, the device 40 may be configured so that the streamed media content is spatially related to the entire scene continuously, or may be limited only to a section 28 thereof. In any case, the device 40 includes a selection unit 56 for selecting an appropriate media segment 58 from a plurality 46 of media segments available on the server 20, and a retrieval unit 60 for retrieving the selected media segment from the server 20 by a respective request, such as an HTTP request. As mentioned above, the selection unit 56 can use the media presentation description to calculate addresses of the media segments selected by the retrieval unit 60, using these addresses when retrieving the selected media segment 58. For example, the calculation rule for calculating the address indicated in the media presentation description may depend on the quality parameter Q, the tile T, and some time segment t. The address may be, for example, a URL.

[0032] Also as mentioned above, selector 56 is configured to perform the selection such that the selected media segment comprises, and is encoded within, at least one spatial section of a time-varying spatial scene. The spatial section may spatially continuously cover the entire scene. Figure 2 shows at 61 an exemplary case in which device 40 adapts spatial section 62 of scene 30 to overlap and surround view section 28. However, as already mentioned above, this is not necessarily the case; the spatial section may continuously cover the entire scene 30.

[0033] Furthermore, the selector 56 performs the selection so that the selected media segment has sections 62 encoded therein with spatially unequal quality. More precisely, a first portion 64 of the spatial section 62, indicated by hatching in FIG. 2, is encoded into the selected media segment with a predetermined quality. This quality may be, for example, the highest quality offered by the server 20 or a "good" quality. The device 42 moves or adapts the first portion 64, for example, to spatially follow the time-varying view section 28. For example, the selector 56 selects the current temporal segment 54 of those tiles that inherit the current position in the view section 28. In doing so, the selector 56 can optionally keep the number of tiles that make up the first portion 64 constant, as will be described with respect to further embodiments below. In any case, a second portion 66 of the section 62 is encoded into the selected media segment 58 with another quality, such as a lower quality. For example, the selector 56 selects a media segment that corresponds to a temporal segment of the current tile that is spatially adjacent to the tiles of the portion 64 and belongs to a lower quality. For example, selector 56 could primarily select the media segment corresponding to portion 66 to address the chance that view section 28 may move too quickly to leave portion 64 and overlapping portion 66 before the time interval corresponding to the current time segment ends, and selector 56 could re-spatially position portion 64. In this situation, the portion of section 28 that protrudes into portion 66 may nevertheless be presented to the user, i.e., with reduced quality.

[0034] It is not possible for device 40 to assess what degradation in quality may occur by pre-presenting the user with reduced-quality scene content along with the scene content in high-quality portion 64. In particular, a transition between these two qualities may occur that is clearly visible to the user. At the very least, such a transition may be visible depending on the current scene content in section 28. The severity of the adverse effect of such a transition in the user's field of view is a characteristic of the scene content as provided by server 20 and may not be predicted by device 40.

[0035] Thus, according to the embodiment of FIG. 2 , apparatus 40 includes a derivation unit 66 that derives a predetermined relationship to be satisfied between the quality of portion 64 and the quality of portion 66. Derivation unit 66 derives this predetermined relationship from information that may be included within the media segment, such as in a transport box within media segment 58, and / or that may be included in signaling obtained from server 20, such as in a media presentation description, or in a proprietary signal transmitted from server 20, such as a SAND message. An example of what information 68 may look like is provided below. The predetermined relationship 70 derived by derivation unit 66 based on information 68 is used by selector 56 to appropriately perform the selection. For example, a restriction on the selection of the qualities of portions 64 and 66 compared to a completely independent selection of the qualities of portions 64 and 66 affects the distribution of bandwidth available for retrieving media content related to portion 62 onto portions 64 and 66. In any case, selector 56 selects media segments such that the qualities with which portions 64 and 66 are ultimately encoded into the retrieved media segment satisfy the predetermined relationship. An example of what the predetermined relationship may look like is also provided below.

[0036] The media segments selected and ultimately retrieved by the search unit 60 are ultimately forwarded to one or more decoders 42 for decoding.

[0037] According to a first example, the signaling mechanism embodied by information 68, for example, includes information 68 indicating to device 40, which may be a DASH client, which quality combinations are acceptable for the video content provided. For example, information 68 may be a list of quality pairs indicating to a user or device 40 that different regions 64 and 66 may be mixed with a maximum quality (or resolution) difference. Device 40 may be configured to necessarily use a particular quality level for portion 64, such as the highest quality level provided by server 10, and to derive the quality level that portion 66 can encode into the selected media segment from information 68, which portion 66 includes, for example, in the form of a list of quality levels for portion 68.

[0038] Information 68 may indicate a tolerance for the measure (magnitude) of the difference between the quality of part 68 and the quality of part 64. As a "measure" of the quality difference the quality index of media segment 58 may be used, by which the same are distinguished in the media presentation description and by which its address is calculated using the calculation rules described in the media presentation description. In MPEG-DASH a corresponding attribute indicating the quality would be for example @qualityRanking. When performing the selection, device 40 may take into account restrictions on the selectable quality level pairs at which parts 64 and 66 can be encoded into the selected media segment.

[0039] However, instead of this difference measure, assuming that bitrate typically increases monotonically with increasing quality, the quality difference can also be measured, for example, by the bitrate difference, i.e., the tolerable difference in the bitrates at which portions 64 and 66 are respectively encoded into corresponding media segments. Information 68 can indicate pairwise options that are permissible for the quality at which portions 64 and 66 are encoded into selected media segments. Alternatively, information 68 simply indicates the tolerable quality for encoding portion 66, thereby indirectly indicating the tolerable quality or tolerable quality difference, assuming that main portion 64 is encoded using some default quality, such as, for example, the highest possible or available quality. For example, information 68 can be a list of permissible representation IDs, or indicate a minimum bitrate level for the media segment related to portion 66.

[0040] However, instead, a more gradual quality difference may be desired, in which case, instead of quality pairs, quality groups (two or more qualities) may be indicated, where the quality difference increases depending on the distance to section 28, i.e., the viewport. That is, information 68 may indicate acceptable values ​​for the measure of the difference between the quality of parts 64 and 66 in a manner that depends on the distance to view section 28. This may be done by a list of pairs, each pair consisting of a respective distance to the view section and a corresponding tolerance value for measuring the quality difference beyond the respective distance. Below the respective distance, the quality difference must be smaller. That is, each pair indicates, for the corresponding distance, that parts in part 66 farther away than the corresponding distance in part 28 may have a quality difference to the quality of part 64 that exceeds the corresponding tolerance value of this list item.

[0041] The tolerable value may increase as the distance to the view section 28 increases. The acceptance of the just-described quality differences often depends on the time for which these different qualities are displayed to the user. For example, content with a large quality difference may be tolerable if it is displayed for only 200 microseconds, while content with a small quality difference may be tolerable if it is displayed for 500 microseconds. Thus, by way of further example, information 68 may, for example, include, in addition to the above-mentioned quality combinations or tolerable quality differences, also the time intervals within which the combinations / quality differences are tolerable. In other words, information 68 may indicate the tolerable or maximum tolerable difference between the qualities of portions 66 and 64, along with an indication of the maximum tolerable time interval within which portion 66 may be displayed in view section 28 simultaneously with portion 64.

[0042] As already mentioned, the acceptance of quality differences depends on the content itself. For example, the spatial location of different tiles 50 influences acceptance. Quality differences in uniform background areas containing low-frequency signals are expected to be more tolerable than quality differences in foreground objects. Furthermore, temporal location also influences the acceptance rate of content changes. Thus, according to another example, the signals forming information 68 are transmitted to device 40 intermittently, for example, per DASH representation or period. That is, the predetermined relationship indicated by information 68 may be updated intermittently. Additionally and / or alternatively, the signaling mechanism realized by information 68 may be spatially variable. That is, information 68 may be made spatially dependent, for example, by an SRD parameter in DASH. That is, information 68 regarding different spatial regions of scene 30 may indicate different predetermined relationships.

[0043] The embodiment of device 40, as described with respect to FIG. 2, concerns the fact that device 40 desires to minimize quality degradation due to prefetched portions 66 within retrieved portion 62 of video content 30, and to be able to temporarily view portions 62 and 64 in section 28 before being able to change their positions to accommodate changes in their positions in section 28. That is, in FIG. 2, portions 64 and 66, whose quality is limited by information 68 as far as their possible combinations are concerned, are different parts of portion 62, and the transition between both portions 64 and 66 is continuously shifted or adapted to track or navigate moving view section 28. According to an alternative embodiment shown in FIG. 3, device 40 uses information 68 to control the possible combinations of the quality of portions 64 and 66, which, according to the embodiment of FIG. 3, are defined as portions distinct or distinguishable from each other, for example, in a manner defined in a media presentation description, i.e., independent of their positions in view section 28. The positions of portions 64 and 66 and the transitions between them may be constant or may vary over time. If they vary over time, the variation is due to changes in the content of scene 30. For example, portion 64 corresponds to an area of ​​interest where higher quality expenditure is worthwhile, while portion 66 is an area where quality degradation due to, for example, low bandwidth conditions should be considered before considering quality degradation for portion 64.

[0044] In the following, further embodiments for advantageous implementation of the device 40 are described. In particular, Figure 4 shows the device 40 in a manner that structurally corresponds to Figures 2-3, where the operating mode has been modified to correspond to the second aspect of the present invention.

[0045] 4, the apparatus 40 comprises a selector 56, a searcher 60, and a deriver 66. The selector 56 selects from a plurality 46 of media segments 58 provided by the server 20, and the searcher 60 retrieves the selected media segment from the server. Figure 4 assumes that the apparatus 40 operates as illustrated with respect to Figures 2-3. That is, the selector 56 performs the selection such that the selected media segment 58 has a spatial portion 62 of the scene 30 encoded such that this spatial portion follows a view portion 28 that changes its spatial position over time. However, a corresponding variation of the same aspect of the present application is described later with respect to Figure 5, in which for each time instant t, the selected and retrieved media segment 58 has the entire scene or a certain spatial section 62 encoded therein.

[0046] 2 and 3, the selector 56 selects a media segment 58, and a first portion 64 within the section 62 is encoded into the selected and retrieved media segment 58 at a predetermined quality, while a second portion 66 of the section 62, spatially adjacent to the first portion 64, is encoded into the selected media segment at a lower quality compared to the predetermined quality of the portion 64. A variant is shown in FIG. 6 in which the selector 56 restricts the selection and retrieval to media segments relating to a moving template that tracks the position of the viewport 28, such that the first portion 64 completely covers the portion 62 while being surrounded by an unencoded portion 72, and the media segment is encoded completely into the section 62 at the predetermined quality. In any case, the selector 56 performs the selection such that the first portion 64 follows the view section 28, whose spatial position changes over time.

[0047] 4 to 6, information 74 is provided by server 20 to device 40 to assist device 40 in setting the size, or size and / or position, of first portion 64 or section 62, respectively, in dependence on information 74. With regard to the possibility of transmitting information 74 from server 20 to device 40, the same applies as described above in relation to FIGS. 2 and 3. That is, the information may be included in media segment 58, such as in its event box, or transmission in a media presentation description, or a proprietary message transmitted from server 20 to device 40, such as a SAND message, may be used for this purpose.

[0048] 4-6, the selector 56 is configured to set the size of the first portion 64 depending on information 74 originating from the server 20. In the embodiment shown in FIGS. 4-6, the size is set in units of tiles 50, although the situation may be slightly different when using other concepts for providing the scene 30 at the server 20 with spatially varying quality, as already mentioned above with respect to FIG.

[0049] According to one example, the information 70 may include, for example, a probability for a given movement speed of the viewport of the view section 28. The information 74 may occur in a media presentation description made available to the client device 40, which may be, for example, a DASH client, as already mentioned above, or some in-band mechanism may be used to convey the information 74, such as an event box, i.e., an EMSG or SAND message in the case of DASH. The information 74 may also be included in any container format, such as the ISO file format or a transport format beyond MPEG-DASH, such as MPEG-2TS. It may also be conveyed within the video bitstream, such as in an SEI message, as described below. In other words, the information 74 may indicate a predetermined value for a measure of the spatial speed of the view section 28. In this way, the information 74 indicates the size of the portion 64 in the form of a scaling or increment relative to the size of the view section 28. That is, the information 74 starts with a certain "base size" for the portion 64 necessary to cover the size of the section 28 and increases this "base size" appropriately, incrementally, by scaling, etc. For example, the aforementioned speed of movement of the view section 28 may be used to correspondingly scale the perimeter of the current view section 28 position after this time interval, e.g., to determine the furthest position of the perimeter of the view section 28 along any possible spatial direction, and to determine the latency in adjusting the spatial position of the portion 64, e.g., the duration of the temporal segment 54 corresponding to the temporal length of the media segment 58. The speed time applied to the perimeter of the current position of the omnidirectional viewport 28 may be such a worst-case perimeter and may be used to determine the magnification of the portion 64 relative to some minimum magnification of the portion 64 assuming a non-moving viewport 28.

[0050] The information 74 may also be relevant to the evaluation of user behavior statistics. Later, suitable embodiments for performing such an evaluation process will be described. For example, the information 74 may indicate the maximum speed for a certain percentage of users. For example, the information 74 may indicate that 90% of users move at a speed slower than 0.2 rad / s and 98% of users move at a speed lower than 0.5 rad / s. The information 74 or a message conveying it may be defined so that probability-speed pairs are defined, or a message informing the maximum speed is defined for a certain percentage of users, e.g., 99% of users at all times. The movement speed signaling 74 may further include direction information, i.e., angle in 2D or 2D plus depth and light field applications in 3D. The information 74 may indicate different probability-speed pairs for different movement directions.

[0051] In other words, information 74 can apply to a given period, such as the time length of a media segment. It can include trajectory-based (X percentile, average user path), speed-based pairs (X percentile, speed), distance-based pairs (X percentile, caliber / diameter / recommended), area-based pairs (X percentile, recommended preferred area), or maximum boundary values ​​for one of the paths, speeds, distances, or preferred areas. Instead of relating information to percentiles, a simple frequency ranking can be used, such as most users moving at a certain speed, followed by many users moving at an even faster speed. Additionally or alternatively, information 74 is not limited to indicating the speed of view section 28, but can also indicate a preferred area to look at to indicate portions 62 and / or 64, respectively, of view section 28 that are sought to be tracked, with or without an indication of the percentage of users following the display or whether the display matches the user's viewing speed / view section most frequently recorded, and with or without an indication of the statistical significance of the display, with or without temporal persistence. Information 74 may indicate another measure of the speed of view section 28, such as a measure of the distance traveled by view section 28 within a period of time, such as the duration of a media segment, or more particularly, the duration of time segment 54. Alternatively, information 74 may be shown to distinguish particular movement directions in which view section 28 may move. This relates to both an indication of the speed or velocity of view section 28 in a particular direction, as well as an indication of the distance traveled by view section 28 relative to a particular movement direction. Furthermore, an extension of portion 64 may be indicated by information 74 either omnidirectionally or directly in a manner that distinguishes between different movement directions. Furthermore, all of the examples just outlined may be modified in that information 74 indicates these values ​​along with the percentage of users for whom these values ​​are sufficient to describe statistical behavior in moving view section 28. In this regard, it should be noted that view speed, i.e., view speed section 28, may be any value and is not limited to, for example, a user's head velocity value.Rather, view section 28 may be moved in response to, for example, the user's eye movements, in which case the view velocity may be significant. View section 28 may also be moved in response to the movement of another input device, such as by the movement of a tablet or the like. Because all of these "input possibilities" that allow a user to move section 28 result in different expected velocities of view section 28, information 74 may be designed to distinguish between different concepts for controlling the movement of view section 28. That is, information 74 may indicate or be displayed to indicate the size of portion 64 to indicate different sizes for different ways of controlling the movement of view section 28, and device 40 would use the size indicated by information 74 for correct view section control. That is, view section 28 is controlled by the user, i.e., device 40 determines whether view section 28 is being controlled by head movement, eye movement, tablet movement, etc., and knows how to set the size according to the portion of information 74 that corresponds to such view portion control.

[0052] In general, movement speed can be signaled per content, per period, per representation, per segment, per SRD position, per pixel, per tile, e.g., at any temporal or spatial granularity, etc. As outlined, movement speed can also be differentiated by head movement and / or eye movement. Furthermore, information 74 about the user's movement probability may be conveyed as recommendations for high-resolution prefetching, i.e., video regions outside the user's viewport, or spherical extents.

[0053] 7a-7c briefly summarize some of the options described for information 74 in methods used by device 40 to modify the size and / or position of portions 64 or 62, respectively. According to the option shown in FIG. 7a, device 40 expands the circumference of section 28 by a distance corresponding to the product of signal rate v and a time length Δt, which may correspond to a duration corresponding to the time length of time segments 54 encoded in individual media segments 50a. Additionally and / or alternatively, the positions of portions 62 and / or 64 may be positioned further away from the current position of section 28 or from the current positions of portions 62 and / or 64 in the direction of signal rate or movement as indicated by information 74, the greater the velocity. The velocity and direction may be derived from studying or estimating recent developments or changes in the indication of recommended preferred regions by information 74. Instead of applying v×Δt in all directions, the velocity may be indicated by different information 74 for different spatial directions. An alternative shown in FIG. 7d directly indicates that information 74 can indicate the distance by which the circumference of view section 28 is enlarged, this distance being indicated by parameter s in FIG. 7b. Also, a cross-sectional enlargement that varies with direction may be applied. FIG. 7c shows that the enlargement of the circumference of section 28 can be indicated by information 74 in terms of an area increase, for example, in the form of the ratio of the area of ​​the enlarged section compared to the original area of ​​section 28. In any case, the circumference of area 28 after enlargement, indicated by 76 in FIGS. 7a-7c, can be used by selector 56 to dimension or set the dimensions of portion 64 so that portion 64 covers the entire area within enlarged portion 76 by at least a predetermined amount thereof. Obviously, the larger section 76, the greater the number of tiles, for example, within portion 64. According to a further alternative, section 74 can directly indicate the size of portion 64, for example, in the form of the number of tiles that make up portion 64.

[0054] The latter possibility of signaling the size of portion 64 is shown in Figure 5. The embodiment of Figure 5 can be modified in the same way as the embodiment of Figure 4 was modified by the embodiment of Figure 6, i.e. the entire area of ​​section 62 can be retrieved from server 20 by segment 58 in the quality of portion 64.

[0055] In any case, at the end of FIG. 5, information 74 distinguishes between view sections 28 of different sizes, i.e., between different fields of view seen by view section 28. Information 74 simply indicates the size of portion 64 according to the size of the view section 28 that the device 40 is currently aiming at. Thereby, without the need for a device such as device 40 to calculate or otherwise infer the size of portion 64, as described with respect to FIGS. 4, 6, and 7, regardless of any movement of section 28, the service of server 20 can be used by devices having different fields of view or view sections 28 of different sizes, as long as portion 64 covers view section 28. As is clear from the description of FIG. 1, for example, it is almost easy to evaluate how many certain number of tiles are sufficient to completely cover a field of view of a certain size, i.e., regardless of the direction of view section 28 for spatial positioning 30. Here, information 74 alleviates this situation, and device 40 can simply examine the value of the size of portion 64 to be used for the size of view section 28 applied to device 40 within information 74. That is, according to the embodiment of FIG. 5, media presentation descriptions made available to a DASH client, or some curved mechanisms such as event boxes or SAND messages, can include information 74 regarding spherical ranges or sets of representations or sets of tiles for each field of view. An example can be a tiled offering having M representations as shown in FIG. 1. Information 74 can indicate, for example, from the cubic representations tiled in a 6×4 tile as shown in FIG. 1, the recommended number n<M (referred to as representations) of tiles to download to cover the field of view of a given end device, and it is considered sufficient to cover a 90°×90° field of view with 12 tiles. Since the field of view of the end device does not necessarily exactly align with the tile boundaries, this recommendation cannot be trivially generated by device 40 alone. Device 40 can use information 74, for example, by downloading at least N tiles, i.e., media segments 58 regarding N tiles.Another way to utilize the information is to emphasize the quality of the N tiles in section 62 that are closest to the end device's current viewing center, i.e., use those N tiles to construct portion 64 of section 62.

[0056] With reference to FIG. 8a, an embodiment relating to a further aspect of the present invention will be described. Here, FIG. 8a shows a client device 10 and a server 20 communicating with each other according to any of the possibilities described above with reference to FIGS. 1 to 7. That is, the device 10 can be embodied according to any of the embodiments described with reference to FIGS. 2-7. Alternatively, it can simply operate without these details, as described above with reference to FIG. 1. Preferably, however, the device 10 is embodied according to any of the embodiments described above with reference to FIGS. 2-7 or a combination thereof, and further inherits the mode of operation just described with reference to FIG. 8a. In particular, the device 10 is internally configured as described above with reference to FIGS. 2-8, that is, the device 40 includes a selection unit 56, a search unit 60, and optionally a derivation unit 66. The selection unit 56 performs selection for the purpose of uneven streaming, i.e., selecting media segments such that the media content is encoded in the selected and retrieved media segments with spatially varying quality and / or with unencoded portions. In addition, however, the device 40 constitutes a log message sender 80 which sends, for example, logged log messages to the server 20 and to an evaluation device 82 . An instantaneous magnitude or statistic measuring the spatial position and / or movement of the first portion 64 instantaneous magnitudes or statistics measuring the quality of the time-varying spatial scene as encoded in the selected media segment and as visible to the view section 28; and / or an instantaneous magnitude or statistic measuring the quality of the first portion or the quality of the time-varying spatial scene 30 as encoded in the selected media segment and as visible in the view section 28;

[0057] The motivations are as follows:

[0058] As mentioned before, a user reporting mechanism is needed to be able to derive statistics such as most interesting regions, speed-probability pairs, etc. Additional DASH metrics are needed to those defined in Annex D of ISO / IEC 23009-1.

[0059] The DASH client will return to the metrics server (which can be the same as the DASH server or something else) the characteristics of the end device in terms of FoV. One metric would be the client's FoV as a DASH metric. TIFF2025131610000002.tif20161

[0060] One metric is ViewportList: DASH clients send back to the metrics server (which can be the same as the DASH server or another server) the viewports that each client was monitoring at that time. An instantiation of such a message looks like this: TIFF2025131610000003.tif62161

[0061] Regarding viewport (region of interest) messages, DASH clients can be asked to report every time a viewport change occurs, potentially with a given granularity (whether or not to avoid reporting very small movements) or a given periodicity. Such messages can be included in the MPD as an attribute @reportViewPortPeriodicity or an element or descriptor. They can also be indicated out-of-band, in a SAND message or in other ways.

[0062] The viewport can also notify at a tile granularity.

[0063] Additionally or alternatively, the log message may report on other current scene-related parameters that change in response to user input, such as the current user distance from the scene center and / or the current view depth, or any of the parameters discussed below with respect to FIG. 10.

[0064] Another metric is ViewportSpeedList, which allows DASH clients to display the movement speed for a particular viewport when movement occurs. TIFF2025131610000004.tif84161

[0065] This message is only sent if the client performs a viewport movement, although the server can indicate that the message should only be sent if the movement is significant, as in the previous case. Such a setting can indicate a pixel size, angle, or the amount that needs to change in order for the message to be sent, as in @minViewportDifferenceForReporting.

[0066] Another important aspect of VR-DASH services where asymmetric quality is provided as described above is to evaluate how quickly a user switches from an asymmetric representation or a set of unequal quality / resolution representations for a viewport to another representation or set of representations that are more appropriate for the other viewport. Such a metric allows the server to derive statistics that help understand the relevant factors affecting QoE. Such a metric might look like this: TIFF2025131610000005.tif87161

[0067] Alternatively, the aforementioned periods can be given as average values. TIFF2025131610000006.tif60161

[0068] As with other DASH metrics, all of these metrics can additionally include the time the measurement was taken. TIFF2025131610000007.tif20161

[0069] In some cases, content of uneven quality is downloaded, and if poor quality (or a mix of good and poor quality) content is displayed for a long enough time (maybe just a few seconds), the user may become frustrated and terminate the session. As a condition for ending the session, the user can send a message with the quality displayed over the last X time intervals. TIFF2025131610000008.tif50161

[0070] Alternatively, it can report the difference in maximum quality, or the max#quality and min#quality of the viewport.

[0071] As becomes clear from the above discussion regarding Fig. 8a, in order for a tile-based DASH streaming service operator to configure and optimize its service in a meaningful way (e.g., in terms of resolution ratio, bitrate and segment duration), it is advantageous if the service operator can derive statistics that require a client reporting mechanism, examples of which are described above. In addition to Annex D of Non-Patent Document 1 ([A1]), additional DASH metrics to those defined above are provided below.

[0072] Imagine a tile-based streaming service using stereoscopically projected video, as shown in FIG. 1. Client-side reconstruction is illustrated in FIG. 8b, where small circles 198 indicate the projection of a two-dimensional distribution of horizontally and vertically conformally distributed view directions within the client's viewport 28 onto the image area covered by individual tiles 50. Hatched tiles indicate high-resolution tiles and thus form the high-resolution portion 64, while unhatched tiles 50 indicate low-resolution tiles and thus form the low-resolution portion 66. It can be seen that partially lower-resolution tiles are presented to the user when the viewport 28 has changed since the last segment selection and download update, thereby determining the resolution of each tile on a cube where the projection surface or pixel array of tiles 50 is encoded in the downloadable segments 58.

[0073] While the above description rather generally presents feedback or log messages indicating the quality of the video presented to the user in the viewport, the following outlines more specific and advantageous metrics that may be applied in this regard. The metrics described here are reported from the client side and are sometimes referred to as effective viewport resolution. They are intended to indicate to the service operator the effective resolution of the client's viewport. If the reported effective viewport resolution indicates that the user was only presented with a resolution that was towards the resolution of the low-resolution tiles, the service operator may change the tiling configuration, resolution ratio, or segment length accordingly to achieve a higher effective viewport resolution.

[0074] One embodiment would be the average number of pixels in the viewport 28, measured in a projected view in which there is a pixel array of tiles 50 coded into a segment 58. The measurement may distinguish or be specified relative to the horizontal 204 and vertical 206 directions in relation to the covered field of view (FoV) of the viewport 28. The table below shows possible examples for appropriate syntax and semantics that may be included in a log message to know a rough viewport quality measurement. TIFF2025131610000009.tif42131

[0075] The horizontal and vertical decomposition can be suspended by instead using a scalar value for the average pixel count. Together with an indication of the aperture or size of the viewport 28, which may also be reported to the recipient of the log message, i.e., the evaluator 82, the average count indicates the pixel density within the viewport.

[0076] It may be advantageous to make the FoV considered for the metric smaller than the FoV of the viewport actually presented to the user, thereby excluding areas toward the borders of the viewport that are used only for peripheral vision and therefore do not affect subjective quality perception. This alternative is indicated by the dashed line 202 surrounding pixels in such a central portion of the viewport 28. Reporting the FoV 202 considered for the reported metric relative to the overall FoV of the viewport 28 may also be indicated to the log message receiver 82. The following table shows the corresponding extension of the previous example. TIFF2025131610000010.tif66132

[0077] According to a further embodiment, the average pixel density is not measured by averaging the quality in a spatially uniform manner within the projection plane, as is the case in the examples described so far with the EffectiveFoVResolutionH / V, but rather this averaging is weighted non-uniformly across the pixels, i.e., the projection plane. The averaging may be performed in a spherically uniform manner. As an example, the averaging may be performed uniformly over sample points distributed as in a circle 198. In other words, the averaging may be performed by weighting the local density with a weight that decreases quadratically with increasing local projection plane distance and increases according to the sine of the local slope of the projection relative to the line connecting it to the viewpoint. The message includes an optional (flag-controlled) step to accommodate the inherent oversampling of some available projections (e.g., equirectangular projections) by, for example, using a uniform spherical sampling grid. Some projections do not have significant oversampling issues, and forcing the computational elimination of oversampling may lead to unnecessary complexity issues. This should not be limited to equirectangular projections. In reporting, it is not necessary to distinguish between horizontal and vertical resolution, but it is possible to combine them. One embodiment is shown below. TIFF2025131610000011.tif136132

[0078] Applying conformal uniformity to the averaging, FIG. 8e shows points 302 equiangularly distributed horizontally and vertically on a sphere 304 centered at viewpoint 306 projected onto a tile's projection plane 308, here a cube, as long as it is within viewpoint 28. This results in pixel density averaging of pixels 308 arrayed in columns and rows within the projection plane, setting a local weight of pixel density depending on the local density of points 302 onto projection plane 198. An entirely similar approach is depicted in FIG. 8f. Here, points 302 are equally spaced within a viewport plane perpendicular to view direction 312, i.e., evenly distributed horizontally and vertically in rows and columns, and their projection onto projection plane 308 defines points 198, whose local density controls the weight that local pixel density 308 (which varies due to high-resolution and low-resolution tiles within viewport 28) contributes to the average. In the above example, such as the updated table, the alternative in FIG. 8f can be used instead of the one shown in FIG. 8e.

[0079] In the following, embodiments of further types of log messages are described for a DASH client 10 (see FIG. 1) having multiple media buffers 300 as exemplarily shown in FIG. 8a, i.e., a DASH client 10 forwarding downloaded segments 58 for subsequent decoding by one or more decoders 42. The distribution of segments 58 across buffers can be performed in different ways. For example, distribution can be performed such that certain regions of a 360 video are downloaded separately from one another or buffered in separate buffers after downloading. The following example illustrates different distributions by showing which tiles T (with logarithms totaling 25), indexed #1 to #24 as shown in FIG. 1, are encoded into individually downloadable representations R #1 to #P at which qualities Q of #1 to #M (1 being the highest and M being the lowest), and how these P-representations R can (optionally) be grouped into adaptation sets A in the MPD, indexed #1 to #S, and how segments 58 of P-representations R can be distributed across buffers B, indexed #1 to #N. TIFF2025131610000012.tif119152

[0080] Here, representations are provided by a server and advertised for download in the MPD, each of them relating to one tile 50, i.e. one section of a scene. The representations relating to one tile 50 but encoding this tile 50 with different qualities will be summarized in an adaptation set, where the grouping is arbitrary but precisely this grouping is used for the association to buffers. Thus, according to this example, there will be one buffer per tile 50, in other words per viewport (view section) encoding. Another set of representations and distributions are: TIFF2025131610000013.tif117154

[0081] According to this example, each representation will cover the entire area, but with higher quality regions focused on one hemisphere, while lower quality regions will be used for the other hemisphere. Representations that differ in the exact quality used in this way, i.e., with equal positions in the high-quality hemisphere, are simply combined into one adaptive set and distributed according to this property across, here for example, six buffers.

[0082] Therefore, in the following description, we assume that such a distribution to buffers according to different viewport encodings, video sub-regions such as tiles, etc., associated with Adaptation Sets, etc., is applied. Figure 8c shows the buffer fullness levels over time for two separate buffers, e.g., Tile 1 and Tile 2, in a tile-based streaming scenario, as shown in the last, but smallest, table. Enabling a client to report the fullness levels of all its buffers allows a service operator to correlate the data with other streaming parameters to understand the impact of his service configuration on the Quality of Experience (QoE).

[0083] The benefit therefrom is that buffer fullness of multiple client-side media buffers is reported with metrics that can be identified and associated with buffer types, e.g. Tile Viewport ·Region AdaptationSet ·Representation Low quality version of the whole content

[0084] One embodiment of the present invention is given in Table 1, which defines metrics for reporting buffer level status events for each buffer along with their identification and association. Table 1: List of buffer levels TIFF2025131610000014.tif101130

[0085] A further embodiment using viewport dependent encoding is as follows.

[0086] In a viewport-dependent streaming scenario, a DASH client downloads and pre-buffers several media segments associated with a particular viewing direction (viewport). If the amount of pre-buffered content is too large and the client changes its viewing direction, the portion of the pre-buffered content that plays after the viewport change will not be displayed, and the respective media buffer will be purged. This scenario is depicted in Figure 8d.

[0087] Another embodiment relates to traditional video streaming scenarios with multiple representations (quality / bitrate) of the same content, and possibly spatially uniform quality at which the video content is encoded. In that case the distribution is: TIFF2025131610000015.tif69152

[0088] That is, here, each representation covers the entire scene, which may not be, for example, a panoramic 360 scene, but may be of different quality, i.e., spatially uniform quality, and these representations will be distributed separately over the buffer. All examples shown in the last three tables should be treated as non-limiting in how the segments 58 of the representations provided by the server are distributed over the buffer. Different methods exist, and the rules can be based on the membership of the segments 58 to the representations, the membership of the segments 58 to the adaptation set, the direction of the locally increased quality of the spatially non-uniform coding of the scene to which each segment belongs, and the quality with which the scene is coded into each segment belongs as follows:

[0089] The client can maintain a buffer for each representation and, when available throughput increases, decide to flush the remaining lower quality / bitrate media buffer before playback and download higher quality media segments for the duration within the existing lower quality / bitrate buffer. Similar embodiments can be built for streaming based on tiles and viewport dependent encoding.

[0090] Service operators may be interested in understanding the amount and type of data downloaded without being presented, as this introduces costs without benefit on the server side and reduces quality on the client side. Therefore, the present invention provides reporting metrics that correlate two events, "media download" and "media presentation," for easy interpretation. The present invention avoids the need to analyze redundantly reported information about the download and playback status of each media segment, and allows efficient reporting of purge events only. The present invention also includes identifying buffers as described above and associating them with types. One embodiment of the present invention is shown in Table 2. Table 2: List of deletion events TIFF2025131610000016.tif105164

[0091] FIG. 9 illustrates a further embodiment of how device 40 may be advantageously implemented. Device 40 of FIG. 9 can correspond to any of the examples described above with respect to FIGS. 1-8. That is, it may include a lock messenger, possibly as described above with respect to FIG. 8a, but need not have, and may use, information 68 as described above with respect to FIGS. 2 and 3 or information 74 as described above with respect to FIGS. 5-7c, but need not have. With respect to FIG. 9, however, unlike the description of FIGS. 2-8, it is assumed that a tile-based streaming approach is actually applied. That is, scene content 30 is provided by server 20 in the tile-based manner described as an option above with respect to FIGS. 2-8.

[0092] 9. Although the internal structure of apparatus 40 may differ from that shown in FIG. 9, apparatus 40 is illustratively shown to comprise selection unit 56 and search unit 60 already discussed above with respect to FIGS. 2 to 8, and optionally a derivation unit 66. However, apparatus 40 further comprises a media presentation description analyzer 90 and an adaptation unit 92. The MPD analyzer 90 is for deriving, from the media presentation description obtained from server 20, at least one version in which time-varying spatial scene 30 is provided for tile-based streaming, and for each of the at least one version, an indication of benefit requirements for each version of the time-varying spatial scene to benefit from tile-based streaming. The meaning of "version" will become clear from the following description. In particular, adaptation unit 92 matches the benefit requirements thus obtained with the device capabilities of apparatus 40 or another device interacting with apparatus 40, such as the decoding capabilities of one or more decoders 42, the number of decoders 42, etc. The background or idea underlying the concept of FIG. 9 is as follows. In a tile-based approach, imagine that, assuming a fixed size of view section 28, a fixed number of tiles must be included in section 62. Furthermore, it can be assumed that media segments belonging to one tile form one media or video stream that must be decoded by a separate decoding instantiation, separate from the decoding of media segments belonging to another tile. Thus, a moving aggregation of a certain number of tiles in section 62, whose corresponding media segments are selected by selector 56, requires the presence of respective decoding resources, e.g., in the form of a corresponding number of decoding instances, i.e., a certain decoding capability, such as a corresponding number of decoders 42. If such a number of decoders does not exist, the service provided by server 20 may not be useful to the client. Therefore, the MPD provided by server 20 may indicate "benefit requirements," i.e., the number of decoders required to use the provided service. However, server 20 may provide MPDs for different versions.That is, different MPDs for different versions may be available by the server 20, or the MPD provided by the server 20 may be internally configured to distinguish between different versions for which the service may be used. For example, the versions may differ in field of view, i.e., the size of the field of view section 28. Fields of view of different sizes may manifest as different numbers of tiles in section 62 and therefore may have different benefit requirements, for example, in that these versions may require different numbers of decoders. Other examples are possible as well. For example, versions with different fields of view may include the same number of media segments 46, but according to another example, different versions in which a scene 30 is provided for tile streaming on the server 20 may even differ in the number of media segments 46 included according to the corresponding version. For example, the tile division according to one version may be coarser than the tile division of a scene according to another version, thereby requiring, for example, a smaller number of decoders.

[0093] The matcher 92 either matches the benefiting requirements and therefore selects the corresponding version, or rejects all versions entirely.

[0094] However, useful requirements may further relate to the profile / level that one or more decoders 42 must be able to handle. For example, the DASH MPD includes multiple locations that allow for indicating a profile. A typical profile describes the attributes, elements that may be present in the MPD, and the video or audio profile of the streams provided for each representation.

[0095] A further example of a beneficial requirement concerns, for example, the client-side ability to move the viewport 28 throughout the scene. A beneficial requirement could indicate a necessary viewport speed that should be available for the user to move the viewport in order to be able to truly enjoy the provided scene content. A conformer would, for example, check whether this requirement is met with an inserted user input device such as, for example, an HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in a directional sense, a set of "sufficient types of input devices" could be indicated as a beneficial requirement.

[0096] In spherical video tile streaming services, there are too many configuration parameters that can be dynamically set, such as the number of quality levels and tiles. If tiles are independent bitstreams that need to be decoded by separate decoders, a large number of tiles might make it impossible for a hardware device with fewer decoders to simultaneously decode all bitstreams. A possible solution would be to leave this as a degree of freedom and have the DASH device analyze all possible representations and count the number of decoders required to decode all representations or a given number that covers the device's FoV, thereby determining whether the DASH client can consume the content. However, a more intelligent solution for interoperability and capability negotiation would be to use signaling in the MPD that maps to a type of profile, which is used as a promise to the client that the provided VR content can be consumed if the profile is supported. Such signaling should be in the form of a URN, such as urn::dash-mpeg::vr::2016, which can be packed either at the MPD level or in an adaptation set. This profiling implies that N decoders of X profile are sufficient to consume the content. Depending on the profile, a DASH client can ignore or accept the MPD or parts of the MPD (Adaptation Set). Furthermore, there are some mechanisms, such as Xlink or MPD Chaining, that do not include all information, such that little signaling regarding the selection is available. In such a situation, a DASH client cannot derive whether it can consume the content. Regardless of whether it makes sense to implement Xlink or MPD Chaining or a similar mechanism, it is necessary to expose the decoding capabilities in terms of the number of decoders and the profile / level of each decoder using such an urn (or something similar) so that a DASH client can do so.The signaling may also imply different operating points such as N decoders with X profiles / levels, or Z decoders with Y profiles / levels.

[0097] FIG. 10 further illustrates that any of the above embodiments and explanations presented with respect to FIGS. 1-9 regarding clients, devices 40, servers, etc., can be extended to the extent that the services provided vary not only in terms of time but also in terms of other parameters. For example, FIG. 10 illustrates a variation of FIG. 1 in which multiple available media segments provided on a server describe scene content 30 for different positions of a view center 100. While the schematic diagram shown in FIG. 10 depicts the scene center as varying only along one direction, X, it is clear that the view center can vary along two or more spatial directions, such as two or three dimensions. This corresponds, for example, to a user's change in user position within a virtual environment. Depending on the user's position within the virtual environment, the available views change, and the scene 30 changes accordingly. Thus, in addition to the media segments describing the scene 30 subdivided into tiles and temporal segments and different qualities, additional media segments describe different content of the scene 30 for different positions of the scene center 100. 1 to 9 , the media presentation description will include such calculation rules that depend on one or more additional parameters in addition to the parameters described above with respect to FIGS. 1 to 9 . The parameter X may be quantized to any of the levels at which the respective scene representation is encoded by the corresponding media segment in the plurality 46 in the server 20.

[0098] Alternatively, X can be a parameter that defines the view depth, i.e., the distance in the radial direction from the scene center 100. Providing a scene in different versions with different view center portions X allows a user to "walk" through the scene, while providing a scene in different versions with different view depths allows a user to "radially zoom" back and forth through the scene.

[0099] Therefore, in the case of multiple non-concentric viewports, the MPD can use additional signaling regarding the position of the current viewport, such as at the segment, representation or period level.

[0100] Non-concentric spheres: The spatial relationship of the different spheres should be signaled in the MPD for the user to navigate. This can be done using coordinates (x,y,z) in any units relative to the diameter of the sphere. The diameter of the sphere should also be displayed for each sphere. The sphere is "enough enough" to be used with additional space added so that the content is fine for the user in the center. If the user navigates beyond the indicated diameter, another sphere should be used to display the content.

[0101] An exemplary signaling of viewports can be done relative to a given center point in space. Each viewport is indicated relative to that center point. In MPEG DASH, this can be indicated, for example, in an AdaptationSet element. TIFF2025131610000017.tif102133

[0102] Finally, FIG. 11 illustrates that information such as or similar to that described above with reference to reference numeral 74 may be present in the video bitstream 110 in which the video 112 is encoded. A decoder 114 decoding such video 110 may use the information 74 to determine the size of a focus region 116 within the video 112 upon which decoding power for decoding the video 110 should be focused. The information 74 may be conveyed, for example, within an SEI message of the video bitstream 110. For example, the focus area may be decoded exclusively, or the decoder 114 may be configured to start decoding each picture of the video at the focus area rather than, for example, the upper-left picture corner, and / or the decoder 114 may cease decoding each picture of the video upon decoding the focus region 116. Additionally or alternatively, the information 74 may simply be present in the data stream to be forwarded to a subsequent renderer or viewport control or a client's streaming device or segment selector to determine which spatial sections to cover or which segments to download or stream to cover with improved or predetermined quality. Information 74 may indicate, for example, a preferred area, as a recommendation for positioning view section 62 or section 66 to coincide with, cover, or track this area, as outlined above, and may be used by a client segment selector. As with the discussion of Figures 4-7c, information 74 may set absolute dimensions of region 116, such as the number of tiles, or may set the speed of moving region 116, e.g., according to user input, thereby scaling region 116 to increase with increasing speed instructions, in order to spatiotemporally track interesting content of the video.

[0103] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0104] The signals generated above, such as streaming signals, MPDs or any other of the signals mentioned above, can be stored in a digital storage medium or transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0105] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. They can be implemented using digital storage media, such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or FLASH memories, on which electronically readable control signals are stored. They cooperate (or can cooperate) with a programmable computer system so that the respective methods are executed. Thus, the digital storage media can be computer-readable.

[0106] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0107] Generally, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer, which program code can be stored on, for example, a machine-readable carrier.

[0108] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0109] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0110] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0111] A further embodiment of the inventive methods is, therefore, a data stream or sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transmitted via a data communication connection, for example the Internet.

[0112] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0113] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0114] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0115] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0116] The apparatus described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0117] The devices described herein, or any components of the devices described herein, may be implemented at least in part in hardware and / or software.

[0118] The methods described herein can be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0119] The methods described herein, or any components of the apparatus described herein, may be performed at least in part by hardware and / or software.

[0120] The above-described embodiments are merely illustrative for explaining the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims, and not by the specific details presented for purposes of illustrating and describing the embodiments herein. [Prior art documents] [Non-patent literature]

[0121] [Non-Patent Document 1] [A1] ISO / IEC 23009-1:2014, Information technology -- Dynamic adaptive streaming over HTTP (DASH) -- Part 1: Media presentation description and segment formats

Claims

1. An apparatus for streaming media content relating to a time-varying spatial scene (30), comprising: selecting (56) a media segment from a plurality (46) of media segments (58) available on the server (20); the selected media segments (60) from the server (20), wherein the device: performing said selection such that said selected media segment has at least a spatial section (62) of said time-varying spatial scene (30) encoded in such a way that a first portion (64) of said spatial section is encoded into said selected media segment with a predetermined quality, and a second portion (66) of said time-varying spatial scene that is spatially adjacent to said first portion (64) is encoded into said selected media segment with a further quality that satisfies a predetermined relationship with respect to said predetermined quality; and The predetermined relationship is derived from information (68) contained in the selected media segments and / or signaling obtained from the server (20). An apparatus configured to:

2. 2. The apparatus of claim 1, wherein each media segment of the plurality (46) of media segments is encoded in an associated spatiotemporal portion of the time-varying spatial scene (30) at an associated one of a set of quality levels.

3. 3. The apparatus of claim 2, wherein each spatiotemporal portion of the time-varying spatial scene (30) encoded into the plurality of media segments (46) is a time segment (54) of the time-varying spatial scene (30) in each of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

4. 4. Apparatus according to any one of claims 1 to 3, wherein said information (64) indicates a tolerance value for a measure of difference between said further quality and said predetermined quality.

5. 5. The device of claim 4, wherein the information (68) indicates the tolerance for the measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section (28).

6. 6. The apparatus of claim 4, wherein the information indicates the tolerance for the measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section as a list of pairs of a respective distance to the view section and a corresponding tolerance for the measure of the difference beyond the respective distance.

7. 7. The apparatus of claim 4, wherein the information (64) indicates the tolerance for the measure of difference between the further quality and the predetermined quality, whereby the tolerance increases as the distance to the view section increases.

8. 8. The apparatus according to claim 4, wherein the information (64) indicates a tolerance value for the measure of difference between the further quality and the predetermined quality together with an indication of the maximum allowable time interval that the second part may have within the view section together with the first part.

9. 9. The apparatus of claim 8, wherein the information (64) indicates a further tolerance value for the measure of difference between the further quality and the predetermined quality, together with an indication of a further maximum allowable time interval that the second portion may be within the view section along the first portion.

10. 10. Apparatus according to any one of claims 4 to 9, characterized in that the information (64) is time-varying and / or spatially-varying.

11. 11. The device according to any one of claims 1 to 10, wherein the information (64) indicates pairs of allowed simultaneous settings for the further quality and the predetermined quality.

12. 12. The apparatus of claim 1, wherein the apparatus is configured to perform the selection such that the first portion follows a time-varying view section (28) of the time-varying spatial scene (30).

13. 13. The apparatus of claim 12, wherein the apparatus is configured such that the spatial position of the time-varying view section (28) varies according to user input.

14. 14. Apparatus according to any preceding claim, wherein the apparatus is configured to determine the first portion to correspond to a region of interest.

15. The apparatus of claim 14 , wherein the apparatus is configured to retrieve information about the ROI from the server.

16. 1. A streaming server for media content relating to time-varying spatial scenes, comprising: making available a plurality of media segments for retrieval by the device, whereby said device is enabled to select a media segment for retrieval having at least a spatial section of a time-varying spatial scene encoded in such a way that a first portion of the spatial section is encoded in the selected media segment with a predetermined quality and a second portion of the time-varying spatial scene that is spatially adjacent to the first portion is encoded in the selected media segment with a further quality; 5. A streaming server configured to include information regarding a predetermined relationship within said media segments and / or signal information having a predetermined relationship to be fulfilled by said further quality relative to a predetermined quality by signaling to said device.

17. 17. The streaming server of claim 16, wherein each media segment of the plurality of media segments encodes an associated spatiotemporal portion of the time-varying spatial scene (30) at an associated one of a set of quality levels.

18. 18. The streaming server of claim 17, wherein each spatiotemporal portion of the time-varying spatial scene (30) encoded into the plurality of media segments is a time segment (54) of the time-varying spatial scene (30) in a respective one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

19. Streaming server according to any of claims 16 to 18, wherein said information (64) indicates a tolerance value for a measure of difference between said further quality and said predetermined quality.

20. 20. The streaming server of claim 19, wherein the information (68) indicates the tolerance for a measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section (28).

21. 21. A streaming server according to claim 19 or 20, wherein the information indicates the tolerance for a measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section (28) by a list (28) of pairs of respective distances to the view section (28) and corresponding tolerances for a measure of difference beyond the respective distances.

22. 22. A streaming server according to any one of claims 19 to 21, wherein the information indicates the tolerance value for a measure of difference between the further quality and the predetermined quality, whereby the tolerance value increases as the distance to the view section increases.

23. 23. A streaming server according to any one of claims 19 to 22, wherein the information (64) indicates a tolerance value for a measure of difference between the further quality and the predetermined quality together with an indication of a maximum allowed time interval that the second part may be within the view section together with the first part.

24. 24. The streaming server of claim 23, wherein the information (64) indicates a further tolerance on a measure of difference between the further quality and the predetermined quality together with an indication of a further maximum allowed time interval that the second part may be within the view section together with the first part.

25. Streaming server according to any of claims 19 to 24, wherein said information (64) is time and / or spatially varying.

26. Streaming server according to any of claims 19 to 25, wherein said information (64) indicates pairs of allowed simultaneous settings for said further quality and said predetermined quality.

27. 27. A streaming server according to any of claims 19 to 26, wherein the device is configured to transmit information about a ROI to the device.

28. - information about calculating addresses of a plurality of media segments, whereby an apparatus using said information can select and retrieve a media segment from a plurality of media segments encoding at least one spatial section of a time-varying spatial scene in such a way that a first part of said spatial section is encoded in said selected media segment with a predetermined quality and accordingly a second part of the time-varying spatial scene that is spatially adjacent to said first part is encoded with a further quality by said selected media segment; and information (64) about a predetermined relationship to be satisfied by a further quality with respect to said predetermined quality.

29. 30. The media presentation description of claim 28, wherein each media segment (46) of the plurality of media segments encodes therein an associated spatiotemporal portion of the time-varying spatial scene (30) at an associated one of a set of quality levels.

30. 30. The media presentation description of claim 29, wherein each spatiotemporal portion of the time-varying spatial scene (30) encoded in the plurality of media segments (46) is a time segment (54) of the time-varying spatial scene (30) in a respective one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

31. 31. A media presentation description according to any of claims 28 to 30, wherein said information (64) indicates a tolerance for a measure of difference between said further quality and said predetermined quality.

32. 32. The media presentation description of claim 31 , wherein the information (68) indicates the tolerance for a measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section (28).

33. 33. The media presentation description of claim 31, wherein the information indicates the tolerance for a measure of difference between the further quality and the predetermined quality in a manner dependent on the distance to the view section (28) by a list of pairs of a respective distance to the view section (28) and a corresponding tolerance for the measure of difference beyond the respective distance.

34. 34. The media presentation description of any of claims 31 to 33, wherein the information (64) indicates a tolerance for a measure of difference between the further quality and the predetermined quality, such that the greater the distance to the view section, the greater the tolerance.

35. 35. The media presentation description of any of claims 31 to 34, wherein the information (64) indicates the tolerance for a measure of difference between the further quality and the predetermined quality together with an indication of a maximum allowed time interval that the second part may be in a view section with the first part.

36. 36. The media presentation description of claim 35, wherein the information (64) indicates a further tolerance on a measure of the difference between the further quality and the predetermined quality together with an indication of a maximum further allowed time interval that the second part that is in the view section together with the first part can tolerate.

37. 37. A media presentation description according to any of claims 31 to 36, wherein said information (64) is time-varying and / or spatially-varying.

38. 38. A media presentation description according to any of claims 31 to 37, wherein said information (64) indicates pairs of allowed simultaneous settings for said further quality and said predetermined quality.

39. 39. The media presentation description of any of claims 31 to 38, wherein the media presentation description includes information about a region of interest.

40. An apparatus for streaming media content relating to a time-varying spatial scene (30), comprising: configured to select a media segment from a plurality (46) of media segments (58) available on the server (20); The device comprises: a first part (64) of said spatial section (62) is encoded into a selected media segment with a predetermined quality, and a second part (66; 72) of the spatial scene, spatially adjacent to said first part (64) and varying over time, is not encoded into said selected media segment or is encoded into a selected media segment with a quality lower than said predetermined quality, in such a way that the first portion (64) follows a time-varying view portion (28) of the time-varying spatial scene (30); performing said selection such that said selected media segment encodes therein at least a spatial section (62) of said time-varying spatial scene (30); An apparatus configured to set the size and / or position of the first portion (64) in response to information (74) contained in the selected media segment and / or signaling obtained from the server.

41. 41. The apparatus of claim 40, wherein the information (74) indicates a size in the form of an increment relative to or scaling of the time-varying view section size.

42. 42. Apparatus according to claim 40 or 41, wherein said information (74) indicates a predetermined value for a measure of spatial velocity of said view section.

43. for a default percentile of users whose measured spatial velocity does not exceed a predetermined value; and / or together with a percentile value indicating the percentile of users whose measured spatial velocity does not exceed a predetermined value; and / or together with a percentile value indicating the percentile of users whose view section falls within a given region; and / or along with a hint indicating one or more types of user input controlling movement of the view section for which the predetermined value is applicable; 43. The apparatus of claim 42, wherein the information (74) indicates the predetermined value for the measure of the spatial velocity of the view section.

44. 44. Apparatus according to claim 42 or 43, configured to set the size so that the larger the predetermined value for the measure of the spatial velocity of the view section d is, the larger the size is.

45. 45. Apparatus according to claim 40 or 44, wherein said information (74) indicates a predetermined value for a measure of probability for the direction of movement of said view section.

46. 46. ​​The apparatus of claim 45, configured to configure the first portion to extend in each direction as the likelihood of each direction of movement increases. 。

47. 47. Apparatus according to any of claims 40 or 46, wherein said information (74) indicates said predetermined spatial velocity in a time-varying and / or space-varying and / or direction-varying manner.

48. 48. The apparatus of claim 40 or 47, wherein the first portion (64) and the spatial portion (62) are configured to coincide.

49. 49. The apparatus of claim 40 or 48, wherein each of the spatiotemporal portions of the time-varying spatial scene (30) encoded into the plurality of media segments is a time segment (54) of the time-varying spatial scene (30) in a respective one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

50. 50. The apparatus of claim 49, wherein each media segment of the plurality of media segments has an associated spatiotemporal portion of the time-varying spatial scene encoded therein at an associated one of a set of quality levels.

51. 41. The apparatus of claim 40, wherein the apparatus is configured to set the size independently of the size of the time-varying view section in response to the information.

52. 41. The device of claim 40, wherein the information includes different values ​​for the size for different size options of the time-changing view section, and the device uses the value included in the information for the size option that matches the actual size of the time-changing view section.

53. 53. Apparatus according to claim 51 or 52, wherein the information indicates a size in number of tiles.

54. 1. A streaming server for media content relating to time-varying spatial scenes, comprising: making available a plurality of media segments for retrieval by the device, whereby the device is capable of selecting a media segment for retrieval having at least a spatial section of a time-varying spatial scene encoded therein in a manner such that a first portion of said spatial section is encoded in said selected media segment with a predetermined quality, and accordingly a second portion of said time-varying spatial scene that is spatially adjacent to the first portion is not encoded in said selected media segment or is encoded in said selected media segment with an even lower quality than said predetermined quality, such that said first portion follows a time-varying view section of said time-varying spatial scene; and a streaming server configured to notify the device of a method for setting the size and / or position of said first portion within a media segment and / or by signaling said device.

55. A signal defining a media presentation description, information about calculating addresses of a plurality of media segments, wherein an apparatus using said information is able to select and retrieve a media segment from said plurality of media segments thus encoding at least a spatial section of said time-varying spatial scene, whereby a first part of said spatial section is encoded into said selected media segment with a predetermined quality, and accordingly a second part of said time-varying spatial scene that is spatially adjacent to said first part is not encoded into said selected media segment or is encoded into said selected media segment with a lower quality compared to said predetermined quality, such that said first part follows a time-varying view section of said time-varying spatial scene; and information about the size and / or how to set the size of said first portion within the media segment and / or by signaling to said device. A signal that defines a media presentation description.

56. 56. The signal of claim 55, wherein the information (74) indicates the size in the form of an increment to or scaling of the size of the time-varying view section.

57. 57. A signal according to claim 55 or 56, wherein said information (74) indicates a predetermined value for a measure of spatial velocity in said view section.

58. for a default percentile of users whose measured spatial velocity does not exceed a predetermined value; and / or together with a percentile value indicating the percentile of users whose measured spatial velocity does not exceed a predetermined value; and / or together with a percentile value indicating the percentile of users whose view section is within a given region; and / or together with a hint indicating one or more types of user input controlling movement of the view section for which the predetermined value is applicable.

58. The signal of claim 57, wherein the information (74) indicates the predetermined value for the measure of the spatial velocity of the view section.

59. 59. A signal according to claim 57 or 58, configured to effect said setting such that the greater the predetermined value for said measure of said spatial velocity of said view section, the greater the size.

60. 60. A signal according to claim 55 or 59, wherein the information (74) indicates a predetermined value for a measure of probability for the direction of movement of the view section.

61. 61. A signal as claimed in any one of claims 60, configured to configure the first portion to extend in each direction as each direction of movement becomes more likely.

62. 62. A signal according to claim 55 or 61, wherein said information (74) is indicative of said predetermined spatial velocity in a time-varying and / or spatially and / or directionally varying manner.

63. 63. A signal as claimed in either claim 55 or 62, wherein the first portion (64) and the spatial portion (62) are configured to coincide.

64. 64. The signal of claim 55 or 63, wherein each of the spatiotemporal portions of the time-varying spatial scene (30) encoded into the plurality of media segments is a time segment (54) of the time-varying spatial scene (30) in a respective one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.

65. 65. The signal of claim 64, wherein each media segment of the plurality of media segments encodes an associated spatiotemporal portion of the time-varying spatial scene at an associated one of a set of quality levels.

66. 56. The signal of claim 55, wherein the device is configured to the size in response to the information in a manner that is independent of the size of the time-varying view section.

67. 56. The signal of claim 55, wherein the information includes different values ​​for the size for different size options of the time-varying view section, and the device uses the value included in the information for the size option that matches the actual size of the time-varying view section.

68. 67. A signal according to claim 65 or 66, wherein the information indicates a size in number of tiles.

69. A video bitstream having video encoded into the video bitstream including signaling of focal regions of one or more sizes within the video on which decoding power for decoding the video should be focused, and recommended preferred view section regions of the video.

70. Deriving signaling (74) of the size of the focal region in the video from the video bitstream; A decoder for decoding the video from the video bitstream, configured to concentrate decoding power for decoding the video on the focal region (116).

71. 71. A decoder according to claim 70, configured to decode the focus region exclusively.

72. 71. The decoder of claim 70 configured to start decoding each picture of the video in the focus region.

73. 71. The decoder of claim 70, configured to cease decoding each picture of the video upon decoding the focus region.

74. 74. A decoder according to any of claims 70 to 73, wherein the signalling indicates the size in absolute terms or the decoder is configured to scale the size of the focus region according to a parameter included in the signalling.

75. An apparatus (30) for streaming media content relating to a time-varying spatial scene, comprising: In at least one version in which a time-varying spatial scene is provided for tile-based streaming, and for each of at least one version, an indication of a benefit requirement for each version of the time-varying spatial scene to benefit from the tile-based streaming. deriving (90) from a description of the media presentation; The device is configured to match (92) requirements for benefiting from at least one version with device capabilities of the device or another device interacting with the device regarding the tile-based streaming.

76. 76. The apparatus of claim 75, wherein the benefit requirement and the device capabilities relate to decoding capabilities.

77. 77. Apparatus according to claim 75 or 76, wherein the benefit requirements and the capabilities of the apparatus relate to a number of available decoders.

78. 78. Apparatus according to any of claims 75 to 77, wherein the benefit requirements and device capabilities relate to level and / or profile descriptors.

79. 80. An apparatus as described in any one of claims 75 to 78, wherein the benefit requirements and capabilities of the apparatus are related to the type of input device for moving a view section across the time-varying spatial scene (30), or the speed at which the view section is moved across the time-varying spatial scene (30) using the input device.

80. The device comprises: selecting a media segment from a plurality of media segments available on the server by calculating an address of the selected media segment using calculation rules included in the media presentation description; retrieving the selected media segment from the server (20) using the calculated address; the device is configured to perform the selection such that the selected media segment has encoded therein at least one spatial section (62) of a time-varying spatial scene (30); According to which a first part (64) of said spatial section (62) is encoded into said selected media segment with a predetermined quality, and accordingly a second part (66; 72) of said time-varying spatial scene is spatially adjacent to said first part (64) and is either not encoded into said selected media segment or is encoded into said selected media segment with a quality lower than said predetermined quality, and 80. Apparatus according to any one of claims 75 to 79, wherein the first portion (64) is configured to follow a time-varying view section (28) of the time-varying spatial scene (30).

81. a streaming server for streaming media content relating to said time-varying spatial scene (30), comprising: In at least one version, in which the time-varying spatial scene is provided for tile-based streaming, for each of the at least one version, an indication of benefit requirements for the respective version of the time-varying spatial scene to benefit from the tile-based streaming; deriving a media content stream from the streaming server, thereby enabling a device to stream the media content from the streaming server; providing a media presentation description (90); A streaming server configured to match (92) requirements for benefiting from at least one version with device capabilities of the device or another device interacting with the device regarding the tile-based streaming.

82. information about at least one version in which the time-varying spatial scene is provided for tile-based streaming; and for each of at least one version, an indication of benefit requirements for the respective version of the time-varying spatial scene to benefit from the tile-based streaming.

83. 83. The media presentation description of claim 82, wherein the benefit requirements and the device capabilities relate to decoding capabilities.

84. 84. The media presentation description of claim 82 or 83, wherein the benefit requirements and the device capabilities relate to a number of available decoders.

85. 85. A media presentation description according to any of claims 82 to 84, wherein the benefiting requirements and the device capabilities relate to level descriptors and / or profile descriptors.

86. 86. The media presentation description of any of claims 82 to 85, wherein the benefiting requirements and device capabilities relate to a type of input device for moving a view section across the time-varying spatial scene (30), or to a speed at which the view section is moved across the time-varying spatial scene (30) using the input device.

87. 87. The media presentation description of any of claims 82 to 86, further comprising computation rules for enabling the device to select a media segment from a plurality (46) of media segments (58) available on the server (20) by calculating an address of the selected media segment using the included computation rules.

88. An apparatus for streaming media content relating to a time-varying spatial scene (30), comprising: calculating an address of a media segment depending on the spatial viewport position and the at least one parameter, said media segment describing a time-varying spatial scene (30) and said at least one parameter; retrieving the media segment using the calculated address; A device configured as follows.

89. 89. The apparatus of claim 88, wherein the at least one parameter comprises one or more coordinates of a center of field of view and / or a depth of field.

90. 1. A media presentation description comprising: a media segment (30) describing a time-varying spatial scene; a computation rule for computing an address of a media segment depending on a spatial viewport position and at least one parameter; and the media segment (30) describing a time-varying spatial scene; and the at least one parameter for retrieving the media segment using the computed address.

91. 91. A streaming server enabling a device to stream media content relating to a time-varying spatial scene (30) from a server, the streaming server being configured to provide the media presentation description of claim 90.

92. 1. An apparatus for streaming media content relating to a time-varying spatial scene, comprising: configured to select a media segment from a plurality of media segments available on the server; the apparatus encodes a first portion of a time-varying spatial scene therein with improved quality compared to a spatial neighbor of the first portion or in a manner such that a spatial neighbor of the first portion is not encoded into the selected media segment; instantaneous measurements measuring the spatial position and / or movement of the first portion; and / or statistics, such as time averages, measuring the spatial position and / or movement of the first portion; and / or an instantaneous measurement measuring the quality of the time-varying spatial scene insofar as it is encoded in the selected media segment and visible in the view section; and / or an indication of a set of buffers (300) of devices involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment among the set of buffers, and the instantaneous buffer fullness of each set of buffers; and / or a measure (42) of the amount of the selected media segment that has not yet been output from the device's buffer to be decoded; and / or statistics, such as temporal averages, measuring the quality of the time-varying spatial scene insofar as it is encoded in the selected media segment and visible in the view section; and / or an instantaneous measurement measuring the quality of the first portion or the quality of the time-varying spatial scene as encoded in the selected media segment and visible in a view section; and / or a statistic, such as a temporal average, measuring the quality of the first portion or the quality of the time-varying spatial scene insofar as it is encoded in the selected media segment and visible in a view section; and / or the field of view covered by the view section, and / or instantaneous measurements measuring the user position or viewing depth relative to the scene center (100); and / or Statistics such as time averages that measure the user position or view depth relative to the scene center (100) The device is configured to emit a log message recording the

93. 93. The apparatus of claim 92, wherein the quality of the first portion or the quality of the time-varying spatial scene, as encoded into the selected media segment and visible within the view section, is measured as a duration during which a lower quality portion is displayed in the view section along with a higher quality portion.

94. 94. The apparatus of claim 92 or 93, configured to perform the selection such that the first portion (64) of the time-varying spatial scene populates the view section (28).

95. A measure measuring the average density of pixels falling into the view section (28) in which the time-varying spatial scene is encoded into the selected media segment.

95. The apparatus of claim 92, further comprising: configuring the transmission log message to record the instantaneous measurements measuring quality of the time-varying spatial scene insofar as the quality of the selected media segment is encoded into the selected media segment and displayed in a view section as one of:

96. 96. The apparatus of claim 95, wherein the measurement is configured to measure the average density of pixels by averaging the pixel densities in a spatially uniform manner about a pixel grid of a picture encoded in the selected media segment.

97. 96. The apparatus of claim 95, wherein the measurement is configured to measure the average density of pixels by averaging the pixel densities in a spatially non-uniform manner about a pixel grid of a picture encoded in the selected media segment.

98. Averaging the pixel density in a spatially uniform manner with respect to a pixel grid of a picture encoded in the selected media segment, or Averaging the pixel density in a spatially non-uniform manner with respect to a pixel grid of a picture encoded in the selected media segment.

96. The apparatus of claim 95, wherein the outgoing log message is configured to indicate whether the measurement value measures an average pixel density.

99. 99. The apparatus of claim 97 or 98, wherein averaging pixel density in a spatially non-uniform manner comprises: Averaging in a spherically uniform manner, or spatially uniformly averaging with respect to a viewport plane (310) perpendicular to a central viewing direction (312) of the view section (28); Devices corresponding to.

100. 100. The apparatus of claim 95, wherein the measurement is configured to measure an average density of pixels by averaging the pixel densities in a manner that limits the averaging to a central subsection of the view section (28), or to apply (202) a higher average weight around the central subsection compared to edge portions (204) of the view section.

101. 101. Apparatus according to any of claims 95 to 100, wherein the measurements are configured to measure average pixel density separately along a horizontal field of view cross-sectional axis (204) and a vertical field of view cross-sectional axis (206), respectively.

102. 102. Apparatus according to any of claims 95 to 101, configured to transmit log messages intermittently.

103. 103. A device according to any of claims 95 to 102, configured to send log messages at a rate controlled by a manifest file on the basis of which the device performs selection of the media segments for download.

104. Each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, the representations differing in one or more of the following: a scene section (50) encoded within said time-varying spatial scene; the quality in which the time-varying spatial scene is encoded spatial quality variations in which the temporally varying spatial scene is encoded; wherein the device stores, in the form of an association for each buffer, the description of distribution rules to be applied in distributing the selected media segments among a set of buffers; or The screen section The above quality spatial quality distribution expression 104. An apparatus according to any one of claims 95 to 103, configured to send log messages that are logged in the form of one or a combination of two or more of the following:

105. 105. Apparatus according to any of claims 95 to 104, wherein the representations are classified into adaptation sets according to one or more of the following: a scene section of a time-varying spatial scene is encoded therein; spatial quality variations in which the temporally varying spatial scene is encoded; The apparatus is configured to compose an outgoing log message that records the description of the distribution rules applied in distributing the selected media segments to a set of buffers in association with one adaptation set for each buffer or records the description of the distribution rules applied in distributing the selected media segments to a set of buffers in association with one representation for each buffer.

106. 106. An apparatus according to any of claims 95 to 105, wherein the apparatus is configured to send a log message recording a measure of the amount of the selected media segments that have not been output from a buffer of the apparatus to undergo decoding (42) in the form of a temporal measurement.

107. 107. The apparatus of claim 106, wherein the apparatus is configured such that a measurement of the amount of the selected media segment not output from a buffer of the apparatus for decoding (42) is defined independently of and / or in the form of a time measurement in units of time shorter than milliseconds and / or a time length of the media segment (58).

108. 108. The apparatus of any of claims 95 to 107, wherein the apparatus is configured to send log messages recording a measure of the amount of selected media segments not output from the buffer of the apparatus for decoding (42) in a format categorized by one or more of the following: a buffer of said decoder in which each media segment is buffered; a scene section encoded in each said media segment; the quality with which the time-varying spatial scene is encoded into the respective media segments; The apparatus, wherein the time-varying spatial scene is a spatial quality distribution encoded into the respective media segments.

109. A method for streaming media content relating to a time-varying spatial scene (30), comprising: selecting (56) a media segment from a plurality (46) of media segments (58) available on the server (20); retrieving (60) the selected media segment from the server (20), wherein the selected media segment has at least a spatial section (62) of the time-varying spatial scene (30) encoded according to: a first portion (64) of the spatial section is encoded into the selected media segment with a predetermined quality; and whereby a second portion (66) of the time-varying spatial scene is encoded into the selected media segment with a further quality that is spatially adjacent to the first portion (64) and satisfies a predetermined relationship with respect to the predetermined quality; The method further comprises deriving the predetermined relationship from information (68) contained in the selected media segments and / or signaling obtained from the server (20).

110. 1. A method for streaming media content relating to a time-varying spatial scene, comprising: - rendering a plurality of media segments available for retrieval by the device, whereby the device can select a media segment for retrieval comprising at least a spatial section of an encoded time-varying spatial scene in such a way and whereby a first portion of the spatial section is encoded in the selected media segment with a predetermined quality, and a second portion of the time-varying spatial scene, spatially adjacent to the first portion, is encoded in the selected media segment with a further quality; and / or by signaling to a device information about a predetermined relationship that is to be satisfied by a further quality within said media segment.

111. A method for streaming media content relating to a time-varying spatial scene (30), comprising: selecting a media segment from a plurality (46) of media segments (58) available on a server (20); The selected media segment According to which a first part (64) of a spatial section (62) is encoded in a selected media segment with a predetermined quality, and a second part (66; 72) of a time-varying spatial scene is spatially adjacent thereto, the first part (64) being either not encoded in the selected media segment or being encoded in a selected media segment with a quality lower than the predetermined quality, and The first portion (64) follows a time-varying view portion (28) of the time-varying spatial scene (30); and selecting the media segment from the server (20) such that the selected media segment has encoded therein at least one spatial section (62) of a time-varying spatial scene (30), such that the selected media segment is selected from the media segment from the server (20) such that the selected media segment is encoded therein; The method further includes setting a size of the first portion (64) in response to information (74) contained in the selected media segment and / or signaling obtained from the server.

112. 1. A method for streaming media content relating to a time-varying spatial scene, comprising: rendering a plurality of media segments available for retrieval by the device, thereby enabling the device to select a media segment for retrieval having at least a spatial section of a time-varying spatial scene encoded in such a way that a first portion of the spatial section is encoded in the selected media segment with a predetermined quality and that a second portion of the time-varying spatial scene spatially adjacent to said first portion is either not encoded in the selected media segment or is encoded in the selected media segment with a lower quality relative to the predetermined quality, so that the first portion follows a time-varying view section of the time-varying spatial scene; and c. signaling information about how to set the size of the first portion within the media segment and / or by signaling to the device.

113. 1. A method for decoding video from a video bitstream, comprising: deriving signaling of a size of a focus region within said video from a video bitstream; and concentrating decoding power to the focal area for decoding the video.

114. 1. A device-implemented method for streaming media content relating to a time-varying spatial scene (30), comprising: at least one version in which the time-varying spatial scene is provided for tile-based streaming; and and for each of the at least one version, an indication of a requirement for each version of the time-based varying spatial scene to benefit from tile-based streaming. deriving from a Media Presentation Description; and matching (92) requirements for tile-based streaming to benefit from at least one version with device capabilities of the device or another device interacting with the device.

115. 1. A method for streaming media content relating to a time-varying spatial scene, comprising: selecting a media segment from a plurality of media segments available on a server; retrieving the selected media content from the server and performing the selection such that the selected media segment has a first portion of a time-varying spatial scene encoded therein with improved quality compared to a spatial neighbor of the first portion or in a manner such that a spatial neighbor of the first portion is not encoded in the media segment; The method comprises: instantaneous measurements measuring the spatial position and / or movement of the first portion; and / or statistics, such as time averages, measuring the spatial position and / or movement of the first portion; and / or an instantaneous measurement of the quality of said time-varying spatial scene insofar as it is encoded in the selected media segment and is visible in the view section; and / or an indication of the set of buffers (300) of the devices involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment among the set of buffers, and the instantaneous buffer fullness of each set of buffers; and / or a measure (42) of the amount of the selected media segment that has not yet been output from the device's buffer to be decoded; and / or statistics, such as temporal averages, measuring the quality of the time-varying spatial scene insofar as it is encoded in the selected media segment and visible in the view section; and / or an instantaneous measurement measuring the quality of the first portion or the quality of the time-varying spatial scene as encoded in the selected media segment and visible in a view section; and / or a statistic, such as a temporal average, measuring the quality of the first portion or the quality of the time-varying spatial scene insofar as it is encoded in the selected media segment and visible in a view section; and / or the field of view covered by said field of view section, and / or an instantaneous measurement (100) measuring the user position or viewing depth relative to the scene center; and / or Statistics such as time averages that measure the user position or view depth relative to the scene center (100) The method further comprises sending a log message recording the

116. A method for streaming media content relating to a time-varying spatial scene (30), comprising: at least one version in which a time-varying spatial scene is provided for tile-based streaming; for each of at least one version, an indication of benefit requirements for each version of the time-varying spatial scene to benefit from the tile-based streaming; derivable, thereby allowing a device to stream the media content from a streaming server. providing a media presentation description (90); A method comprising: matching (92) requirements for tile-based streaming to benefit from at least one version with device capabilities of the device or another device interacting with the device.

117. 91. A method for enabling a device to stream media content relating to a time-varying spatial scene (30) from the server, comprising providing a media presentation description as claimed in claim 90.

118. 118. Computer program having a program code for performing the method according to any of claims 109 to 117, when the program runs on a computer.

Citation Information

Patent Citations

  • Video Data Stream Concept

    JP2015526006A

  • Top region of interest in the image

    JP2019519981A

  • Omnidirectional video transmission method, omnidirectional video reception method, omnidirectional video transmission device, and omnidirectional video reception device

    JP2019525675A

  • Determining a region of interest on the basis of a HEVC-tiled video stream

    WO2015197815A1

  • IEC23009-1