Spatially unequal streaming
By using spatial uneven method to select and encode media segments in VR streaming, the problem of slow response to viewport changes is solved, and user experience and bandwidth utilization efficiency is improved.
Patent Information
- Application Number
- CN202510525003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-08
- Filing Date
- 2017-10-11
- Publication Date
- 2025-07-22
AI Technical Summary
The existing VR streaming technology cannot respond quickly to viewport changes during user interaction, causing users to see a mixture of high-quality and low-quality areas in the viewport, affecting visible quality and bandwidth consumption.
By selecting and encoding media clips to stream video content in a spatially uneven manner based on predetermined relationships and signal prompts on the server side, ensuring that high-quality areas always cover the user's viewport and reducing quality in adjacent areas to reduce bandwidth consumption and computational complexity.
It improves users' visible quality in VR environment, reduces bandwidth consumption and processing complexity, and enhances the applicability and response speed of streaming.
Smart Images

Figure CN120358336A_ABST
Abstract
Description
[0001] This application is a divisional application of the applicant Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V., with an application date of October 11, 2017, an application number of 202210671217.X, and an invention title of "Spatial Uneven Streaming". Technical Field
[0002] This application relates to spatial uneven streaming such as occurs in virtual reality (VR) streaming. Background Art
[0003] VR streaming typically involves the transmission of extremely high-resolution videos. The resolution ability of the human fovea is about 60 pixels per degree. Considering the transmission of a full sphere of 360°×180°, the transmission can be completed by sending a resolution of about 22k×11k pixels. Since sending this high resolution will result in extremely high bandwidth requirements, another solution is to only send the viewport shown at the head-mounted display (HMD), which has a field of view of 90°×90°, thus generating a video of about 6k×6k pixels. The compromise between sending the full video at the highest resolution and only sending the viewport is to send the viewport at a high resolution and send some adjacent data (or the rest of the spherical video) at a lower resolution or lower quality.
[0004] In the context of DASH, omnidirectional videos (also known as spherical videos) can be provided in the way that the DASH client controls the previously described hybrid resolution or hybrid quality videos. The DASH client only needs to know the information describing how the content is provided.
[0005] An example can be to provide different representations with different projections, which have asymmetric characteristics, such as different qualities and distortions for different parts of the video. Each representation will correspond to a given viewport and will encode the viewport at a higher quality / resolution than the rest of the content. Knowing the orientation information (the direction of the viewport where the content has been encoded at a higher quality / resolution), the DASH client can dynamically select one or another representation to match the user's viewing direction at any time.
[0006] A more flexible option for the DASH client to select this asymmetric characteristic for omnidirectional videos is when the video is split into several spatial zones, where each zone can be obtained at a different resolution or quality. One option can be to split the video into rectangular zones (also known as tiles) based on a grid, but other options are foreseeable. In this case, the DASH client will need some signaling about the different qualities at which different zones are provided, and the DASH client can download different zones at different qualities so that the quality of the viewport shown to the user is better than the other unshown content.
[0007] In any of the previous scenarios, when user interaction occurs and the viewport has changed, the DASH client takes some time to react to the user's movement and download content in a way that matches the new viewport. During the time between the user's movement and the DASH client adapting its requests to match the new viewport, the user will see some areas of high quality and some areas of low quality in the viewport simultaneously. Although the acceptable quality / resolution differences are content-dependent, the quality seen by the user is reduced in any case.
[0008] Therefore, a concept that would have the effect of alleviating or more effectively manifesting or even increasing the visible quality of the user with respect to the partial rendering of spatial scene content streamed via adaptive streaming could be beneficial. Summary of the Invention
[0009] Accordingly, an object of the present invention is to provide a concept for streaming spatial scene content in a spatially non-uniform manner such that the visible quality of the user is increased, or the processing complexity or required bandwidth at the streaming extraction site is reduced, or to provide a concept for streaming spatial scene content in a way that increases the applicability to other application scenarios.
[0010] This object is achieved by the apparatus for streaming media content of a spatial scene varying over time, the streaming server for media content of a spatial scene varying over time, the media presentation description, the signal defining the media presentation description, the video bitstream, the decoder for decoding the video from the video bitstream, the streaming server for streaming media content of a spatial scene varying over time, the streaming server for allowing the apparatus to stream media content of a spatial scene varying over time from the server, the method for streaming media content of a spatial scene varying over time, the method for decoding the video from the video bitstream, the method for allowing the apparatus to stream media content of a spatial scene varying over time from the server, and the computer program having program code.
[0011] A first aspect of the present application is based on the following discovery: If the selected and extracted media segments and / or the signals obtained from the server provide the extraction device with a hint about a predetermined relationship that the quality used to encode different parts of a spatially varying scene over time adheres to, streaming media content (such as video) about a spatially varying scene over time in a spatially unequal manner can be improved in terms of visible quality and / or computational complexity at the streaming reception site under comparable bandwidth consumption. Otherwise, the extraction device may not know in advance how the juxtaposition of parts encoded with different qualities into the selected and extracted media segments negatively affects the overall visible quality experienced by the user. The information contained in the media segments and / or the signals obtained from the server (such as, within a manifest file (media presentation description) or additional streaming-related control messages from the server to the client (such as SAND messages)) enables the extraction device to make an appropriate selection among the media segments provided at the server. In this way, virtual reality streaming or partial streaming of video content can become more robust with respect to quality degradation that occurs due to insufficient distribution of available bandwidth over this spatial section of the spatially varying scene presented to the user.
[0012] On the other hand, the present invention is based on the discovery that streaming media content (such as video) of a spatially varying scene over time in a spatially non-uniform manner (such as using a first quality for a first portion and a lower second quality or leaving a second portion non-streamed in a second portion) by determining the size and / or position of the first portion depending on information contained in a media segment and / or a signal obtained from a server can result in an improvement in visible quality and / or the bandwidth consumption and / or computational complexity at the extraction side of the streaming becomes less complex. For example, for tile-based streaming, it is contemplated that a spatially varying scene over time can be provided at the server in a tile-based manner, i.e., a media segment can represent a spectral-temporal portion of the spatially varying scene over time, each of which can be a temporal segment of the spatially varying scene within a corresponding tile of the distribution of tiles into which the spatial scene is subdivided. In this case, the extraction device (client) makes a decision on how to distribute the available bandwidth and / or computational power in the spatial scene (i.e., at the tile granularity). The extraction device can perform a selection of media segments such that a first portion of the spatial scene (which respectively follows an observation section that tracks the temporal changes of the spatial scene) is encoded into the selected and extracted media segments at a predetermined quality, which can be, for example, the highest quality achievable under the current bandwidth and / or computational power conditions. For example, a second portion of the spatially adjacent spatial scene may not be encoded into the selected and extracted media segments, or may be encoded into the media segments at another quality that is reduced relative to the predetermined quality. In this case, counting the number of adjacent tiles is computationally complex or even infeasible, the aggregation of which completely covers the observation section over time, regardless of the orientation of the observation section. Depending on the projection selected to map the spatial scene onto individual tiles, the angular scene coverage of each tile can vary within this scene, and the fact that individual tiles can overlap even makes the calculation of counting adjacent tiles that are sufficient to cover the observation section in terms of space (regardless of the orientation of the observation section) more difficult. Thus, in this case, the aforementioned information can indicate the size of the first portion as a count N of tiles or the number of tiles, respectively. By this measure, the device will be able to track the observation section over time by selecting those media segments having a co-located aggregation of N tiles encoded at a predetermined quality. The fact that the aggregation of these N tiles sufficiently covers the observation section can be ensured by the information indicating N. Another example can be the information contained in the media segment and / or the signal obtained from the server, which indicates the size of the first portion relative to the size of the observation section itself. For example, this information can to some extent set a "safety zone" or prefetch zone around the actual observation section in order to account for the movement of the observation section over time. The greater the speed at which the observation section over time moves across the spatial scene, the greater the safety zone should be.Thus, the foregoing information can indicate the size of the first portion in a manner that varies with the size of the observation segment over time (such as in an incremental or scaled manner). An extraction device that sets the size of the first portion based on this information will be able to avoid quality degradation that could otherwise occur due to unextracted or low-quality portions of the spatial scene being visible in the observation segment. Here, it is irrelevant whether this scene is provided in a tile-based manner or in some other way.
[0013] Related to the just-mentioned aspect of the present application, the video bitstream encoding a video can be decoded with increased quality provided that the video bitstream has a signaling regarding the size of the focused region within the video, and the decoding capabilities for decoding the video should be concentrated on the focused region. By this measure, a decoder that decodes a video from the bitstream can concentrate or even limit its decoding capabilities for the decoded video to the portion having the size of the focused region signaled in the video bitstream, thereby knowing (for example) that the portion decoded in this way is decodable by the available decoding capabilities and spatially covers the desired segment of the video. For example, the size of such signaled focused region can be selected to be large enough to cover the size of the observation segment and the movement of this observation segment, thus taking into account the decoding latency when decoding the video. Or, in other words, the signaling of the recommended preferred observation segment region of the video contained in the video bitstream can allow the decoder to process this region in a better way, thus allowing the decoder to concentrate its decoding capabilities accordingly. Whether or not region-specific decoding capability concentration is performed, the region signaling can be forwarded to the platform where it is selected which media segments are to be downloaded, i.e., where to place the quality-increased portions and how to size the quality-increased portions.
[0014] The first and second aspects of the present application are closely related to the third aspect of the present application. According to the third aspect, the fact that a large number of extraction devices stream media content from a server is utilized to obtain information, and the information can subsequently be used to appropriately set the information of the foregoing type, thereby allowing the size or size and / or position of the first part to be set, and / or the predetermined relationship between the first quality and the second quality to be appropriately set. Therefore, according to this aspect of the present application, the extraction device (client) issues a log message that records one of the following: an instantaneous measurement result or statistical value of measuring the spatial position and / or movement of the first part; an instantaneous measurement result or statistical value of measuring the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section; and an instantaneous measurement result or statistical value of measuring the quality of the first part or the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section. The instantaneous measurement result and / or statistical value may be provided with time information related to the time when the corresponding instantaneous measurement result or statistical value was obtained. The log message may be sent to the server where the media segment is located, or to some other device that evaluates the incoming log message, in order to update the current setting of the foregoing information used to set the size or size and / or position of the first part based on the log message, and / or to derive the predetermined relationship based on the log message.
[0015] According to another aspect of the present application, in particular, by providing a media presentation description for streaming media content (such as video) of a spatially varying scene over time in a tile-based manner is more effective in avoiding useless streaming trials. The media presentation description includes: at least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene to benefit from the tile-based streaming. By this measure, the extraction device can match the beneficial requirements of the at least one version with the device capabilities of the extraction device itself or another device that interacts with the extraction device regarding tile-based streaming. For example, the benefit requirements may be related to decoding capability requirements. That is, if the decoding capability for decoding the streamed / extracted media content will not be sufficient to decode all the media segments required to cover the viewing section of the spatially varying scene over time, then attempting to stream and present the media content will waste time, bandwidth, and computing power, and thus, it may be more effective not to attempt to stream and present the media content in any case. For example, if (for example) the media segments related to a particular tile form a media stream (such as a video stream) separate from the media segments regarding another tile, the decoding capability requirements may, for example, indicate the number of decoder instances required for the corresponding version. For example, the decoding capability requirements may also be regarding other information, such as a particular portion of the decoder instances necessary for a predetermined decoding profile and / or level, or may indicate a particular minimum capability of the user input device to move the viewport / section for viewing the scene fast enough. Depending on the scene content, low mobility may not be sufficient for the user to view the portion of the scene of interest.
[0016] Another aspect of the present invention relates to an extension of streaming media content of a spatially varying scene over time. In particular, the idea according to this aspect is that the spatial scene can actually vary not only over time but also with respect to at least one other parameter (e.g., view and position, viewing depth, or some other physical parameter). The extraction device can use adaptive streaming in this context by: calculating the addresses of media segments that describe the spatially varying scene varying in time and in the at least one parameter depending on the viewport direction and the at least one other parameter; and extracting the media segments from the server using the calculated positions.
[0017] The aspects outlined above in the present application and the advantageous implementations that are the subject of the dependent claims can be combined individually or together. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The preferred embodiments of the present application are described below with reference to the accompanying drawings, in which:
[0019] Figure 1 A demonstration schematic diagram, which illustrates the systems of the client and the server for virtual reality applications as examples of situations where the embodiments described in the following figures can be advantageously used;
[0020] Figure 2 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment of the present application for describing possible operation modes of the client device, where the server 10 provides the device with information about acceptable or tolerable quality variations within the media content presented to the user;
[0021] Figure 3 Show Figure 2 modifications, where the portion with increased quality does not concern the part of the viewing section tracking the viewport, but rather the region of interest of the media scene content signaled from the server to the client;
[0022] Figure 4 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment, where the server provides information on how to set the size or size and / or position of the portion with increased quality, or the size or size and / or position of the actually extracted section of the media scene;
[0023] Figure 5 Show Figure 4 a variant, where the information sent by the server directly indicates the size of portion 64, rather than scaling the portion depending on the expected movement of the viewport;
[0024] Figure 6 Show Figure 4 a variant, according to which the extracted section has a predetermined quality and its size is determined by information originating from the server;
[0025] Figures 7a to 7c Show an illustration of Figure 4 and Figure 6 the manner in which information 74 causes the size of the portion extracted with a predetermined quality to increase through corresponding magnification of the size of the viewport;
[0026] Figure 8a Show an illustration of an embodiment in which the client device sends log messages to the server or a specific evaluator for evaluating these log messages in order to (for example) arrive at appropriate settings for Figures 2 to 7c the types of information discussed;
[0027] Figure 8bSchematic diagram of a tile - based cube projection of a 360 - degree scene onto tiles, and examples of how some of the tiles are covered by exemplary positions of the viewport. The small circles indicate positions in the equiangularly distributed viewports, and the shaded tiles are encoded in the downloaded segment at a higher resolution than the non - shaded tiles;
[0028] Figure 8c and Figure 8d Schematic diagram showing how the buffer fill levels (vertical axis) of different buffers of a client can evolve along a time axis (horizontal), where Figure 8c it is assumed that the buffer will be used to buffer the representation of a specific tile, while Figure 8d it is assumed that the buffer will be used to buffer the omnidirectional representation of a scene encoded into it with non - uniform quality (i.e., increasing in a certain direction specific to the corresponding buffer);
[0029] Figure 8e and Figure 8f Three - dimensional diagram showing different pixel density measurements within the viewport 28, differing in terms of uniformity in the sense of a sphere or an observation plane;
[0030] Figure 9 Block diagram of a client device and a schematic illustration of a media segment selection procedure when the device detects information from a server to evaluate whether a specific version of tile - based streaming provided by the server is acceptable for the client device;
[0031] Figure 10 Schematic diagram showing a plurality of media segments provided by a server according to an embodiment, allowing the media scene to depend not only on time but also on another non - temporal parameter (i.e., here illustratively the scene center position);
[0032] Figure 11 Schematic diagram showing a video bitstream containing information for manipulating or controlling the size of a focus region within a video encoded into the bitstream, and an example of a video decoder capable of utilizing this information. Detailed Description
[0033] For ease of understanding the description of the embodiments of the present application with respect to various aspects of the present application, Figure 1 an example of an environment in which the embodiments described subsequently of the present application can be applied and advantageously used is shown. In particular, Figure 1Disclosed is a system consisting of a client 10 and a server 20 that interact via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) can be used for the communication 22 between the client 10 and the server 20. However, the embodiments outlined subsequently should not be construed as being limited to the use of DASH, and likewise, terms such as Media Presentation Description (MPD) should be understood broadly so as to also cover manifest files that are different from those in DASH.
[0034] Figure 1 A system configured to implement a virtual reality application is described. That is, the system is configured to present to a user wearing a head-up display 24 (i.e., via an internal display 26 of the head-up display 24) an observation segment 28 of a spatially varying scene 30 that changes over time, the segment 28 corresponding to the orientation of the head-up display 24 exemplarily measured by an internal orientation sensor 32 (such as an inertial sensor of the head-up display 24). That is, the segment 28 presented to the user forms a segment of the spatially varying scene 30, the spatial position of which corresponds to the orientation of the head-up display 24. In Figure 1 the case where, the spatially varying scene 30 that changes over time is depicted as an omnidirectional video or a spherical video, but Figure 1 the description and the embodiments explained subsequently can also be easily transferred to other examples, such as presenting a segment in a video, where the spatial position of the segment 28 is determined by the intersection of face access or eye access with a virtual or real projection wall or the like. Additionally, the sensor 32 and the display 26 can be included by different devices (such as a remote control and a corresponding television), respectively, or the sensor and the display can be part of a handheld device (such as a mobile device, such as a tablet computer or a mobile phone). Finally, it should be noted that some of the embodiments described later can also be applied to the situation where the area 28 presented to the user always covers the entire spatially varying scene 30 that changes over time, where the non-uniformity during the presentation of the spatially varying scene is related to, for example, an unequal distribution of quality in the spatial scene.
[0035] Additional details regarding the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are described in Figure 1 and are described below. However, these details should not be regarded as limiting the embodiments explained subsequently, but should actually serve as examples of how to implement any of the embodiments explained subsequently.
[0036] In particular, as Figure 1As shown, server 20 may include a memory 34 and a controller 36, such as a suitably programmed computer, an application specific integrated circuit, etc. The memory 34 has media segments stored thereon, and the media segments represent a spatially varying scene 30 that changes over time. Specific examples will be outlined in more detail below with respect to Figure 1 . The controller 36 answers requests sent by the client 10 by re - sending the requested media segments to the client 10, and the media presentation description may send information about itself to the client 10. Details regarding this are also stated below. The controller 36 may extract the requested media segments from the memory 34. Other information may also be stored in this memory, such as the media presentation description or parts thereof, which are sent from the server 20 to the client 10 in other signals.
[0037] As Figure 1 shown, the server 20 may optionally further include a stream modifier 38 that modifies the media segments sent from the server 20 to the client 10 in response to a request from the client 10, so as to produce a media data stream at the client 10 that forms a single media stream decodable by an associated decoder, but the media segments extracted by the client 10 in this way are actually aggregated from several media streams. However, the presence of this stream modifier 38 is optional.
[0038] Figure 1 The client 10 is illustratively depicted as including a client device or controller 40 and one or more decoders 42 and a reprojection unit 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device (such as an FPGA or an application specific integrated circuit), etc. The client device 40 is responsible for selecting the segments to be extracted from the server 20 from among a plurality of 46 media segments provided at the server 20. For this purpose, the client device 40 first extracts a manifest or a media presentation description from the server 20. From the manifest or the media presentation description, the client device 40 obtains the calculation rules for calculating the addresses of the media segments corresponding to a specific desired spatial portion of the spatial scene 30 among the plurality of 46 media segments. The client device 40 extracts the thus - selected media segments from the server 20 by sending corresponding requests to the server 20. These requests contain the calculated addresses.
[0039] The media segments extracted by the client device 40 in this way will be forwarded by the client device 40 to one or more decoders 42 for decoding. In Figure 1In the example, the media segments thus extracted and decoded represent only the spatial section 48 in the spatial scene 30 that changes over time for each time unit, but as already indicated above, this may be different depending on, for example, the observation section 28 to be presented that always covers other aspects of the entire scene. The reprojection device 44 may optionally reproject the observation section 28 to be displayed to the user and cut out the observation section from the extracted and decoded scene content of the selected, extracted, and decoded media segments. For this purpose, as Figure 1 shown in, the client device 40 may, for example, continuously track the spatial position of the observation section 28 and update the spatial position in response to user orientation data from the sensor 32, and notify the reprojection device 44, for example, of this current spatial position of the scene section 28 and the reprojection mapping to be applied to the extracted and decoded media content so as to be mapped to the area forming the observation section 28. The reprojection device 44 may accordingly apply the mapping and interpolation to, for example, a regular grid of pixels to be displayed on the display 26.
[0040] Figure 1 Illustrates the case where the spatial scene 30 has been mapped to the tiles 50 using cube mapping. The tiles are thus depicted as rectangular sub-regions of a cube onto which the scene 30 in the form of a sphere has been projected. The reprojection device 44 reverses this projection. However, other examples may also be applied. For example, instead of cube projection, a projection onto a truncated cone or a non-truncated cone may be used. Furthermore, although Figure 1 the tiles are depicted as non-overlapping with respect to covering the spatial scene 30, the subdivision into tiles may involve mutual tile overlap. And as will be outlined in more detail below, it is also not mandatory for the scene 30 to be spatially subdivided into tiles 50 (as will be further explained below, each tile forms a representation).
[0041] Therefore, as Figure 1 depicted in, the entire spatial scene 30 is spatially subdivided into tiles 50. In Figure 1 the example, each of the six faces of the cube is subdivided into 4 tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52, as Figure 1 depicted in. To be more precise, the server 20 provides more than one video 52 for each tile 50, and these videos have different qualities Q#. Even further, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# respectively form or are encoded into one of the media segments of a plurality of 46 media segments stored in the memory 34 of the server 20.
[0042] Even to emphasize again, Figure 1The examples of tile-based streamification described herein are merely examples that may deviate in many ways from it. For example, although Figure 1 may seem to indicate that a media segment of a higher-quality representation of scene 30 is for a tile that is consistent with the tile to which the media segment belongs, and the tile encodes scene 30 at quality Q1 therein, this consistency is not required and tiles of different qualities may even correspond to tiles of different projections of scene 30. Additionally, although not discussed so far, it is possible that Figure 1 the media segments corresponding to different quality levels depicted in
[0043] differ in terms of spatial resolution and / or signal-to-noise ratio and / or temporal resolution, etc.
[0044] Finally, different from the tile-based streamification concept (according to which the media segments that can be individually extracted by device 40 from server 20 are spatially subdivided into tiles 50 of scene 30), the media segments provided at server 20 may alternatively (for example) each encode scene 30 therein in a spatially complete manner with a spatially varying sampling resolution, however, the sampling resolution reaches a maximum at different spatial positions in scene 30. For example, this situation can be achieved by providing at server 20 a sequence of segments 54 related to the projection of scene 30 onto a truncated cone, and the frustum of the truncated cone can be oriented in mutually different directions, resulting in resolution peaks in different orientations.
[0045] After the systems of server 20 and client 10 have been more generally explained, the functionality of client device 40 will be described in more detail with respect to an embodiment according to the first aspect of the present application. For this purpose, refer to Figure 2 which shows device 40 in more detail. As explained above, device 40 is used for streamifying media content of a spatial scene 30 that changes over time. As regarding Figure 1As explained, the apparatus 40 may be configured such that the streamed media content is spatially continuous with respect to the entire scene, or only with respect to a section 28 of the scene. In any case, the apparatus 40 includes: a selector 56 for selecting an appropriate media segment 58 from a plurality of 46 media segments available on the server 20; and an extractor 60 for extracting the selected media segment from the server 20 by a corresponding request (such as an HTTP request). As described above, the selector 56 may use the media presentation description to calculate the addresses of the selected media segments, and the extractor 60 uses these addresses when extracting the selected media segment 58. For example, the calculation rules indicated in the media presentation description for calculating the addresses may depend on the quality parameter Q, the tile T, and a certain time segment t. For example, the address may be a URL.
[0046] As also discussed above, the selector 56 is configured to perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene encoded therein. The spatial section may cover the entire scene continuously in space. Figure 2 An illustrative case where the apparatus 40 adapts a spatial section 62 of the scene 30 to overlap and surround the viewing section 28 is illustrated at 61. However, as mentioned above, this need not be the case, and the spatial section may cover the entire scene 30 continuously.
[0047] In addition, the selector 56 performs the selection such that the selected media segment has a section 62 encoded therein in a spatially non-uniform quality manner. More precisely, a first part 64 of the spatial section 62 (at Figure 2The middle part (indicated by the shaded line) is encoded into the selected media segment with a predetermined quality. This quality can be, for example, the highest quality provided by server 20 or can be a "good" quality. For example, device 42 moves or adapts the first part 64 in a manner that spatially follows the observation segment 28 over time. For example, selector 56 selects the current time segment 54 of those tiles that inherit the current position of the observation segment 28. After such selection, as will be explained below with respect to other embodiments, selector 56 can optionally keep the number of tiles that make up the first part 64 constant. In any case, the second part 66 of segment 62 is encoded into the selected media segment 58 with another quality, such as a lower quality. For example, selector 56 selects a media segment corresponding to the current time segment of tiles that are spatially adjacent to part 64 and belong to the lower quality tiles. For example, to address the possible moment when the observation segment 28 moves too fast to leave part 64 before the end of the time interval corresponding to the current time segment and overlaps with part 66 and selector 56 will be able to reconfigure part 64 spatially, selector 56 mainly selects the media segment corresponding to part 66. In this case, the part of segment 28 that protrudes into part 66 can still be presented to the user (i.e., with reduced quality).
[0048] Device 40 is not able to evaluate for the user which negative quality degradation can be caused by presenting the content of the reduced quality scenario and the content of the scenario within the higher quality part 64 to the user in advance. In particular, a transition between these two qualities that can be clearly visible to the user is generated. At least, this transition depends on the current scenario content within segment 28 being visible. The severity of the negative impact of this transition within the user's field of view is a characteristic of the scenario content provided by server 20 and may not be predicted by device 40.
[0049] Therefore, according to Figure 2In an embodiment, the apparatus 40 includes a derivator 66 that derives a predetermined relationship that will be satisfied between the quality of portion 64 and the quality of portion 66. The derivator 66 derives this predetermined relationship from information that may be included in a media segment (such as within a delivery block in media segment 58) and / or included in a signaling from the server 20 (such as in a media presentation description or an exclusive signal sent from the server 20 (such as within a SAND message, etc.)). Examples of what the information 68 may look like are presented below. The predetermined relationship 70 derived by the derivator 66 based on the information 68 is used by the selector 56 to perform the selection appropriately. For example, compared to a completely independent selection of the quality of portions 64 and 66, the constraints in selecting the quality of portions 64 and 66 affect the distribution of the available bandwidth for extracting the media content of section 62 onto portions 64 and 66. In any case, the selector 56 selects a media segment such that the quality used to encode portions 64 and 66 into the last extracted media segment satisfies the predetermined relationship. Examples of what the predetermined relationship may look like are also presented below.
[0050] The media segment selected and finally extracted by the extractor 60 is finally forwarded to one or more decoders 42 for decoding.
[0051] For example, according to a first example, the signaling mechanism embodied by the information 68 relates to information 68 that indicates to the apparatus 40 (which may be a DASH client) which quality combinations are acceptable for the provided video content. For example, the information 68 may be a list of quality pairs that indicates to the user or the apparatus 40 the maximum quality (or resolution) difference by which different sections 64 and 66 can be mixed. The apparatus 40 may be configured to inevitably use a particular quality level (such as the highest quality level provided at the server 10) for portion 64 and derive the quality level used to encode portion 66 into the selected media segment from the information 68, where the information is included in the form of (for example) a list of quality levels for portion 68.
[0052] The information 68 may indicate a tolerable value for a measure of the difference between the quality of portion 68 and the quality of portion 64. The quality index of the media segment 58 may be used as a "measure" of the quality difference, and media segments may be distinguished in the media presentation description by the quality index, and the address of the media segment is calculated using the calculation rules described in the media presentation description by the quality index. In MPEG-DASH, the corresponding attribute indicating quality may be (for example) @qualityRanking. The apparatus 40 may consider the constraints in the available quality level pairs that can be used to encode portions 64 and 66 into the selected media segment when performing the selection.
[0053] However, instead of this difference metric, the quality difference may alternatively be measured (for example) in terms of bitrate difference (i.e., the tolerable difference in the bitrates used to encode portions 64 and 66 into the corresponding media segments), assuming that the bitrate generally increases monotonically with quality. Information 68 may indicate the allowed pairings of options for the quality used to encode portions 64 and 66 into the selected media segments. Alternatively, information 68 simply indicates the allowed quality for encoding portion 66, thereby indirectly indicating the allowed or tolerable quality difference, assuming that the main portion 64 is encoded using some default quality (such as the highest quality that may be available). For example, information 68 may be a list of acceptable representation IDs or may indicate a minimum bitrate level for the media segment related to portion 66.
[0054] However, alternatively, a more gradual quality difference may be required, where instead of quality pairs, quality groups (more than two qualities) may be indicated, and depending on the distance from section 28 (viewport), the quality difference may increase. That is, in a manner depending on the distance from the viewing section 28, information 68 may indicate the tolerable values of the metric of the difference in quality between portions 64 and 66. This may be done through a list of pairs of the corresponding distances from the viewing section and the corresponding tolerable values of the metric of the quality difference for distances beyond the corresponding distance. Below the corresponding distance, the quality difference must be lower. That is, each pair may indicate for the corresponding distance that a portion of portion 66 that is farther from section 28 than the corresponding distance may have a quality difference from the quality of portion 64 that exceeds the corresponding tolerable value of this list entry.
[0055] The tolerable value may increase with the distance from the viewing section 28. The acceptability of the quality difference just discussed often depends on the time for which these different qualities are presented to the user. For example, content with a high quality difference may be acceptable if the content is presented for only 200 microseconds, while content with a lower quality difference may be acceptable if the content is presented for 500 microseconds. Thus, according to another example, in addition to, for example, the aforementioned quality combinations, or in addition to the allowed quality differences, information 68 may also include a time interval during which the combination / quality difference may be acceptable. In other words, information 68 may indicate the tolerable or maximum allowed difference in quality between portions 66 and 64, and an indication of the maximum allowed time interval during which portion 66 may be presented simultaneously with portion 64 within the viewing section 28.
[0056] As previously mentioned, the acceptability of quality differences depends on the content itself. For example, the spatial position of different tiles 50 affects acceptability. Quality differences in a uniform background region with low-frequency signals are expected to be more acceptable than those in foreground objects. In addition, due to changes in the content, the temporal position also affects acceptability. Thus, according to another example, the signal forming information 68 is sent intermittently (such as, every representation or period in DASH) to the device 40. That is, the predetermined relationship indicated by information 68 can be updated intermittently. Additionally and / or alternatively, the signaling mechanism implemented by information 68 can vary in space. That is, information 68 can be made spatially dependent, such as through the SRD parameter in DASH. That is, for different spatial regions of the scene 30, different predetermined relationships can be indicated by information 68.
[0057] As regarding Figure 2 described, an embodiment of the device 40 relates to the fact that the device 40 desires to keep the quality degradation caused by the pre-fetched part 66 within the extracted segment 62 of the video content 30 that is briefly visible in the segment 28 as low as possible before being able to change the positions of the segment 62 and the part 64 so as to adapt the segment and the part to the position change caused by the segment 28. That is, in Figure 2 the parts 64 and 66 are different parts of the segment 62, the quality of which is restricted until their possible combination is of interest to the information 68, and the transition between the two parts 64 and 66 is continuously shifted or adjusted so as to track or overtake the moving viewing segment 28. According to Figure 3 the alternative embodiment shown in, the device 40 uses the information 68 to control the possible combination of the quality of the parts 64 and 66, however, according to Figure 3 the embodiment of, the parts 64 and 66 are defined as being different or distinct from each other in a manner defined, for example, in the media presentation description (i.e., in a manner independent of the position of the viewing segment 28). The positions of the parts 64 and 66 and the transition therebetween can be constant or vary in time. If they vary in time, such variations are due to changes in the content of the scene 30. For example, the part 64 can correspond to the region of interest that is worthy of consuming higher quality, while the part 66 is the part for which quality degradation due to, for example, low bandwidth conditions should be considered before considering the quality degradation of the part 64.
[0058] In the following, another embodiment of an advantageous implementation of the device 40 is described. In particular, Figure 4 a device 40 is shown, which is structurally corresponding to Figure 2 and 3 but the operating mode is changed so as to correspond to the second aspect of the present application.
[0059] That is, the apparatus 40 includes a selector 56, an extractor 60, and a derivator 66. The selector 56 selects from among a plurality of media segments 58 provided by the server 20, and the extractor 60 extracts the selected media segment from the server. Figure 4 Assume that the apparatus 40 operates as depicted and described with respect to Figure 2 and Figure 3 i.e., the selector 56 performs the selection such that the selected media segment 58 encodes a spatial segment 62 of the scene 30 in such a way that the spatial segment follows an observation segment 28 whose spatial position varies in time. However, a variant corresponding to the same aspect of the present application will subsequently be described with respect to Figure 5 wherein, for each time instant t, the selected and extracted media segment 58 encodes the entire scene or a constant spatial segment 62 therein.
[0060] In any case, similar to the description with respect to Figure 2 and 3 the selector 56 selects the media segment 58 such that a first portion 64 within the segment 62 is encoded into the selected and extracted media segment at a predetermined quality, while a second portion 66 of the segment 62 (which is spatially adjacent to the first portion 64) is encoded into the selected media segment at a quality reduced with respect to the quality of the portion 64. A variant is described in Figure 6 where the selector 56 restricts the selection and extraction of media segments with respect to a moving template to tracking the position of the viewport 28, and wherein the media segment has a segment 62 encoded therein at a predetermined quality such that the first portion 64 completely covers the segment 62 while being surrounded by an uncoded portion 72. In any case, the selector 56 performs the selection such that the first portion 64 follows the observation segment 28 whose spatial position varies in time.
[0061] In this case, it is also not easy for the client 40 to predict how large the segment 62 or the portion 64 should be. Depending on the scene content, most users may perform similar actions when moving the observation segment 28 across the scene 30, and thus, the same actions apply to the intervals of the observation segment 28, which is likely to move across the scene 30 at approximately this speed. Therefore, according to the embodiment of Figures 4 to 6 information 74 is provided by the server 20 to the apparatus 40 to help the apparatus 40 set the size or size and / or position of the first portion 64, or the size or size and / or position of the segment 62, respectively, depending on the information 74. Regarding the possibility of transmitting the information 74 from the server 20 to the apparatus 40, as described above with respect to Figure 2 and Figure 3The described situation applies. That is, the information can be included within media segment 58, such as within the event chunk of the media segment, or for this purpose, a media presentation description or a transmission within an exclusive message (such as a SAND message) sent from the server to the device 40 can be used.
[0062] That is, according to Figures 4 to 6 the embodiment of, the selector 56 is configured to set the size of the first portion 64 depending on the information 74 originating from the server 20. In Figures 4 to 6 the embodiment illustrated, the size is set in units of tiles 50, but as described above with respect to Figure 1 it may be slightly different when using another concept of providing a scene 30 with spatially varying quality at the server 20.
[0063] According to an example, the information 70 may (for example) include the probability of a given movement speed of the viewport of the viewing section 28. As already indicated above, the information 74 may result in a media presentation description being available for the client device 40 (which may be, for example, a DASH client), or some in-band mechanism may be used to convey the information 74, such as an event chunk, that is, an EMSG or a SAND message in the case of DASH. The information 74 may also be included in any container format, such as the ISO file format or a transport format beyond MPEG-DASH (such as MPEG-2TS). The information may also be conveyed in the video bitstream (such as, in an SEI message as described later). In other words, the information 74 may indicate a predetermined value for the measure of the spatial speed of the viewing section 28. In this way, the information 74 indicates the size of the portion 64, either in the form of a scaling relative to the size of the viewing section 28 or in the form of an increment relative to the size of the viewing section 28. That is, the information 74 starts from the "basic size" of the portion 64 necessary to cover the size of the section 28 and appropriately (such as incrementally or proportionally) increases this "basic size". For example, the aforementioned movement speed of the viewing section 28 can be used to correspondingly scale the perimeter of the current position of the viewing section 28 in order to determine (for example) the furthest position of the perimeter of the viewing section 28 in any spatial direction that is feasible after this time interval, for example, to determine the delay when adjusting the spatial position of the portion 64, such as the duration of the time segment 54 corresponding to the time length of the media segment 58. The speed multiplied by this duration plus the perimeter of the current position of the omnidirectional viewport 28 can thus result in this worst-case perimeter and can be used to determine the magnification of the portion 64 relative to a certain minimum expansion of the portion 64 assuming a non-moving viewport 28.
[0064] The information 74 can even be about the evaluation of statistical data on user behavior. Subsequently, embodiments suitable for feeding such an evaluation program are described. For example, the information 74 can indicate the maximum speed for a certain percentage of users. For example, the information 74 can indicate that 90% of the users move at a speed below 0.2 radians per second and 98% move at a speed below 0.5 radians per second. The information 74 or the message carrying the information can be defined such that a probability-speed pair is defined or the message can be defined to signal the maximum speed of a fixed percentage of users (e.g., always 99% of the users). The movement speed signaling 74 can additionally include direction information, i.e., an angle in 2D, or depth in 2D plus 3D (also referred to as light field applications). The information 74 can indicate different probability-speed pairs for different movement directions.
[0065] In other words, the information 74 can be applied to a given time span, such as the time length of a media segment. The information can consist of a trajectory-based (x percentage, average user path) or speed-based pair (x percentage, speed) or distance-based pair (x percentage, pore / diameter / preferred) or area-based pair (x percentage, recommended preferred area) or a single maximum boundary value of a path, speed, distance, or preferably area. Instead of associating the information with a percentage, a simple frequency grading can be made according to the fact that most users move at a specific speed, the second most users move at another speed, etc. Additionally or alternatively, the information 74 is not limited to indicating the speed of the observation segment 28, but can equally indicate the preferred areas to be observed separately, in order to guide the attempt to track parts 62 and / or 64 of the observation segment 28, with or without an indication of the statistical significance of the indication (such as the percentage of users who have complied with that indication or an indication of whether the indication is consistent with the most frequently recorded user observation speed / observation segment), and with or without an indication of the time duration of the indication. The information 74 can indicate another measure of the speed of the observation segment 28, such as a measure of the travel distance of the observation segment 28 over a specific time period (such as within the time length of a media clip, or more specifically within the time length of the time segment 54). Alternatively, the information 74 can be notified in a way that differentiates between specific movement directions in which the observation segment 28 can travel. This applies both to indicating the rate or speed of the observation segment 28 in a specific direction and to indicating the travel distance of the observation segment 28 with respect to a specific movement direction. Additionally, the extension of part 64 can be signaled directly by the information 74 omnidirectionally or in a way that differentiates between different movement directions. Additionally, all of the examples outlined above can be modified, where the information 74 indicates these values as well as the percentage of users for whom these values are sufficient to explain the statistical behavior when moving the observation segment 28. In this regard, it should be noted that the observation speed (i.e., the speed of the observation segment 28) can be quite large and is not limited to (for example) the speed value of the user's head. In fact, the observation segment 28 can move depending on (for example) the eye movement of the user, in which case the observation speed can be significantly larger. The observation segment 28 can also move according to the movement of another input device (such as according to the movement of a tablet computer, etc.). Since all of these "input possibilities" that enable the user to move the segment 28 result in different expected speeds of the observation segment 28, the information 74 can even be designed such that the information differentiates between different concepts for controlling the movement of the observation segment 28. That is, the information 74 can indicate the size of part 64 in a way that indicates different sizes for different methods of controlling the movement of the observation segment 28, and the device 40 can use the size indicated by the information 74 for correct observation segment control.That is, the device 40 obtains knowledge about the manner in which the viewing section 28 is controlled by the user, that is, checks whether the viewing section 28 is controlled by head movement, eye movement, or tablet computer movement or similar movement, and sets the size according to a portion of the information 74 corresponding to such viewing section control.
[0066] Generally, the movement speed can be signaled according to content, time period, representation, segment, according to the SRD position, according to pixels, according to tiles (e.g., at any temporal or spatial granularity, etc.). The movement speed can also distinguish head movement and / or eye movement, as outlined just above. Additionally, the information 74 about the user movement probability can be conveyed as a recommendation regarding high-resolution prefetching (i.e., the video region outside the user's viewport, or the sphere coverage).
[0067] Figures 7a to 7c Briefly outline some of the options as explained regarding the information 74 in terms of its use by the device 40 to respectively modify the size and / or the position of the portion 64 or the portion 62. According to Figure 7a the option shown in, the device 40 magnifies the perimeter of the section 28 by a distance corresponding to the product of the signaled speed v and the duration Δt, where the duration can correspond to a time period corresponding to the time length of the time segment 54 encoded in the individual media segment 50a. Additionally and / or alternatively, the greater the speed, the further the position of the portion 62 and / or 64 can be placed away from the current position of the section 28, or the current position of the portion 62 and / or 64 can be placed in the direction of the signaled speed or movement, as signaled by the information 74. The speed and direction can be derived from the recent developments or changes in measuring or extrapolating the recommended preferred regions indicated by the information 74. Instead of applying v×Δt omnidirectionally, the speed can be signaled differently by the information 74 for different spatial directions. Figure 7b The alternative shown in depicts that the information 74 can directly indicate the distance to magnify the perimeter of the viewing section 28, which is indicated by the parameter s in Figure 7b . Again, magnification with a change in the direction of the section can be applied. Figure 7c shows that the magnification of the perimeter of the section 28 can be indicated by an area increase by the information 74, such as in the form of a ratio of the area of the magnified section to the original area of the section 28. In any case, the perimeter of the region 28 after magnification (indicated by 76 in Figures 7a to 7c ) can be used by the selector 56 to size or set the size of the portion 64 such that the portion 64 covers at least the entire region within the magnified section 76 by a predetermined amount. Obviously, the larger the section 76, for example, the larger the number of tiles within the portion 64. According to another alternative, the section 74 can directly indicate the size of the portion 64, such as in the form of the number of tiles constituting the portion 64.
[0068] In Figure 5A further possibility of signaling the size of the signaling part 64 is depicted. Figure 5 The embodiments of Figure 4 can be modified in a manner similar to the embodiments of Figure 6 i.e., the entire area of the section 62 can be retrieved from the server 20 in terms of the quality of the part 64 through the segment 58.
[0069] In any case, at the end of Figure 5 the information 74 differentiates between different sizes of the viewing section 28, i.e., different fields of view seen by the viewing section 28. The information 74 simply indicates the size of the part 64 depending on the size of the viewing section 28 currently targeted by the device 40. This enables the service of the server 20 to be used by devices having different fields of view or different sizes of the viewing section 28 without a device such as the device 40 having to cope with calculating or otherwise guessing the size of the part 64 such that the part 64 is sufficient to cover the viewing section 28 regardless of any movement of the section 28 (as discussed with respect to Figure 4 Figure 6 and FIG. 7). As will become clear from the description of Figure 1 it is easy to evaluate, for example, which constant number of tiles may be sufficient to completely cover a particular size of the viewing section 28 (i.e., a particular field of view) regardless of the orientation of the viewing section 28 for spatial positioning 30. Here, the information 74 alleviates this situation and the device 40 can simply look up the value of the size of the part 64 within the information 74 for the size of the viewing section 28 of the device 40. That is, according to the embodiments of Figure 5 a media presentation description (such as an event chunk or a SAND message) available for use by a DASH client or some interested agency may include the information 74 regarding the sphere coverage or the field of view of a set of representations or a set of tiles respectively. One example could be providing M representations of tiles as depicted in Figure 1 The information 74 may indicate a recommended number n < M tiles (referred to as representations) to be downloaded for covering a given terminal device field of view. For example, in a cube representation tiled into 6×4 tiles as depicted in Figure 1 it is considered that 12 tiles are sufficient to cover a 90°×90° field of view. Due to the fact that the terminal device field of view may not always align perfectly with the tile boundaries, this recommendation cannot be trivially generated by the device 40 itself. The device 40 may use the information 74 by downloading, for example, at least N tiles, i.e., the media segment 58 is related to N tiles. Another way of using the information could be to focus on the quality of the N tiles closest to the current viewing center of the terminal device within the section 62, i.e., using the N tiles to form the part 64 of the section 62.
[0070] Regarding Figure 8a embodiments of another aspect of the present application are described. Here,Figure 8a The client device 10 and the server 20 are shown, and the two are based on the above Figure 1 7 to communicate with each other. That is, the devices 10 may communicate with each other according to the Figure 2 7, or may be implemented without these details as described above with respect to Figure 1 However, the device 10 is advantageously configured according to the above description. Figure 2 7 or any combination thereof, and further inherits the present Figure 8a In particular, the device 10 is understood internally as described above with respect to Figure 2 8, that is, the device 40 includes a selector 56, an extractor 60 and optionally a deriver 66. The selector 56 performs the selection for targeting unequal streaming, that is, selecting the media segments in a way that the media content is encoded into the selected and extracted media segments in a way that the quality varies spatially and / or there are unencoded parts. However, in addition to this, the device 40 also includes a log message sender 80 that sends log messages recorded in (for example) the following to the server 20 or the evaluation device 82:
[0071] measuring instantaneous measurements or statistics of the spatial position and / or movement of the first portion 64,
[0072] measuring instantaneous measurements or statistics of the quality of the time-varying spatial scene up to encoding into the selected media segment and up to being visible in the observation section 28, and / or
[0073] An instantaneous measurement or statistical value of the quality of the first portion or of the temporally varying spatial scene 30 up to the encoding into the selected media segment and up to the visible observation section 28 is measured.
[0074] The motivation is as follows.
[0075] In order to be able to derive statistics, such as the most interesting regions or speed-probability pairs, a reporting mechanism from the user is required as described previously. Additional DASH metrics to the statistics defined in Annex D of ISO / IEC 23009-1 are required.
[0076] One metric may be the client's field of view which is a DASH metric, where the DASH client sends characteristics of the terminal device regarding the field of view back to a metric server (which may be the same as the DASH server or another server).
[0077] Keywords Type Description EndDeviceFoVH Integer Horizontal field of view of the terminal device, in degrees EndDeviceFoVV Integer Vertical field of view of the terminal device, in degrees
[0078] One metric can be the ViewportList, where the DASH client sends back to the metric server (which can be the same as the DASH server or another server) in a timely manner the viewports seen by each client. The instantiation of this message can be as follows.
[0079]
[0080] For the viewport (region of interest) message, the DASH client may be required to report when a viewport change occurs, possibly with a given granularity (to avoid or not avoid reporting very small movements) or a given periodicity. This message can be included in the MPD as the attribute @reportViewPortPeriodicity or as an element or descriptor. This message can also be indicated out-of-band, such as using a SAND message or any other means.
[0081] The viewport can also be signaled at the tile granularity.
[0082] Additionally or alternatively, the log message can report on other current scene-related parameters that change in response to user input, such as any of the parameters discussed below Figure 10 such as the current user distance from the scene center and / or the current viewing depth.
[0083] Another metric can be the ViewportSpeedList, where the DASH client indicates the movement speed of a given viewport when a movement occurs.
[0084]
[0085] This message can be sent only when the client performs a viewport movement. However, as in the previous case, the server can indicate that the message should be sent only when the movement is significant. This configuration can be somewhat similar to @minViewportDifferenceForReporting, for signaling the size in pixels or degrees or any other quantity that needs to be changed for the message being sent.
[0086] Another important aspect of the VR-DASH service (where the asymmetric quality as described above is provided) is to evaluate how fast a user switches from an asymmetric representation or a set of unequal quality / resolution representations of a viewport to a more adequate another representation or set of representations of another viewport. Using this metric, the server can derive statistics that help it understand the relevant factors affecting the QoE. This metric can look as follows.
[0087]
[0088]
[0089] Alternatively, the durations described previously can be given as an average value.
[0090]
[0091] For other DASH metrics, all such metrics can additionally have a measurement of the time elapsed.
[0092] t Real-time Time at which the parameter is measured
[0093] In some cases, it may occur that if unequal quality content is downloaded and poor quality (or a mixture of good and poor quality) is presented for a sufficient length of time (which may be only a few seconds), the user will be unhappy and leave the session. Under the condition of leaving the session, the user can send a message with the quality presented in the most recent x time interval.
[0094]
[0095] Alternatively, the maximum quality difference can be reported or the maximum and minimum quality of the viewport can be reported.
[0096] As is clear from the above description regarding Figure 8a it is advantageous for tile - based DASH streaming service operators to be able to derive statistics that exemplify the client reporting mechanisms described above in order to set and optimize their services in a meaningful way (e.g., regarding resolution ratio, bitrate, and segment duration). Additional DASH metrics are stated below, in addition to the metrics defined above and in Appendix D of the document "ISO / IEC 23009 - 1:2014, Information technology--Dynamic adaptive streaming over HTTP (DASH)--Part 1: Media presentation description and segment formats".
[0097] Suppose a tile - based streaming service uses a video with a cube projection as depicted in Figure 1 The reconstruction on the client side is in Figure 8bIt is described in [description], where the small circles 198 indicate the projection of the observation directions that are horizontally and vertically equiangularly distributed within the client's viewport 28 onto the two-dimensional distribution on the image area covered by the individual tiles 50. The tiles marked with hatching indicate high-resolution tiles, thus forming the high-resolution portion 64, while the tiles 50 shown without hatching represent low-resolution tiles, thus forming the low-resolution portion 66. It can be seen that as the viewport 28 changes, the user is partially presented with low-resolution tiles because the resolution of each tile on the cube determined by the most recent update of segment selection and download, where the projection plane or pixel array of the tile 50 encoded into the downloadable segment 58 falls on the cube.
[0098] Although the above description actually generally (especially) indicates feedback or log messages that indicate the quality of the video presented to the user in the viewport, hereinafter, more specific and advantageous metrics applicable in this regard will be outlined. The metrics now described can be reported back from the client side and are called the effective viewport resolution. It is speculated that this metric indicates to the service operator the effective resolution in the client's viewport. In the case where the reported effective viewport resolution indicates a resolution where the user is only presented with the resolution towards the low-resolution tiles, the service operator can accordingly change the tiling configuration, resolution ratio, or segment length to achieve a higher effective viewport resolution.
[0099] One embodiment can be the average pixel count in the viewport 28 measured in the projection plane, where the pixel array of the tile 50 encoded into the segment 58 falls in the projection plane. The measurement can be distinguished with respect to the covered field of view (FoV) of the viewport 28 or specific to the horizontal direction 204 and the vertical direction 206. The following table shows possible examples of the appropriate syntax and semantics that can be included in the log message to signal the outlined viewport quality metrics.
[0100]
[0101] The decomposition in the horizontal and vertical directions can be stopped by alternatively using a scalar value of the average pixel count. Together with an indication of the aperture or size of the viewport 28, also reportable to the recipient of the log message (i.e., the evaluator 82) is the average count that indicates the pixel density within the viewport.
[0102] It can be advantageous to reduce the field of view considered for the metric to less than the field of view of the viewport actually presented to the user, thus excluding regions that are only for peripheral vision towards the boundaries of the viewport and therefore have no impact on the subjective quality perception. This alternative is illustrated by the dashed line 202, which encloses the pixels in this central section of the viewport 28. The reporting of the considered field of view 202 for the reported metric with respect to the total field of view of the viewport 28 can also be communicated to the log message recipient 82. The following table shows the corresponding extension of the previous example.
[0103]
[0104] According to another embodiment, instead of measuring the average pixel density by spatially averaging the mass in a uniform manner in the projection plane (as in the examples actually described so far for the examples containing EffectiveFoVResolutionH / V), it is measured in a manner that weights this averaging non-uniformly for the pixels (i.e., the projection plane). The averaging can be performed in a spherically uniform manner. As an example, the averaging can be performed uniformly with respect to sample points distributed as in circle 198. In other words, the averaging can be performed by weighting the area density with weights that decrease quadratically with increasing local projection plane distance and increase according to the sine of the local tilt of the projection with respect to the line connected to the viewport. The message can include an optional (flag-controlled) step size to adjust for the inherent oversampling in some of the available projections (such as the equirectangular projection), for example by using a uniform spherical sampling grid. Some projections do not have a large oversampling problem, and forcing the calculation to remove the oversampling can create unnecessary complexity problems. This must not be limited to the equirectangular projection. The report does not need to distinguish between horizontal and vertical resolutions, but can combine them. An example is given below.
[0105]
[0106]
[0107] In Figure 8e the application of equiangular uniformity in the averaging is illustrated by showing how the points 302 (in terms of just within the viewport 28) that are equiangularly horizontally and vertically distributed on the sphere 304 centered on the viewport 306 project onto the projection plane 308 of the tile (here a cube) in order to perform the averaging of the pixel density 308 of the pixels arranged in an array (by rows and columns) in the projection area, so as to set the local weights for the pixel density according to the local density of the projection 198 of the points 302 onto the projection plane. Figure 8f A very similar method is depicted in Figure 8f . Here, the points 302 are equally spaced in the viewport plane perpendicular to the viewing direction 312, i.e., horizontally and vertically uniformly in rows and columns, and the projections onto the projection plane 308 define the points 198, and the local density of the points controls the weights, with the local density pixel density 308 varying due to the high and low resolution tiles within the viewport 28 contributing to the averaging with these weights. In the above examples such as the latest table, Figure 8f an alternative to Figure 8e the example depicted in
[0108] In the following, embodiments of another classification of log messages are described, which relate to a DASH client 10 having a plurality of media buffers 300, as Figure 8a illustrated exemplarily in, i.e., the DASH client 10 forwards the downloaded segments 58 to subsequent decoding by one or more decoders 42 (compare Figure 1 ). The distribution of the segments 58 onto the buffers can be done in different ways. For example, the distribution can be made such that specific regions of the 360 video are downloaded separately from each other, or cached after being downloaded into separate buffers. The following examples illustrate different distributions by indicating regarding the following: which tiles T indexed #1 to #24 as shown in Figure 1 are encoded to which individual downloadable representations R#1 to #P in which quality Q out of qualities #1 to #M (1 being the best and M being the worst), and how these P representations R can be grouped into adaptive sets A indexed #1 to #S in the MPD (optional), and how the segments 58 of the P representations R can be distributed onto the buffers of buffers B indexed #1 to #N.
[0109]
[0110] Here, the representations can be provided at the server and announced in the MPD for download, each of these representations regarding one tile 50, i.e., a section of the scene. Representations encoding this tile 50 with different qualities but related to one tile 50 can be summarized in optionally grouped adaptive sets, but precisely, this grouping is used to associate to the buffers. Thus, according to this example, for each tile 50, or in other words, for each viewport (observation section) encoded, there will be one buffer.
[0111] Another set of representations and distribution can be:
[0112]
[0113] According to this example, each representation can cover the entire region, but the high-quality region will be focused on one hemisphere, while the lower quality is used for the other hemisphere. Representations differing only in the exact quality used in this way (i.e., equally in the position of the higher-quality hemisphere) can be collected in one adaptive set and distributed onto the buffers (exemplarily six here) according to this characteristic.
[0114] Therefore, the following description assumes that this distribution to the buffers is applied according to different viewport encodings (video sub-regions such as tiles) associated with adaptive sets or the like. Figure 8cDescribe the buffer fullness levels over time of two separate buffers (e.g., tile 1 and tile 2) in a tile-based streaming scenario, which scenario has been described in the last but not least table. Enabling the client to report the fullness levels of all its buffers allows the service operator to correlate the data with other streaming parameters to understand the quality of experience (QoE) impact of its service settings.
[0115] Advantageously, the buffer fullness of multiple media buffers on the client side can be reported with metrics and identified and associated with buffer types. For example, the association types are as follows:
[0116] ● Tile
[0117] ● Viewport
[0118] ● Region
[0119] ● Adaptation set
[0120] ● Representation
[0121] ● Low-quality version of the full content
[0122] An embodiment of the present invention is given in Table 1, which defines metrics for reporting buffer level status events for each identified and associated buffer.
[0123] Table 1: List of buffer levels
[0124]
[0125]
[0126] Another embodiment using viewport-dependent encoding is as follows.
[0127] In a viewport-dependent streaming scenario, the DASH client downloads and pre-buffers a number of media segments related to a specific viewing orientation (viewport). If the amount of pre-buffered content is too high and the client changes its viewing orientation, the portion of the pre-buffered content that will be played after the viewport change is not presented and the corresponding media buffer is cleared. This scenario is depicted in Figure 8d .
[0128] Another embodiment can be in a traditional video streaming scenario with multiple representations (quality / bitrate) of the same content and the quality used for encoding the video content can be spatially uniform.
[0129] The distribution may thus look as follows:
[0130]
[0131] That is, here, each representation can cover, for example, a complete scene that may not be a panoramic 360 scene with different qualities (i.e., spatially uniform qualities), and these representations can be individually distributed to the buffer. All examples stated in the last three tables should be considered as not restricting the way in which the segments 58 of the representations provided at the server are distributed to the buffer. There are different methods, and the rules can be based on the membership of the segment 58 to the representation, the membership of the segment 58 to the adaptive set, the direction of the locally increased quality in which the spatial non-uniformity of the scene is encoded into the representation to which the corresponding segment belongs, the quality used to encode the scene into the corresponding segment to which it belongs, etc.
[0132] The client can maintain a buffer for each representation and, after experiencing an increase in available throughput, decide to clear the remaining low-quality / bitrate media buffer before playback and download high-quality media segments with a duration into the existing low-quality / bitrate buffer. Similar embodiments can be constructed for tile-based streaming and viewport-dependent coding.
[0133] The service operator may not be interested in understanding how much and what kind of data is downloaded without presenting, because this introduces a cost with no gain on the server side and reduces the quality on the client side. Therefore, the present invention should provide a reporting metric that correlates the two events "media download" and "media presentation" for easy interpretation. The present invention avoids analyzing the information about the download and playback status of each media segment reported at length and only allows efficient reporting of the clear event. The present invention also includes the identification of the buffer as described above and the association to the type. Embodiments of the present invention are given in Table 2.
[0134] Table 2: List of clear events
[0135]
[0136]
[0137] Figure 9 Another embodiment showing how the device 40 can be advantageously implemented Figure 9 The device 40 can correspond to any of the examples stated above with respect to Figure 1 to FIG. 8. That is, the device may include a log transmitter as discussed above with respect to Figure 8a but not necessarily, and can use the information 68 as discussed above with respect to Figure 2 and Figure 3 or the information 74 as discussed above with respect to Figures 5 to 7c but not necessarily. However, different from the description of Figure 2 to FIG. 8, with respect to Figure 9, assume that the tile-based streaming method is truly applied. That is, the scene content 30 is provided at the server 20 in a tile-based manner, and the tile-based manner is discussed as the option for FIGS. Figure 2 to 8 above.
[0138] Although the internal structure of the device 40 may be different from Figure 9 the internal structure depicted in, the device 40 is illustratively shown as including the selector 56 and the extractor 60 discussed above for FIGS. Figure 2 to 8, and optionally includes the inferencer 66. However, additionally, the device 40 includes a media presentation description analyzer 90 and a matcher 92. The MPD analyzer 90 is used to derive from the media presentation description obtained from the server 20: at least one version, for tile-based streaming, in which the spatially varying scene 30 over time is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming for the corresponding version of the spatially varying scene over time. The meaning of "version" will become clear from the following description. In particular, the matcher 92 matches the beneficial requirements thus obtained with the device capabilities of the device 40 or another device that interacts with the device 40 (such as, the decoding capabilities of one or more decoders 42, the number of decoders 42, or the like). Figure 9The background or idea implicit in the concept is as follows. Imagine, assume a specific size of the viewing section 28. Based on the tile-based approach, a specific number of tiles need to be included by the section 62. Additionally, it can be assumed that the media segments belonging to a specific tile form a media stream or video stream, and the stream should be decoded by a separate decoding instance, separately from decoding the media segments belonging to another tile. Thus, the movement aggregation of a specific number of tiles within the section 62 where the corresponding media segments are selected by the selector 56 requires a specific decoding capability, such as the existence of corresponding decoding resources in the form of, for example, a corresponding number of decoding instances (i.e., a corresponding number of decoders 42). If this number of decoders does not exist, the service provided by the server 20 is not available to the client. Thus, the MPD provided by the server 20 can indicate the "useful requirement", that is, the number of decoders required to use the provided service. However, the server 20 can provide MPDs for different versions. That is, different MPDs for different versions can be obtained by the server 20, or the MPD provided by the server 20 can be structured internally to distinguish the different versions that can be served and used. For example, the versions can differ in the field of view (i.e., the size of the viewing section 28). Different sizes of the field of view manifest themselves as different numbers of tiles within the section 62, and thus can differ in the useful requirement because, for example, these versions may require different numbers of decoders. Other examples can also be imagined. For example, although versions with different fields of view may involve the same number of media segments 46, according to another example, the differences between different versions of the scenario 30 provided for tile-based streaming at the server 20 can even lie in the number of 46 media segments involved according to the corresponding version. For example, the tile segmentation according to one version is coarser compared to the tile-hyphenation segmentation of the scenario according to another version, thus requiring, for example, a smaller number of decoders.
[0139] The matcher 92 matches the useful requirement and thus selects the corresponding version or completely rejects all versions.
[0140] However, the useful requirement can additionally focus on the profile / level that one or more decoders 42 must be able to handle. For example, the DASH MPD includes multiple locations that allow indicating the profile. A typical profile description can be the attributes, elements present in the MPD, and the video or audio profile for each representation of the provided media stream.
[0141] Other examples of beneficial requirements concern, for example, the ability to move the viewport 28 across scenarios on the client side. The beneficial requirement may indicate the required viewport speed, which should be available to the user to move the viewport and thus be able to truly enjoy the provided scenario content. The matcher may check, for example, whether this requirement is met, e.g., inserted in a user input device such as the HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in a directional sense, the set of "sufficient types of input devices" may be indicated by the beneficial requirement.
[0142] In the tile streaming service of spherical videos, there are a plethora of configuration parameters that can be set dynamically, such as the number of qualities, the number of tiles. In the case where the tiles are independent bitstreams that need to be decoded by separate decoders, if the number of tiles is too high, a hardware device with several decoders will not be able to decode all the bitstreams simultaneously. The possibility is to keep this as a degree of freedom, and the DASH device parses all possible representations and counts how many decoders are needed to decode all the representations or the given number of the field of view of the overlay device, and thus determines whether the DASH client is likely to consume the content. However, a smarter solution for interoperability and capability negotiation is to use signaling in the MPD mapped to a profile, which is used as a commitment to the client: if the profile is supported, the provided VR content can be consumed. This signaling should be in the form of a URN (such as urn::dash-mpeg::vr::2016) that can be encapsulated at the MPD level or at the adaptation set. This parsing will mean that N decoders at profile X are sufficient to consume the content. Depending on the profile, the DASH client may ignore or accept the MPD or parts (adaptation sets) of the MPD. Additionally, there are several mechanisms that do not include all the information, such as Xlink or MPD links, where little signaling for selection is available. In this case, the DASH client will not be able to determine whether it can consume the content. It is necessary to expose the decoding capabilities regarding the number of decoders and the profile / level of each decoder through this urn (or something similar) so that the DASH client can now make sense of whether to perform an Xlink or MPD link or a similar mechanism. The signaling may also imply different operating points, such as N decoders with profile X / level or Z decoders with profile Y / level.
[0143] Figure 10 Further illustration, regarding the Figures 1 to 9 Any of the above-described embodiments and descriptions presented by the client, device 40, server, etc. can be extended to the following scope: the provided service is extended to a scope where the time-varying spatial scenario changes not only over time but also depends on another parameter. For example, Figure 10 Illustration Figure 1Variant in which a plurality of available media segments are obtained on a server to describe scene content 30 for different positions of the viewing center 100. In Figure 10 In the schematic diagram shown in, the scene center is depicted as varying only along one direction X, but obviously, the viewing center can vary along more than one spatial direction (such as two-dimensionally or three-dimensionally). For example, this corresponds to a change in the user's position in a specific virtual environment. Depending on the user's position in the virtual environment, the available view changes, and accordingly, the scene 30 changes. Thus, in addition to describing that the scene 30 is subdivided into tiles and time segments and media segments of different qualities, other media segments describe different contents of the scene 30 for different positions of the scene center 100. The device 40 or the selector 56 calculates the addresses of the media segments to be extracted within the selection program from among the plurality of 46 media segments respectively depending on the viewing section position and at least one parameter (such as parameter X), and can then use the calculated addresses to extract these media segments from the server. For this purpose, the media presentation description can describe a function depending on the tile index, quality index, scene center position, and time t, and generate the corresponding addresses of the corresponding media segments. Thus, according to Figure 10 the embodiment of, the media presentation description can include this calculation rule, in addition to the parameters described above with respect to Figures 1 to 9 which, the calculation rule also depends on one or more additional parameters. Parameter X can be quantized to any level in a hierarchy, for which level the corresponding scene representation is encoded by the corresponding media segment within the plurality of 46 media segments in the server 20.
[0144] As an alternative, X can be a parameter that defines the viewing depth (i.e., the distance radially from the scene center 100). Although providing scenes in different versions with different X values in the viewing center part allows the user to "walk" through the scene, providing scenes in different versions with different viewing depths can allow the user to "radially zoom" back and forth in the scene.
[0145] For a plurality of non-concentric viewports, the MPD can thus be signaled using another signal of the position of the current viewport. The signaling can be done at the segment, representation, or period level or the like.
[0146] Non-concentric spheres: To enable user movement, the spatial relationship of different spheres should be signaled in the MPD. This signaling can be done by coordinates (x, y, z) in any unit relative to the sphere diameter. Additionally, the diameter of each sphere should be indicated. The sphere can be "good enough" for a user at its center and the additional space for which the content will be good. If the user can move beyond the signaled diameter, another sphere should be used for presenting the content.
[0147] Exemplary signaling of the viewport may be relative to a predefined center point in space. Each viewport will be signaled relative to that center point. In MPEG-DASH, this may be signaled (for example) in the AdaptationSet element.
[0148]
[0149]
[0150] Finally, Figure 11 information that describes information such as or similar to that described above with respect to reference numeral 74 may reside in the video bitstream 110 into which the video 112 is encoded. A decoder 114 that decodes this video 110 may use the information 74 to determine the size of the focus region 116 within the video 112, and the decoding capabilities for decoding the video 110 should be concentrated on the focus region. For example, the information 74 may be conveyed within the SEI information of the video bitstream 110. For example, the focus region may be decoded specifically, or the decoder 114 may be configured to start decoding each image of the video at the focus region rather than (for example) at the upper left image corner, and / or the decoder 114 may stop decoding each image of the video after the focus region 116 has been decoded. Additionally or alternatively, the information 74 may be present in the data stream for use only in forwarding to a subsequent display or viewport control or streaming media device of a client or segment selector for determining which segments to download or stream in order to cover a spatial section completely or with increased or predetermined quality. For example, as outlined above, the information 74 indicates a recommended preferred region as a recommendation for placing the viewing section 62 or section 66 to coincide with or cover or track this region. The information 74 may be used by the segment selector of the client. Just as is true with respect to the description of Figures 4 to 7c the information 74 may absolutely set the size of the region 116 (such as in terms of the number of tiles), or may set, for example, the speed of the region 116 used for region movement according to user input or the like, so as to follow the content of interest in the video in space-time, thereby scaling the region 116 to increase as the indication of the speed increases.
[0151] A first aspect of the present application provides an apparatus for streaming media content regarding a spatially varying scene (30) over time.
[0152] (A1) The apparatus is configured to:
[0153] select media segments from a plurality of media segments available on a server,
[0154] the selected media segments are from the server, wherein the apparatus is configured to:
[0155] Perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene of the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene of the time variation that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and
[0156] derive the predetermined relationship from information included in the selected media segment and / or from a signal obtained from the server.
[0157] (A2), The apparatus according to (A1), wherein each of the plurality of media segments has an associated spatio-temporal part of the spatially varying scene of the time variation encoded therein at an associated quality level within a set of quality levels.
[0158] (A3), The apparatus according to (A1), wherein each of the spatio-temporal parts of the spatially varying scene (30) encoded into the plurality of media segments is a time segment of the spatially varying scene at a corresponding one of tiles (50) into which the spatially varying scene is spatially subdivided.
[0159] (A4), The apparatus according to any one of (A1)-(A3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
[0160] (A5), The apparatus according to (A4), wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section.
[0161] (A6), The apparatus according to (A4) or (A5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section by a list of pairs of corresponding distances from the observation section and corresponding tolerable values for the measure of the difference beyond the corresponding distances.
[0162] (A7), The apparatus according to any one of (A4)-(A6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.
[0163] (A8), The apparatus according to any one of (A4)-(A7), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval during which the second part can be together with the first part within the observation section.
[0164] (A9), The apparatus according to any one of (A8), wherein the information indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval during which the second part can be together with the first part within the observation section.
[0165] (A10), The apparatus according to any one of (A4)-(A9), wherein the information is time-varying and / or spatially varying.
[0166] (A11), The apparatus according to any one of (A1)-(A10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0167] (A12), The apparatus according to any one of (A1)-(A11), wherein the apparatus is configured to perform the selection such that the first part follows a time-varying observation section of the time-varying spatial scene.
[0168] (A13), The apparatus according to (A12), wherein the apparatus is configured to cause the spatial position of the time-varying observation section to change according to a user input.
[0169] (A14), The apparatus according to any one of (A1)-(A13), wherein the apparatus is configured to determine the first part to correspond to a region of interest.
[0170] (A15), The apparatus according to any one of (A14), wherein the apparatus is configured to extract information about the region of interest from the server.
[0171] The present invention also provides a streaming media server for media content of a time-varying spatial scene.
[0172] (B1) The streaming media server is configured to:
[0173] Make multiple media segments available for extraction by a device, enabling the device to select a media segment for extraction, the media segment having at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and
[0174] Signal information about a predetermined relationship in the media segment and / or by signaling to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
[0175] (B2), The streaming media server according to (B1), wherein each of the multiple media segments has an associated spatio-temporal part of the spatially varying scene with the time variation encoded therein at an associated quality level within a set of quality levels.
[0176] (B3), The streaming media server according to (B2), wherein each of the spatio-temporal parts of the spatially varying scene encoded into the multiple media segments is a time segment of the spatially varying scene at a corresponding one of the tiles into which the spatially varying scene is spatially subdivided.
[0177] (B4), The streaming media server according to any one of (B1)-(B3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
[0178] (B5), The streaming media server according to (B4), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from an observation section.
[0179] (B6), The streaming media server according to (B4) or (B5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section by a list of pairs of the respective distances from the observation section and the corresponding tolerable values for the measures of the differences beyond the respective distances.
[0180] (B7), The streaming media server according to any one of (B4)-(B6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.
[0181] (B8), The streaming media server according to any one of (B4)-(B7), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second portion can be together with the first portion within the observation section.
[0182] (B9), The streaming media server according to (B8), wherein the information indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval within which the second portion can be together with the first portion within the observation section.
[0183] (B10), The streaming media server according to any one of (B4)-(B9), wherein the information is time-varying and / or spatially varying.
[0184] (B11), The streaming media server according to any one of (B4)-(B10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0185] (B12), The streaming media server according to any one of (B4)-(B11), wherein the server is configured to send information about the region of interest to the device.
[0186] This application also provides a media presentation description.
[0187] (C1), The media presentation description includes:
[0188] Information about calculating addresses of a plurality of media segments such that the device can use the information to select and extract media segments from the plurality of media segments, wherein at least a spatial section of the spatially varying temporal scene encoded therein is encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying temporal scene that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and
[0189] Information about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.
[0190] (C2), The media presentation description according to (C1), wherein each media segment of the plurality of media segments has an associated spatio-temporal portion of the spatially varying temporal scene encoded therein at an associated quality level within a set of quality levels.
[0191] (C3), the media presentation description according to (C2), wherein each of the spatio-temporal portions of the spatially varying scene that changes over time encoded into the plurality of media segments is a temporal segment of the spatially varying scene that changes over time at a respective one of the tiles into which the spatially varying scene that changes over time is spatially subdivided.
[0192] (C4), the media presentation description according to any one of (C1)-(C3), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.
[0193] (C5), the media presentation description according to (C4), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner that depends on the distance from the viewing section.
[0194] (C6), the media presentation description according to (C4) or (C5), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner that depends on the distance from the viewing section by a list of pairs of the respective distances from the viewing section and the corresponding tolerable values for the measure of the difference beyond the respective distances.
[0195] (C7), the media presentation description according to any one of (C4)-(C6), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
[0196] (C8), the media presentation description according to any one of (C4)-(C7), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
[0197] (C9), the media presentation description according to (C8), wherein the information indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
[0198] (C10), the media presentation description according to any one of (C4)-(C9), wherein the information is time-varying and / or spatially varying.
[0199] (C11), a media presentation description according to any one of (C4)-(C10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0200] (C12), a media presentation description according to any one of (C4)-(C11), wherein the media presentation description includes information about a region of interest.
[0201] This application also provides an apparatus for streaming media content of a spatial scene that changes over time.
[0202] (D1) The apparatus is configured to:
[0203] Select media segments from among a plurality of media segments available on a server,
[0204] wherein the apparatus is configured to:
[0205] Perform the selection such that the selected media segment has at least a spatial section of the spatial scene that changes over time encoded therein, encoded in such a way that:
[0206] A first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatial scene that changes over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and
[0207] Such that the first part follows an observation section that changes over time of the spatial scene that changes over time, and
[0208] Set the size and / or position of the first part depending on information contained in the selected media segment and / or the action of a signal obtained from the server.
[0209] (D2), the apparatus according to (D1), wherein the information indicates the size in the form of an increment relative to the size of the observation section that changes over time or a scaling of the size of the observation section that changes over time.
[0210] (D3), the apparatus according to (D1) or (D2), wherein the information indicates a predetermined value for a measure of the spatial speed of the observation section.
[0211] (D4), the apparatus according to (D3), wherein the information indicates the predetermined value for the measure of the spatial speed of the observation section for:
[0212] A default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or
[0213] A percentage value indicating the percentage of users whose spatial speed of the viewing section does not exceed the predetermined value, and / or
[0214] A percentage value indicating the percentage of users for whom the viewing section is within a predetermined area, and / or
[0215] A hint indicating one or more types of user input for controlling the movement of the viewing section to which the predetermined value is applicable.
[0216] (D5), the device according to (D3) or (D4), the device being configured to perform the setting such that
[0217] The higher the predetermined value d of the measure of the spatial speed for the viewing section, the larger the size.
[0218] (D6), the device according to any one of (D1) or (D5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement of the viewing section.
[0219] (D7), the device according to (D6), the device being configured to perform the setting such that
[0220] The higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.
[0221] (D8), the device according to any one of (D1) or (D7), wherein the information indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.
[0222] (D9), the device according to any one of (D1) or (D8), the device being configured such that the first part coincides with the spatial section.
[0223] (D10), the device according to any one of (D1) or (D9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.
[0224] (D11), the device according to (D10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level within a set of quality levels.
[0225] (D12), the apparatus according to (D1), wherein the apparatus is configured to set the size in a manner independent of the size of the observation section that varies with time depending on the information.
[0226] (D13), the apparatus according to (D1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the apparatus uses the values included in the information for the size option that fits the actual size of the time-varying observation section.
[0227] (D14), the apparatus according to (D12) or (D13), wherein the information indicates the size in terms of the number of tiles.
[0228] This application also provides a streaming media server for media content of a spatial scene that varies with time.
[0229] (E1), the streaming media server is configured to:
[0230] Make a plurality of media segments available for extraction by a device, so that the device can select the media segments for extraction, and the media segments at least have a spatial section of the spatial scene that varies with time encoded therein, and the encoding is such that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatial scene that varies with time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment with another quality that is reduced relative to the predetermined quality, and such that the first part follows the time-varying observation section of the spatial scene that varies with time, and
[0231] Signal information in the media segment and / or by signaling to the device regarding how to set the size and / or position of the first part.
[0232] This application also provides a signal defining a media presentation description.
[0233] (F1) The signal includes:
[0234] Information about calculating addresses of a plurality of media segments such that a device, using the information, can select and extract a media segment from the plurality of media segments, the media segment having at least a spatial section of the spatially varying scene that varies over time, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and such that the first part follows an observation section of the temporal variation of the spatially varying scene, and
[0235] Information in the media segment and / or via a signal to the device on how to set the size and / or position of the first part.
[0236] (F2), according to the signal of (F1), wherein the information indicates the size in the form of an increment relative to the size of the observation section of the temporal variation or a scaling of the size of the observation section of the temporal variation.
[0237] (F3), according to the signal of (F1) or (F2), wherein the information indicates a predetermined value of a measure of the spatial velocity for the observation section.
[0238] (F4), according to the signal of (F3), wherein the information indicates the predetermined value of the measure of the spatial velocity for the observation section for:
[0239] A default percentage of users for whom the measured spatial velocity does not exceed the predetermined value, and / or
[0240] A percentage value indicating the percentage of users for whom the measured spatial velocity does not exceed the predetermined value, and / or
[0241] A percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or
[0242] A hint indicating one or more types of user input for controlling the movement of the observation section to which the predetermined value is applicable.
[0243] (F5), according to the signal of (F3) or (F4), the signal being configured to perform the setting such that
[0244] wherein the higher the predetermined value d of the measure of the spatial velocity for the observation section, the larger the size.
[0245] (F6), according to the signal of any one of (F1)-(F5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement for the observation section.
[0246] (F7), according to the signal of (F6), the signal is configured to perform the setting such that
[0247] the higher the probability of the corresponding moving direction, the more the first part extends into the corresponding direction.
[0248] (F8), according to the signals of (F1)-(F7), wherein the information indicates the predetermined spatial velocity in a time-varying and / or spatially varying and / or directionally varying manner.
[0249] (F9), according to the signal of any one of (F1) or (F8), is configured such that the first part is consistent with the spatial section.
[0250] (F10), according to the signal of any one of (F1) or (F9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.
[0251] (F11), according to the signal of (F10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level in a set of quality levels.
[0252] (F12), according to the signal of (F11), wherein the device is configured to set the size in a manner independent of the size of the time-varying observation section depending on the information.
[0253] (F13), according to the signal of (F1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.
[0254] (F14), according to the signal of (F11) or (F12), wherein the information indicates the size in terms of the number of tiles.
[0255] This application also provides a video bitstream.
[0256] (G1) having video encoded therein, the video bitstream includes a signal effect for one or more of the size of the focused area within the video to which the decoding ability for decoding the video should be concentrated and the recommended preferred observation section area of the video.
[0257] The present application also provides a decoder for decoding video from a video bitstream.
[0258] (H1) The decoder is configured to:
[0259] derive a signal indicative of the size of a focused region within the video from the video bitstream, and
[0260] concentrate the decoding capability for decoding the video to the focused region.
[0261] (H2), the decoder according to (H1), the decoder is configured to specifically decode the focused region.
[0262] (H3), the decoder according to (H1), the decoder is configured to start decoding each image of the video at the focused region.
[0263] (H4), the decoder according to (H1), the decoder is configured to stop decoding each image of the video after decoding the focused region.
[0264] (H5), the decoder according to any one of (H1)-(H4), wherein the signal absolutely indicates the size, or the decoder is configured to scale the size of the focused region with a parameter included in the signal.
[0265] The present application also provides a device for streaming media content of a spatial scene that changes over time.
[0266] (I1) The device is configured to:
[0267] derive from a media presentation description:
[0268] at least one version, for tile-based streaming, the spatially changing scene over time is provided in the at least one version,
[0269] for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming of the corresponding version of the spatially changing scene over time,
[0270] match the beneficial requirements of the at least one version to the device capabilities of the device or another device that interacts with the device regarding the tile-based streaming.
[0271] (I2), the device according to (I1), wherein the beneficial requirements and the device capabilities relate to decoding capabilities.
[0272] (I3), The apparatus according to (I1) or (I2), wherein the beneficial requirements and the apparatus capabilities relate to the number of available decoders.
[0273] (I4), The apparatus according to any one of (I1)-(I3), wherein the beneficial requirements and the apparatus capabilities relate to layer and / or profile descriptors.
[0274] (I5), The apparatus according to any one of (I1)-(I4), wherein the beneficial requirements and the apparatus capabilities relate to the type of input device for moving an observation section across a spatially varying scene over time, or to the speed of moving the observation section across the spatially varying scene over time using the input device.
[0275] (I6), The apparatus according to any one of (I1)-(I5), wherein the apparatus is configured to:
[0276] Select media segments from a plurality of media segments available on the server, the selection being made by using a calculation rule included in the media presentation description to calculate the addresses of the selected media segments,
[0277] Extract the selected media segments from the server using the calculated addresses,
[0278] wherein the apparatus is configured to perform the selection such that the selected media segments have at least a spatial section of the spatially varying scene varying over time encoded therein, encoded in such a way that:
[0279] A first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene varying over time that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and
[0280] such that the first part follows the observation section of the spatially varying scene varying over time.
[0281] The present application also provides a streaming media server for streaming media content regarding a spatially varying scene varying over time.
[0282] (J1) The streaming media server is configured to:
[0283] Provide a media presentation description from which
[0284] At least one version, for tile-based streaming, the spatially varying scene varying over time is provided in the at least one version,
[0285] For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatial scene that varies over time for benefiting from the tile-based streaming,
[0286] so that the apparatus for streaming the media content from the streaming server can
[0287] match the beneficial requirements of the at least one version with the apparatus capabilities of the apparatus or another apparatus that interacts with the apparatus regarding the tile-based streaming.
[0288] This application also provides a media presentation description.
[0289] (K1) The media presentation description includes:
[0290] information about at least one version, for tile-based streaming, the spatial scene that varies over time is provided in the at least one version,
[0291] For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatial scene that varies over time for benefiting from the tile-based streaming.
[0292] (K2), the media presentation description according to (K1), wherein the beneficial requirements and the apparatus capabilities are regarding decoding capabilities.
[0293] (K3), the media presentation description according to (K1) or (K2), wherein the beneficial requirements and the apparatus capabilities are regarding the number of available decoders.
[0294] (K4), the media presentation description according to any one of (K1)-(K3), wherein the beneficial requirements and the apparatus capabilities are regarding layer and / or profile descriptors.
[0295] (K5), the media presentation description according to any one of (K1)-(K4), wherein the beneficial requirements and the apparatus capabilities are regarding the type of input device for moving the viewing section across the spatial scene that varies over time, or regarding the speed of moving the viewing section across the spatial scene that varies over time using the input device.
[0296] (K6), the media presentation description according to any one of (K1)-(K5), further includes a calculation rule, using which the apparatus can
[0297] select a media segment from among a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rule.
[0298] The present application also provides an apparatus for streaming media content of a spatial scene that changes over time.
[0299] (L1) The apparatus is configured to:
[0300] calculate an address of a media segment depending on a spatial viewport position and at least one parameter, the media segment depicting a spatial scene that changes over time and in the at least one parameter,
[0301] use the calculated address to extract the media segment.
[0302] (L2) The apparatus according to claim (L1), wherein the at least one parameter includes one or more coordinates of a center of view and / or a viewing depth.
[0303] The present application also provides a media presentation description.
[0304] (M1) The media presentation description includes:
[0305] a calculation rule for calculating an address of a media segment depending on a spatial viewport position and at least one parameter, so as to use the calculated address to extract the media segment, the media segment depicting a spatial scene that changes over time and in the at least one parameter.
[0306] The present application also provides a streaming server for allowing an apparatus to stream media content of a spatial scene that changes over time from a server.
[0307] (N1) The streaming server is configured to provide the media presentation description according to (M1).
[0308] The present application also provides an apparatus for streaming media content of a spatial scene that changes over time.
[0309] (O1) The apparatus is configured to:
[0310] select a media segment from a plurality of media segments available on a server,
[0311] wherein the apparatus is configured to:
[0312] encode a first part of the spatial scene that changes over time into the selected media segment with an increased quality compared to a spatial neighborhood of the first part or in such a way that the spatial neighborhood of the first part is not encoded into the selected media segment,
[0313] emit a log message that records:
[0314] an instantaneous measurement result of measuring a spatial position and / or movement of the first part; and / or
[0315] Measure statistical values of the spatial position and / or movement of the first part, such as time averages; and / or
[0316] Measure instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or
[0317] Indications of the set of buffers of the device involved in buffering the selected media segment, descriptions of the distribution rules applied in distributing the selected media segment into the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or
[0318] Measurement results of the amount of the selected media segment that has not been output from the buffer of the device for undergoing decoding; and / or
[0319] Measure statistical values of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time averages; and / or
[0320] Measure the instantaneous measurement results of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or
[0321] Measure statistical values of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time averages; and / or
[0322] The field of view covered by the observation section; and / or
[0323] Measure instantaneous measurement results of the user position or the observation depth relative to the scene center; and / or
[0324] Measure statistical values of the user position or the observation depth relative to the scene center, such as time averages.
[0325] (O2), The device according to (O1), wherein the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.
[0326] (O3), The device according to (O1) or (O2), configured to perform the selection such that the first part of the spatial scene of the time variation pins the observation section.
[0327] (O4), The apparatus according to any one of (O1)-(O3), configured to emit a log message that records the instantaneous measurement result of the quality of the spatial scene whose change over time is measured until encoded into the selected media segment and until visible in the observation section as one of the following:
[0328] The measurement result of the average density of the pixels falling within the observation section, where the spatial scene changing over time is encoded into the selected media segment at the average density.
[0329] (O5), The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result is obtained by averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0330] (O6), The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result is obtained by averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0331] (O7), The apparatus according to (O4), configured to make the emitted log message indicate whether the measurement result measures the average density of the pixels by
[0332] averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment, or
[0333] averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0334] (O8), The apparatus according to (O6) or (O7), wherein averaging the pixel density in a spatially non-uniform manner corresponds to
[0335] averaging in a spherically uniform manner, or
[0336] averaging uniformly in the viewport plane space, where the viewport plane is perpendicular to the central viewing direction of the observation section.
[0337] (O9), The apparatus according to any one of (O4)-(O8), configured to measure the average density of the pixels such that the measurement result is obtained by averaging the pixel density in a way that limits the averaging to the central sub-section of the observation section, or by applying a higher averaging weight to the central sub-section compared to the edge part of the observation section around the central sub-section.
[0338] (O10), The apparatus according to any one of (O4)-(O9) is configured such that the measurement result measures the average density of the pixels in a manner that is separate along the horizontal viewing section axis and the vertical viewing section axis, respectively.
[0339] (O11), The apparatus according to any one of (O4)-(O10) is configured to intermittently emit log messages.
[0340] (O12), The apparatus according to any one of (O4)-(O11) is configured to emit log messages at a rate controlled by a manifest file, and the apparatus performs the selection of the media segments for download based on the manifest file.
[0341] (O13), The apparatus according to any one of (O4)-(O12), wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the spatio-temporal scene that changes over time, and the representations differ in one or more of the following:
[0342] The scene segment of the spatio-temporal scene encoded into the media segment,
[0343] The quality used to encode the spatio-temporal scene into the media segment,
[0344] The spatial quality variation used to encode the spatio-temporal scene into the media segment,
[0345] Wherein the apparatus is configured to emit log messages, and the log messages are recorded in the form of the association of each buffer with one or a combination of two or more of the following in the description of the distribution rules applied in distributing the selected media segments to the set of buffers:
[0346] The scene segment,
[0347] The quality,
[0348] The spatial quality distribution,
[0349] Representation.
[0350] (O14), The apparatus according to any one of (O4)-(O13), wherein the representations are grouped into adaptive sets according to one or more of the following:
[0351] The scene segment of the spatio-temporal scene encoded into the media segment,
[0352] The spatial quality variation used to encode the spatio-temporal scene into the media segment,
[0353] wherein the device is configured to emit a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, either in the form of an association of each buffer with one of the adaptation sets or in the form of an association of each buffer with one of the representations.
[0354] (O15) The device according to any one of (O4)-(O14), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in the form of a time measurement result.
[0355] (O16) The device according to (O15), wherein the device is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding in the form of a time measurement result in a time unit smaller than the time length of the media segment and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.
[0356] (O17) The device according to any one of (O4)-(O16), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding in a format classified by one or more of the following:
[0357] the buffer of the decoder in which the corresponding media segment has been buffered,
[0358] the scene segment encoded into the corresponding media segment,
[0359] the quality used to encode the spatially varying scene over time into the corresponding media segment,
[0360] the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.
[0361] This application also provides a method for streaming media content regarding a spatially varying scene over time.
[0362] (P1) The method includes:
[0363] selecting a media segment from among a plurality of media segments available on a server,
[0364] Extract a selected media segment from the server, perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and
[0365] The method further includes deriving the predetermined relationship from information included in the selected media segment and / or from a signal obtained from the server.
[0366] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.
[0367] (Q1) The method includes:
[0368] Make a plurality of media segments available for extraction by a device, such that the device can select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and
[0369] Signal information regarding the predetermined relationship in the media segment and / or by signal action to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
[0370] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.
[0371] (R1) The method includes:
[0372] Select a media segment from a plurality of media segments available on a server,
[0373] Extract the selected media segment from the server, where the selection is performed such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded such that
[0374] A first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced with respect to the predetermined quality, and
[0375] An observation section that causes the first part to follow the temporal changes of the spatial scene that changes over time, and
[0376] The method further includes setting the size of the first part depending on information included in the selected media segment and / or a signal action obtained from the server.
[0377] This application also provides a method for streaming media content regarding a spatial scene that changes over time.
[0378] (S1) The method includes:
[0379] Making a plurality of media segments available for extraction by a device, so that the device can select a media segment for extraction, the selected media segment having at least a spatial section of the spatial scene that changes over time encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatial scene that changes over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and an observation section that causes the first part to follow the temporal changes of the spatial scene that changes over time, and
[0380] Signaling information in the media segment and / or via a signal action to the device regarding how to set the size of the first part.
[0381] This application also provides a method for decoding video from a video bitstream.
[0382] (T1) The method includes:
[0383] Deriving a signal action of the size of a focus area within the video from the video bitstream, and
[0384] Concentrating the decoding ability for decoding the video on the focus area.
[0385] This application also provides a method for streaming media content regarding a spatial scene that changes over time, performed by a device.
[0386] (U1) The method includes:
[0387] Deriving from a media presentation description:
[0388] At least one version, for tile-based streaming, the spatial scene that changes over time is provided in the at least one version,
[0389] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatial scene that benefits from the time variation for tile-based streaming
[0390] Match the beneficial requirements of the at least one version with the device capabilities of the device or another device that interacts with the device regarding tile-based streaming.
[0391] This application also provides a method for streaming media content of a spatial scene that varies over time.
[0392] (V1) The method includes:
[0393] Select media segments from a plurality of media segments available on a server
[0394] Extract the selected media segments from the server, where the selection is performed such that the selected media segments have an increased quality compared to the spatial neighborhood of the first part or are encoded therein in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segments, for the first part of the spatial scene that varies over time
[0395] The method further includes emitting a log message that records:
[0396] An instantaneous measurement of the spatial position and / or movement of the first part; and / or
[0397] A statistical value of the spatial position and / or movement of the first part, such as a time average; and / or
[0398] An instantaneous measurement of the quality of the spatial scene that varies over time until it is encoded into the selected media segments and until it is visible in the observation section; and / or
[0399] An indication of the set of buffers of the device involved in buffering the selected media segments, a description of the distribution rules applied in distributing the selected media segments to the set of buffers, and the instantaneous buffer fill level of each of the set of buffers; and / or
[0400] A measurement of the amount of the selected media segments yet to be output from the device's buffer for undergoing decoding; and / or
[0401] A statistical value of the quality of the spatial scene that varies over time until it is encoded into the selected media segments and until it is visible in the observation section, such as a time average; and / or
[0402] An instantaneous measurement of the quality of the first part or of the spatial scene whose quality varies over time until it is encoded into the selected media segment and until it is visible in the viewing segment; and / or
[0403] A statistical value, such as a time average, of the quality of the first part or of the spatial scene whose quality varies over time until it is encoded into the selected media segment and until it is visible in the viewing segment; and / or
[0404] The field of view covered by the viewing segment; and / or
[0405] An instantaneous measurement of the user's position relative to the center of the scene or of the viewing depth; and / or
[0406] A statistical value, such as a time average, of the user's position relative to the center of the scene or of the viewing depth.
[0407] The present application also provides a method for streaming media content of a spatial scene that varies over time.
[0408] (W1) The method includes:
[0409] Providing a media presentation description from which can be derived:
[0410] At least one version, for tile - based streaming, of the spatial scene that varies over time is provided in the at least one version,
[0411] For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatial scene that varies over time for benefiting from the tile - based streaming,
[0412] Thereby enabling a device for streaming the media content from a streaming server to
[0413] Match the beneficial requirements of the at least one version with the device capabilities of the device or of another device that interacts with the device regarding the tile - based streaming.
[0414] The present application also provides a method for allowing a device to stream media content of a spatial scene that varies over time from a server.
[0415] (X1) The method includes providing the media presentation description as described in (M1).
[0416] The present application also provides a computer program having program code for performing the method according to any one of claims (P1)-(X1) when the program is executed on a computer.
[0417] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, or an object, or a feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by this apparatus.
[0418] Signals as mentioned above, such as a streamed signal, an MPD, or any other of the mentioned signals, may be stored on a digital storage medium or may be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet.
[0419] Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or in software. The implementation may be performed using a digital storage medium storing an electronically readable control signal, such as a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, which electronically readable control signal cooperates (or is capable of cooperating) with a programmable computer system so as to cause performance of the corresponding method. Thus, the digital storage medium may be computer-readable.
[0420] Some embodiments according to the invention include a data carrier having an electronically readable control signal capable of cooperating with a programmable computer system so as to cause performance of one of the methods described herein.
[0421] Generally, embodiments of the invention may be implemented as a computer program product having program code, which is operable to perform one of the methods when the computer program product is executed on a computer. The program code may, for example, be stored on a machine-readable carrier.
[0422] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.
[0423] In other words, thus, embodiments of the inventive method are computer programs having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0424] Thus, another embodiment of the inventive method is a data carrier (or a digital storage medium, or a computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, the digital storage medium, or the recorded medium is generally tangible and / or non-transitory.
[0425] Accordingly, another embodiment of the method of the present invention is a data stream or signal sequence, which represents a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured to be transmitted (e.g., via a data communication connection such as via the Internet).
[0426] Another embodiment includes a processing component, such as a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0427] Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.
[0428] Another embodiment according to the present invention includes an apparatus or system configured to (e.g., electrically or optically) transmit a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system can include, for example, a file server for transmitting the computer program to the receiver.
[0429] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.
[0430] The apparatuses described herein can be implemented using hardware devices or using a computer or using a combination of hardware devices and a computer.
[0431] The apparatuses described herein or any components of the apparatuses described herein can be implemented at least in part in hardware and / or in software.
[0432] The methods described herein can be performed using hardware devices or using a computer or using a combination of hardware devices and a computer.
[0433] The methods described herein or any components of the apparatuses described herein can be performed at least in part by hardware and / or by software.
[0434] The above embodiments merely illustrate the principles of the present invention. It should be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is only intended to be limited by the scope of the appended patent claims, rather than by the specific details presented by the description and explanation of the embodiments herein.
Claims
1. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select (56) media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the selected media segments are from the server (20), and wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) varying over time encoded therein, encoded such that a first part (64) of the spatial section is encoded into the selected media segment at a predetermined quality and a second part (66) of the spatially varying scene varying over time that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and derive the predetermined relationship from information (68) included in the selected media segment and / or from a signal obtained from the server (20).
2. The apparatus according to claim 1, wherein each of the plurality (46) of media segments has an associated spatio-temporal part of the spatially varying scene (30) varying over time encoded therein at an associated quality level within a set of quality levels.
3. The apparatus according to claim 2, wherein the spatio-temporal part of the spatially varying scene (30) encoded into each of the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of tiles (50) into which the spatially varying scene (30) is spatially subdivided.
4. The apparatus according to any one of claims 1 to 3, wherein the information (68) indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
5. The apparatus according to claim 4, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from an observation section (28).
6. The apparatus according to claim 4 or 5, wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section (28) by way of a list of pairs of respective distances from the observation section (28) and corresponding tolerable values for the measure of the difference beyond the respective distances.
7. The apparatus according to any one of claims 4 to 6, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.
8. The apparatus according to any one of claims 4 to 7, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality, and an indication of a maximum allowable time interval during which the second part can be together with the first part within the observation section.
9. The apparatus according to claim 8, wherein the information (68) indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and the second portion is capable of being with the first portion an indication of another maximum allowable time interval within the observation section.
10. The apparatus according to any one of claims 4 to 9, wherein the information (68) is time-varying and / or spatially varying.
11. The apparatus according to any one of claims 1 to 10, wherein the information (68) indicates an allowable concurrent setting pair for the other quality and the predetermined quality.
12. The apparatus according to any one of claims 1 to 11, wherein the apparatus is configured to perform the selection such that the first portion (64) follows a time-varying observation section (28) of the time-varying spatial scene (30).
13. The apparatus according to claim 12, wherein the apparatus is configured such that the spatial position of the time-varying observation section (28) is changed according to a user input.
14. The apparatus according to any one of claims 1 to 13, wherein the apparatus is configured to determine the first portion (64) to correspond to a region of interest.
15. The apparatus according to claim 14, wherein the apparatus is configured to extract information about the region of interest from the server.
16. A streaming media server for media content of a time-varying spatial scene, configured to: make a plurality of media segments available for extraction by a device, so that the device can select media segments for extraction, the media segments having at least a spatial section of the time-varying spatial scene encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the time-varying spatial scene spatially adjacent to the first portion is encoded into the selected media segment at another quality, and signal information about a predetermined relationship in the media segment and / or by means of a signal to the device, the other quality satisfying the predetermined relationship with respect to the predetermined quality.
17. The streaming media server according to claim 16, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the time-varying spatial scene (30) encoded therein at an associated quality level within a set of quality levels.
18. The streaming media server according to claim 17, wherein each of the spatio-temporal portions of the time-varying spatial scene (30) encoded into the plurality (46) of media segments is a time segment (54) of the time-varying spatial scene (30) at a respective one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.
19. The streaming media server according to any one of claims 16 to 18, wherein the information (68) indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.
20. The streaming media server according to claim 19, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).
21. The streaming media server according to claim 19 or 20, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by means of a list of pairs of corresponding distances from the viewing section (28) and corresponding tolerable values for the measure of the difference beyond the corresponding distances.
22. The streaming media server according to any one of claims 19 to 21, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
23. The streaming media server according to any one of claims 19 to 22, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.
24. The streaming media server according to claim 23, wherein the information (68) indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.
25. The streaming media server according to any one of claims 19 to 24, wherein the information (68) is time-varying and / or spatially varying.
26. The streaming media server according to any one of claims 19 to 25, wherein the information (68) indicates allowed concurrent setting pairs for the other quality and the predetermined quality.
27. The streaming media server according to any one of claims 19 to 26, wherein the server is configured to send information about the region of interest to the device.
28. A media presentation description, comprising: information about calculating the addresses of a plurality of media segments such that a device can use the information to select and extract media segments from the plurality of media segments, wherein at least a spatial section of the spatially varying scene over time is encoded into the media segments in such a way that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatially varying scene over time that is spatially adjacent to the first part is encoded into the selected media segment with another quality, and information (68) about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.
29. The media presentation description according to claim 28, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the spatially varying scene (30) with time variation encoded therein at an associated quality level within a set of quality levels.
30. The media presentation description according to claim 29, wherein each of the spatio-temporal portions of the spatially varying scene (30) with time variation encoded into the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.
31. The media presentation description according to any one of claims 28 to 30, wherein the information (68) indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.
32. The media presentation description according to claim 31, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).
33. The media presentation description according to claim 31 or 32, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by a list of pairs of respective distances from the viewing section (28) and corresponding tolerable values for measures of differences beyond the respective distances.
34. The media presentation description according to any one of claims 31 to 33, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
35. The media presentation description according to any one of claims 31 to 34, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
36. The media presentation description according to claim 35, wherein the information (68) indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of a maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
37. The media presentation description according to any one of claims 31 to 36, wherein the information (68) is time-varying and / or spatially varying.
38. The media presentation description according to any one of claims 31 to 37, wherein the information (68) indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
39. The media presentation description according to any one of claims 31 to 38, wherein the media presentation description includes information about a region of interest.
40. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) varying over time encoded therein, encoded such that: a first part (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second part (66; 72) of the spatially varying scene varying over time that is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first part (64) follows an observation section (28) of the spatially varying scene (30) varying over time, and set a size and / or position of the first part (64) depending on information (74) contained in the selected media segment and / or a signal received from the server.
41. The apparatus according to claim 40, wherein the information (74) indicates the size in the form of an increment relative to a size of the observation section varying over time or a scaling of the size of the observation section varying over time.
42. The apparatus according to claim 40 or 41, wherein the information (74) indicates a predetermined value of a measure of a spatial speed for the observation section.
43. The apparatus according to claim 42, wherein the information (74) indicates the predetermined value of the measure of the spatial speed for the observation section for: a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating a percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating a percentage of users for whom the observation section is in a predetermined region, and / or a hint indicating one or more types of user input for controlling movement of the observation section to which the predetermined value is applicable.
44. The apparatus according to claim 42 or 43, the apparatus being configured to perform the setting such that the higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.
45. The apparatus according to any one of claims 40 or 44, wherein the information (74) indicates a predetermined value of a measure of a probability of a movement direction for the observation section.
46. The apparatus according to claim 45, the apparatus being configured to perform the setting such that the higher the probability of the corresponding movement direction, the more the first part extends into the corresponding direction.
47. The apparatus according to any one of claims 40 or 46, wherein the information (74) indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.
48. The apparatus according to any one of claims 40 or 47, configured such that the first portion (64) coincides with the spatial section (62).
49. The apparatus according to any one of claims 40 or 48, wherein each of the spatio-temporal portions of the spatially varying scene (30) encoding the time variation in the plurality of media segments is a time segment (54) of the spatially varying scene (30) at a corresponding one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.
50. The apparatus according to claim 49, wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying scene encoding the time variation therein at an associated quality level within a set of quality levels.
51. The apparatus according to claim 40, wherein the apparatus is configured to set the size in a manner independent of the size of the time-varying observation section depending on the information (74).
52. The apparatus according to claim 40, wherein the information includes different values of the size for different size options of the time-varying observation section, and the apparatus uses the values included in the information for the size option that fits the actual size of the time-varying observation section.
53. The apparatus according to claim 51 or 52, wherein the information indicates the size in terms of the number of tiles.
54. A streaming media server for media content of a spatially varying scene over time, configured to: make a plurality of media segments available for extraction by an apparatus, such that the apparatus can select media segments for extraction, the media segments having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first portion follows a time-varying observation section of the spatially varying scene over time, and signal in the media segments and / or via a signal to the apparatus information on how to set the size and / or position of the first portion.
55. A signal defining a media presentation description, comprising: Information about calculating addresses of a plurality of media segments such that a device can, using the information, select and extract a media segment from the plurality of media segments, the media segment having at least a spatial section of the spatially varying scene varying over time encoded therein in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene varying over time that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced with respect to the predetermined quality, and such that the first part follows an observation section of the spatially varying scene varying over time, and Information in the media segment and / or acting through a signal to the device on how to set the size and / or position of the first part.
56. The signal according to claim 55, wherein the information (74) indicates the size in the form of an increment with respect to the size of the observation section of the spatially varying scene varying over time or a scaling of the size of the observation section of the spatially varying scene varying over time.
57. The signal according to claim 55 or 56, wherein the information (74) indicates a predetermined value of a measure of the spatial speed for the observation section.
58. The signal according to claim 57, wherein the information (74) indicates the predetermined value of the measure of the spatial speed for the observation section for: a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the observation section is in a predetermined area, and / or a hint of one or more types of user input controlling the movement of the observation section for which the predetermined value is applicable.
59. The signal according to claim 57 or 58, the signal being configured to perform the setting such that the higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.
60. The signal according to any one of claims 55 or 59, wherein the information (74) indicates a predetermined value of a measure of the probability of the direction of movement of the observation section.
61. The signal according to claim 60, the signal being configured to perform the setting such that the higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.
62. The signal according to any one of claims 55 or 61, wherein the information (74) indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.
63. The signal according to any one of claims 55 or 62, configured such that the first part (64) coincides with the spatial section (62).
64. The signal according to any one of claims 55 or 63, wherein each of the spatio-temporal portions of the spatially varying spatio-temporal scene (30) encoded into the plurality of media segments is a temporal segment (54) of the spatially varying spatio-temporal scene (30) at a corresponding one of the tiles (50) into which the spatially varying spatio-temporal scene (30) is spatially subdivided.
65. The signal according to claim 64, wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying spatio-temporal scene encoded therein at an associated quality level within a set of quality levels.
66. The signal according to claim 55, wherein the apparatus is configured to set the size in a manner independent of the size of the temporally varying viewing section depending on the information (74).
67. The signal according to claim 55, wherein the information includes different values for the size for different size options of the time-varying viewing section, and the apparatus uses the values included in the information for the size option that fits the actual size of the time-varying viewing section.
68. The signal according to claim 65 or 66, wherein the information indicates the size in terms of the number of tiles.
69. A video bitstream having video encoded therein, the video bitstream including a signaling function for one or more of the size of a focus region within the video to which decoding capabilities for decoding the video should be concentrated and a recommended preferred viewing section region of the video.
70. A decoder for decoding video from a video bitstream, configured to: derive a signaling function (74) of the size of a focus region within the video from the video bitstream, and concentrate decoding capabilities for decoding the video to the focus region (116).
71. The decoder according to claim 70, the decoder being configured to specifically decode the focus region.
72. The decoder according to claim 70, the decoder being configured to decode each image of the video starting at the focus region where decoding begins.
73. The decoder according to claim 70, the decoder being configured to stop decoding each image of the video after decoding the focus region.
74. The decoder according to any one of claims 70 to 73, wherein the signaling function absolutely indicates the size, or the decoder is configured to scale the size of the focus region with a parameter included in the signaling function.
75. An apparatus for streaming media content regarding a spatially varying spatio-temporal scene (30), configured to: derive (90) from a media presentation description: at least one version, for tile-based streaming, the spatially varying spatio-temporal scene being provided in the at least one version, for each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying spatio-temporal scene to benefit from the tile-based streaming, Match the benefit requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming media (92).
76. The device according to claim 75, wherein the benefit requirements and the device capabilities relate to decoding capabilities.
77. The device according to claim 75 or 76, wherein the benefit requirements and the device capabilities relate to the number of available decoders.
78. The device according to any one of claims 75 to 77, wherein the benefit requirements and the device capabilities relate to layer and / or profile descriptors.
79. The device according to any one of claims 75 to 78, wherein the benefit requirements and the device capabilities relate to the type of input device for moving the viewing section across the spatial scene (30) that changes over time, or to the speed of moving the viewing section across the spatial scene (30) that changes over time using the input device.
80. The device according to any one of claims 75 to 79, wherein the device is configured to: Select media segments from a plurality (46) of media segments (58) available on the server (20), the selection calculating the addresses of the selected media segments by using calculation rules included in the media presentation description, Extract the selected media segments from the server (20) using the calculated addresses, wherein the device is configured to perform the selection such that the selected media segments have at least a spatial section (62) of the spatial scene (30) that changes over time encoded therein, encoded in such a way that: A first part (64) of the spatial section (62) is encoded into the selected media segment with a predetermined quality, and a second part (66; 72) of the spatial scene that changes over time and is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment with another quality reduced relative to the predetermined quality, and such that the first part (64) follows the viewing section (28) of the spatial scene (30) that changes over time.
81. A streaming media server for streaming media content regarding a spatial scene (30) that changes over time, configured to: Provide (90) a media presentation description from which can be derived At least one version, for tile-based streaming, the spatial scene that changes over time is provided in the at least one version, For each of the at least one version, an indication of the benefit requirements for the corresponding version of the spatial scene that changes over time to benefit from the tile-based streaming, Thereby enabling a device streaming the media content from the streaming media server to Match the benefit requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming media (92).
82. A media presentation description, comprising: Information about at least one version, for tile - based streaming, a spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements of the corresponding version for benefiting from the tile - based streaming of the spatially varying scene over time.
83. The media presentation description according to claim 82, wherein the beneficial requirements and the device capabilities relate to decoding capabilities.
84. The media presentation description according to claim 82 or 83, wherein the beneficial requirements and the device capabilities relate to the number of available decoders.
85. The media presentation description according to any one of claims 82 to 84, wherein the beneficial requirements and the device capabilities relate to layer and / or profile descriptors.
86. The media presentation description according to any one of claims 82 to 85, wherein the beneficial requirements and the device capabilities relate to the type of input device for moving an observation section across the spatially varying scene (30) over time, or to the speed of using the input device to move the observation section across the spatially varying scene (30) over time.
87. The media presentation description according to any one of claims 82 to 86, further comprising calculation rules, using which the device can select a media segment from a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rules.
88. A device for streaming media content about a spatially varying scene (30) over time, configured to: calculate the address of a media segment depending on a spatial viewport position and at least one parameter, the media segment describing the spatially varying scene (30) varying over time and in the at least one parameter, use the calculated address to extract the media segment.
89. The device according to claim 88, wherein the at least one parameter includes one or more coordinates of an observation center and / or an observation depth.
90. A media presentation description, comprising: calculation rules for calculating the address of a media segment depending on a spatial viewport position and at least one parameter, in order to use the calculated address to extract the media segment, the media segment describing the spatially varying scene (30) varying over time and in the at least one parameter.
91. A streaming server for allowing a device to stream media content about a spatially varying scene over time from a server, configured to provide the media presentation description according to claim 90.
92. A device for streaming media content about a spatially varying scene over time, configured to: select a media segment from a plurality of media segments available on the server, wherein the device is configured to: encode a first part of the spatially varying scene over time into the selected media segment with increased quality compared to the spatial neighborhood of the first part or in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segment, emit a log message that records: Instantaneous measurement results of the spatial position and / or movement of the first part; and / or Statistical values of the spatial position and / or movement of the first part, such as time average values; and / or Instantaneous measurement results of the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section; and / or Indications of the set of buffers (300) of the device involved in buffering the selected media segment, descriptions of the distribution rules applied in distributing the selected media segment to the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or Measurement results of the amount of the selected media segment yet to be output from the buffer of the device for undergoing decoding (42); and / or Statistical values of the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section, such as time average values; and / or Instantaneous measurement results of the quality of the first part or the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section; and / or Statistical values of the quality of the first part or the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section, such as time average values; and / or The field of view covered by the observation section; and / or Instantaneous measurement results of the user position or the observation depth relative to the scene center (100); and / or Statistical values of the user position or the observation depth relative to the scene center (100), such as time average values.
93. The device according to claim 92, wherein the quality of the first part or the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.
94. The device according to claim 92 or 93, configured to perform the selection such that the first part (64) of the spatially varying scene pins the observation section (28).
95. The device according to any one of claims 92 to 94, configured to issue a log message that records the instantaneous measurement result of the quality of the spatial scene up to the time variation encoded into the selected media segment and up to being visible in the observation section as one of the following: Measurement results of the average density of the pixels falling within the observation section (28), with the spatially varying scene encoded into the selected media segment at the average density.
96. The device according to claim 95, configured such that the measurement result measures the average density of the pixels by averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
97. The apparatus according to claim 95, configured such that the measurement result measures the average density of pixels by averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded in the selected media segment.
98. The apparatus according to claim 95, configured such that the emitted log message indicates whether the measurement result measures the average density of pixels by: averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded in the selected media segment, or averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded in the selected media segment.
99. The apparatus according to claim 97 or 98, wherein averaging the pixel density in a spatially non-uniform manner corresponds to averaging in a spherically uniform manner, or averaging spatially uniformly with respect to a viewport plane (310) that is perpendicular to the central viewing direction (312) of the viewing section (28).
100. The apparatus according to any one of claims 95 to 99, configured such that the measurement result measures the average density of pixels by: averaging the pixel density in a manner that limits the averaging to a central sub-section of the viewing section (28), or applying a higher averaging weight to the central sub-section (202) compared to the edge portion (204) of the viewing section surrounding the central sub-section.
101. The apparatus according to any one of claims 95 to 100, configured such that the measurement result measures the average density of pixels separately along a horizontal viewing section axis (204) and a vertical viewing section axis (206).
102. The apparatus according to any one of claims 95 to 101, configured to intermittently emit log messages.
103. The apparatus according to any one of claims 95 to 102, configured to emit log messages at a rate controlled by a manifest file, the apparatus performing the selection of the media segment for download based on the manifest file.
104. The apparatus according to any one of claims 95 to 103, wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, the representations differing in one or more of the following: the scene section (50) of the time-varying spatial scene encoded in the media segment, the quality used to encode the time-varying spatial scene in the media segment, the spatial quality variation used to encode the time-varying spatial scene in the media segment, wherein the apparatus is configured to emit a log message that records, in the form of an association of each buffer with a combination of one or two or more of the following, a description of the distribution rule applied in distributing the selected media segment to the set of buffers: the scene section, the quality, the spatial quality distribution, representation 105. The apparatus according to any one of claims 95 to 104, wherein the representations are grouped into adaptation sets according to one or more of the following: the scene segments of the spatially varying scene over time encoded into the media segment, the spatial quality variation used to encode the spatially varying scene over time into the media segment, wherein the apparatus is configured to emit a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, either in the form of an association of each buffer with one of the adaptation sets or in the form of an association of each buffer with one of the representations.
106. The apparatus according to any one of claims 95 to 105, wherein the apparatus is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the apparatus and is to undergo decoding (42), in the form of a time measurement result.
107. The apparatus according to claim 106, wherein the apparatus is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffer of the apparatus and is to undergo decoding (42) in the form of a time measurement result in time units less than the time length of the media segment and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.
108. The apparatus according to any one of claims 95 to 107, wherein the apparatus is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the apparatus and is to undergo decoding (42) in a format classified by one or more of the following: the buffer of the decoder in which the corresponding media segment has been buffered, the scene segment encoded into the corresponding media segment, the quality used to encode the spatially varying scene over time into the corresponding media segment, the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.
109. A method for streaming media content regarding a spatially varying scene (30) over time, comprising: selecting (56) media segments from a plurality (46) of media segments (58) available on a server (20), extracting (60) the selected media segments from the server (20), performing the selection such that the selected media segments have at least a spatial segment (62) of the spatially varying scene (30) encoded therein, encoded such that a first part (64) of the spatial segment is encoded into the selected media segment at a predetermined quality and a second part (66) of the spatially varying scene that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and The method further includes deriving the predetermined relationship from information (68) included in the selected media segment and / or from a signal received from the server (20).
110. A method for streaming media content of a spatially varying scene over time, comprising: Making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and Signaling information about a predetermined relationship in the media segment and / or by signaling to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
111. A method for streaming media content of a spatially varying scene (30) over time, comprising: Selecting a media segment from a plurality (46) of media segments (58) available on a server (20), Extracting the selected media segment from the server (20), wherein the selection is performed such that the selected media segment has at least a spatial section (62) of the spatially varying scene (30) over time encoded therein, encoded such that A first portion (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second portion (66; 72) of the spatially varying scene over time that is spatially adjacent to the first portion (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and Causing the first portion (64) to follow an observation section (28) of the spatially varying scene (30) that varies over time, and The method further includes setting the size of the first portion (64) depending on information (74) included in the selected media segment and / or on a signal received from the server.
112. A method for streaming media content of a spatially varying scene over time, comprising: Making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and causing the first portion to follow an observation section of the spatially varying scene that varies over time, and Signal information on how to size the first part in the media segment and / or via signaling to the device.
113. A method for decoding video from a video bitstream, comprising: Deriving a signal for the size of a focus region within the video from the video bitstream, and Concentrating decoding capabilities for decoding the video to the focus region.
114. A method for streaming media content of a spatially varying scene (30) over time, performed by a device, comprising: Deriving (90) from a media presentation description: At least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatially varying scene over time for benefiting from the tile-based streaming, Matching (92) the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming.
115. A method for streaming media content of a spatially varying scene over time, comprising: Selecting a media segment from a plurality of media segments available on a server, Extracting the selected media segment from the server, where the selection is performed such that the selected media segment has the first part of the spatially varying scene encoded therein with increased quality compared to the spatial neighborhood of the first part or in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segment, The method further includes emitting a log message that records: An instantaneous measurement of the spatial position and / or movement of the first part; and / or A statistical value of the spatial position and / or movement of the first part, such as a time average; and / or An instantaneous measurement of the quality of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section; and / or An indication of the set of buffers of the device involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment to the set of buffers, and the instantaneous buffer fullness of each of the set of buffers; and / or A measurement of the amount of the selected media segment yet to be output from the device's buffer for undergoing decoding (42); and / or A statistical value of the quality of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section, such as a time average; and / or An instantaneous measurement of the quality of the first part or of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section; and / or A statistical value of the quality of the first part or of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section, such as a time average; and / or The field of view covered by the viewing section; and / or Instantaneous measurements of a user position or viewing depth relative to a scene center (100); and / or Statistical values, such as time averages, of a user position or viewing depth relative to a scene center (100).
116. A method for streaming media content of a spatially varying scene (30) over time, comprising: Providing (90) a media presentation description from which can be derived: At least one version, for tile-based streaming, the spatially varying scene over time being provided in the at least one version, For each of the at least one version, an indication of beneficial requirements of a corresponding version of the spatially varying scene over time for benefiting from the tile-based streaming, Thereby enabling an apparatus for streaming the media content from a streaming server to Match (92) the beneficial requirements of the at least one version with the apparatus capabilities of the apparatus or another apparatus interacting with the apparatus regarding the tile-based streaming.
117. A method for allowing an apparatus to stream media content of a spatially varying scene (30) over time from a server, comprising providing a media presentation description as claimed in claim 90.
118. A computer program having program code for performing the method according to any one of claims 109 to 117 when the program is executed on a computer.
Citation Information
Patent Citations
Spatially unequal streaming
CN115037917A