Spatially unequal streaming
By using spatial uneven methods to select and encode media segments in VR streaming, the quality instability caused by viewport changes during user interaction is solved, and the user experience and bandwidth utilization efficiency is improved.
Patent Information
- Application Number
- CN202510525021.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-08
- Filing Date
- 2017-10-11
- Publication Date
- 2025-07-22
AI Technical Summary
The existing VR streaming technology cannot respond quickly to viewport changes during user interaction, causing users to see a mixture of high-quality and low-quality areas in the viewport, affecting visible quality and bandwidth consumption.
By selecting and encoding media clips of different quality based on predetermined relationships and signal prompts on the server side, streaming spatial scene content in a spatially uneven manner, ensuring that users maintain a high-quality experience in the viewport.
It improves the user's visible quality in the viewport, reduces the bandwidth consumption and computing complexity of the streaming process, and enhances the adaptability to user interaction.
Smart Images

Figure CN120358338A_ABST
Abstract
Description
[0001] This application is a divisional application of the applicant Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V., with an application date of October 11, 2017, an application number of 202210671217.X, and an invention title of "Spatial Unequal Streaming". Technical Field
[0002] This application relates to spatial unequal streaming such as occurs in virtual reality (VR) streaming. Background Art
[0003] VR streaming typically involves the transmission of extremely high-resolution videos. The resolution ability of the human fovea is about 60 pixels per degree. Considering the transmission of a full sphere of 360°×180°, the transmission can be ended by sending a resolution of about 22k×11k pixels. Since sending this high resolution will generate extremely high bandwidth requirements, another solution is to only send the viewport shown at the head-mounted display (HMD), which has a field of view of 90°×90°, thus generating a video of about 6k×6k pixels. The compromise between sending the full video at the highest resolution and only sending the viewport is to send the viewport at a high resolution and send some adjacent data (or the rest of the spherical video) at a lower resolution or lower quality.
[0004] In the DASH scenario, omnidirectional videos (also known as spherical videos) can be provided in the way controlled by the DASH client with the previously described hybrid resolution or hybrid quality videos. The DASH client only needs to know the information describing how the content is provided.
[0005] An example can be to provide different representations with different projections, which have asymmetric characteristics, such as different qualities and distortions for different parts of the video. Each representation will correspond to a given viewport and will encode the viewport at a higher quality / resolution than the rest of the content. Knowing the orientation information (the direction of the viewport for which the content has been encoded at a higher quality / resolution), the DASH client can dynamically select one or another representation to match the user's viewing direction at any time.
[0006] A more flexible option for the DASH client to select this asymmetric characteristic for omnidirectional videos is when the video is split into several spatial zones, where each zone can be obtained at a different resolution or quality. One option can be to split the video into rectangular zones (also known as tiles) based on a grid, but other options are foreseeable. In this case, the DASH client will need some signaling about the different qualities at which different zones are provided, and the DASH client can download different zones at different qualities so that the quality of the viewport shown to the user is better than the other unshown content.
[0007] In any of the previous scenarios, when user interaction occurs and the viewport has changed, the DASH client takes some time to react to the user's movement and download content in a way that matches the new viewport. During the time between the user's movement and the DASH client adapting its requests to match the new viewport, the user will see some regions of high quality and low quality simultaneously in the viewport. Although the acceptable quality / resolution differences are content-dependent, the quality seen by the user is reduced in any case.
[0008] Therefore, a concept that would have the effect of alleviating or more effectively manifesting or even increasing the visible quality of the user with respect to the partial rendering of spatial scene content streamed via adaptive streaming could be beneficial. Summary of the Invention
[0009] Accordingly, an object of the present invention is to provide a concept for streaming spatial scene content in a spatially non-uniform manner such that the visible quality of the user is increased, or the processing complexity or the required bandwidth at the streaming extraction site is reduced, or to provide a concept for streaming spatial scene content in a way that increases its applicability to other application scenarios.
[0010] This object is achieved by a device for streaming media content of a spatial scene with respect to temporal changes, a streaming server for media content of a spatial scene with respect to temporal changes, a media presentation description, a signal defining the media presentation description, a video bitstream, a decoder for decoding video from the video bitstream, a streaming server for streaming media content of a spatial scene with respect to temporal changes, a streaming server for allowing a device to stream media content of a spatial scene with respect to temporal changes from the server, a method for streaming media content of a spatial scene with respect to temporal changes, a method for decoding video from the video bitstream, a method for allowing a device to stream media content of a spatial scene with respect to temporal changes from the server, and a computer program having program code.
[0011] A first aspect of the present application is based on the following discovery: If the selected and extracted media segments and / or the signals obtained from the server provide a hint to the extraction device regarding a predetermined relationship that the quality used to encode different parts of a spatially varying scene over time complies with, streaming media content (such as video) of a spatially varying scene over time in a spatially unequal manner can be improved in terms of the visible quality at a comparable bandwidth consumption and / or the computational complexity at the streaming reception site. Otherwise, the extraction device may not know in advance how the juxtaposition of the parts encoded with different qualities into the selected and extracted media segments negatively affects the overall visible quality experienced by the user. The information contained in the media segments and / or the signals obtained from the server (such as, within a manifest file (media presentation description) or additional streaming-related control messages from the server to the client (such as SAND messages)) enables the extraction device to make an appropriate selection among the media segments provided at the server. In this way, the virtual reality streaming or partial streaming of video content can become more robust with respect to quality degradation that occurs due to an insufficient distribution of the available bandwidth over this spatial section of the spatially varying scene presented to the user.
[0012] On the other hand, the present invention is based on the discovery that streaming media content (such as video) of a spatially varying scene over time in a spatially non-uniform manner (such as using a first quality for a first part and a lower second quality for a second part or leaving the second part un-streamed) by determining the size and / or position of the first part depending on the information contained in the media segment and / or the signal action obtained from the server can result in an improvement in the visible quality and / or the bandwidth consumption and / or the computational complexity at the extraction side of the streaming becomes less complex. For example, for tile-based streaming, it is contemplated that a spatially varying scene over time can be provided at the server in a tile-based manner, i.e., the media segment can represent a spectral-temporal part of the spatially varying scene over time, each of which can be a temporal segment of the spatially varying scene within the corresponding tile of the distribution of tiles into which the spatial scene is subdivided. In this case, the extraction device (client) makes a decision on how to distribute the available bandwidth and / or computational power in the spatial scene (i.e., at tile granularity). The extraction device can perform a selection of the media segment such that a first part of the spatial scene (which respectively follows an observation section that tracks the temporal changes of the spatial scene) is encoded into the selected and extracted media segment at a predetermined quality, which can be (for example) the highest quality achievable under the current bandwidth and / or computational power conditions. For example, a second part of the spatially adjacent spatial scene may not be encoded into the selected and extracted media segment or may be encoded into the media segment at another quality lower than the predetermined quality. In this case, counting the number of adjacent tiles is computationally complex or even infeasible, and the aggregation of these tiles completely covers the observation section of the temporal change, regardless of the orientation of the observation section. Depending on the projection selected to map the spatial scene onto individual tiles, the angular scene coverage of each tile can vary in this scene, and the fact that individual tiles can overlap even makes the calculation of counting adjacent tiles that are sufficient to cover the observation section in terms of space (regardless of the orientation of the observation section) more difficult. Therefore, in this case, the foregoing information can respectively indicate the size of the first part as the count N of tiles or the number of tiles. By this measure, the device will be able to track the observation section of the temporal change by selecting those media segments having a co-located aggregation of N tiles encoded at the predetermined quality. The fact that the aggregation of these N tiles sufficiently covers the observation section can be ensured by the information indicating N. Another example can be the information contained in the media segment and / or the signal action obtained from the server, which indicates the size of the first part relative to the size of the observation section itself. For example, this information can to some extent set a "safety zone" or a prefetch zone around the actual observation section in order to account for the movement of the observation section of the temporal change. The greater the speed at which the observation section of the temporal change moves across the spatial scene, the greater the safety zone should be.Thus, the aforementioned information can indicate the size of the first part in a manner that varies with the size of the observation section over time (such as in an incremental or scaled manner). An extraction device that sets the size of the first part based on this information will be able to avoid quality degradation that could otherwise occur due to unextracted or low-quality parts of the spatial scene being visible in the observation section. Here, it is irrelevant whether this scene is provided in a tile-based manner or in some other way.
[0013] Related to the just-mentioned aspect of the present application, the video bitstream encoding the video can be decoded with increased quality provided that the video bitstream has a signaling function regarding the size of the focused area within the video, and the decoding capabilities for decoding the video should be concentrated on the focused area. By this measure, a decoder that decodes the video from the bitstream can concentrate or even limit its decoding capabilities for the decoded video to the part having the size of the focused area signaled in the video bitstream, thereby knowing (for example) that the part decoded in this way is decodable with the available decoding capabilities and spatially covers the desired section of the video. For example, the size of such signaled focused area can be selected to be large enough to cover the size of the observation section and the movement of this observation section, thereby taking into account the decoding delay when decoding the video. Or, in other words, the signaling of the recommended preferred observation section area of the video contained in the video bitstream can allow the decoder to process this area in a better way, thereby allowing the decoder to concentrate its decoding capabilities accordingly. Whether or not region-specific decoding capability concentration is performed, the region signaling can be forwarded to the platform where it is selected which media segments to download, i.e., where to place the quality-increased parts and how to size the quality-increased parts.
[0014] The first and second aspects of the present application are closely related to the third aspect of the present application. According to the third aspect, by virtue of the fact that a large number of extraction devices stream media content from a server in order to obtain information, the information can subsequently be used to appropriately set the information of the foregoing type, thereby allowing the size or size and / or position of the first part to be set, and / or the predetermined relationship between the first quality and the second quality to be appropriately set. Thus, according to this aspect of the present application, the extraction device (client) issues a log message that records one of the following: an instantaneous measurement result or statistical value of measuring the spatial position and / or movement of the first part; an instantaneous measurement result or statistical value of measuring the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section; and an instantaneous measurement result or statistical value of measuring the quality of the first part or the quality of the spatial scene that changes over time until it is encoded into the selected media segment and until it is visible in the viewing section. The instantaneous measurement result and / or statistical value may be provided with time information related to the time when the corresponding instantaneous measurement result or statistical value was obtained. The log message may be sent to the server where the media segment is located, or to some other device that evaluates the incoming log message, in order to update the current setting of the foregoing information used to set the size or size and / or position of the first part based on the log message, and / or to derive the predetermined relationship based on the log message.
[0015] According to another aspect of the present application, in particular, it is more effective in avoiding useless streaming trials by providing a media presentation description to stream media content (such as video) of a spatially varying scene over time in a tile-based manner. The media presentation description includes: at least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming of the corresponding version of the spatially varying scene over time. By this measure, the extraction device can match the beneficial requirements of the at least one version with the device capabilities of the extraction device itself or another device that interacts with the extraction device regarding tile-based streaming. For example, the benefit requirements may be related to decoding capability requirements. That is, if the decoding capability for decoding the streamed / extracted media content will not be sufficient to decode all the media segments required to cover the viewing section of the spatially varying scene over time, then attempting to stream and present the media content will waste time, bandwidth, and computing power, and thus, it may be more effective not to attempt to stream and present the media content in any case. For example, if (for example) the media segments related to a particular tile form a separate media stream (such as a video stream) from the media segments related to another tile, the decoding capability requirements may, for example, indicate the number of decoder instances required for the corresponding version. For example, the decoding capability requirements may also be regarding other information, such as a particular portion of the decoder instances necessary to fit a predetermined decoding profile and / or level, or may indicate a particular minimum capability of the user input device to move the viewport / section for viewing the scene fast enough. Depending on the scene content, low mobility may not be sufficient for the user to view the portion of the scene of interest.
[0016] Another aspect of the present invention relates to an extension of streaming media content of a spatially varying scene over time. In particular, the idea according to this aspect is that the spatial scene can actually vary not only in time but also in terms of at least one other parameter (e.g., view and position, viewing depth, or some other physical parameter). The extraction device can use adaptive streaming in this context by: calculating the addresses of media segments that describe the spatially varying scene varying in time and in the at least one parameter depending on the viewport direction and the at least one other parameter; and extracting the media segments from the server using the calculated positions.
[0017] The aspects outlined above in the present application and the advantageous implementations that are the subject of the dependent claims can be combined individually or together. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The preferred embodiments of the present application are described below with reference to the drawings, in which:
[0019] Figure 1 A demonstration schematic diagram, which illustrates a system of a client and a server for virtual reality applications as an example of a situation where the embodiments described in the following figures can be advantageously used;
[0020] Figure 2 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment of the present application for describing possible operation modes of the client device, where the server 10 provides the device with information about acceptable or tolerable quality changes in the media content presented to the user;
[0021] Figure 3 Show Figure 2 modification, the part with increased quality does not care about the part of the viewing section of the tracking viewport, but about the region of interest of the media scene content signaled from the server to the client;
[0022] Figure 4 A block diagram of a client device and a schematic illustration of a media segment selection program according to an embodiment, where the server provides information on how to set the size or size and / or position of the part with increased quality, or the size or size and / or position of the actually extracted section of the media scene;
[0023] Figure 5 Show Figure 4 variant, where the information sent by the server directly indicates the size of part 64, rather than scaling the part depending on the expected movement of the viewport;
[0024] Figure 6 Show Figure 4 variant, according to which the extracted section has a predetermined quality and its size is determined by information from the server;
[0025] Figures 7a to 7c Show an illustration of Figure 4 and Figure 6 the manner in which information 74 increases the size of the part extracted with a predetermined quality by corresponding magnification of the size of the viewport;
[0026] Figure 8a Show a schematic diagram of an embodiment in which the client device sends log messages to the server or a specific evaluator for evaluating these log messages in order to (for example) obtain appropriate settings for the type of information Figures 2 to 7c discussed;
[0027] Figure 8bSchematic illustration of a tile - based cube projection of a 360 - degree scene onto tiles, and examples of how some of the tiles are covered by exemplary positions of the viewport. The small circles indicate positions in the equiangularly distributed viewports, and the shaded tiles are encoded in the downloaded segment at a higher resolution than the non - shaded tiles;
[0028] Figure 8c and Figure 8d Schematic illustration showing how the buffer fullness (vertical axis) of different buffers of a client can evolve along a time axis (horizontal), where Figure 8c it is assumed that the buffer will be used to buffer the representation of a specific tile, and Figure 8d it is assumed that the buffer will be used to buffer the omnidirectional representation of a scene encoded therein with non - uniform quality (i.e., increasing in a certain direction specific to the corresponding buffer);
[0029] Figure 8e and Figure 8f Three - dimensional graph showing different pixel density measurements within the viewport 28, differing in terms of uniformity in the sense of a sphere or an observation plane;
[0030] Figure 9 Block diagram of a client device and a schematic illustration of a media segment selection procedure when the device detects information from a server to evaluate whether a particular version of tile - based streaming provided by the server is acceptable for the client device;
[0031] Figure 10 Schematic illustration showing a plurality of media segments provided by a server according to an embodiment, allowing a media scene to depend not only on time but also on another non - temporal parameter (i.e., here, illustratively, the scene center position);
[0032] Figure 11 Schematic illustration showing a video bitstream containing information for manipulating or controlling the size of a focused region within a video encoded into the bitstream, and an example of a video decoder capable of utilizing this information. Detailed Description
[0033] For ease of understanding the description of the embodiments of the present application with respect to various aspects of the present application, Figure 1 an example of an environment in which the embodiments described subsequently in the present application can be applied and advantageously used is shown. In particular, Figure 1Disclosed is a system consisting of a client 10 and a server 20 that interact via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) can be used for the communication 22 between the client 10 and the server 20. However, the embodiments outlined subsequently should not be construed as being limited to the use of DASH, and likewise, terms such as Media Presentation Description (MPD) should be understood broadly so as to also cover manifest files that are different from those in DASH.
[0034] Figure 1 A system configured to implement a virtual reality application is described. That is, the system is configured to present to a user wearing a head-up display 24 (i.e., via an internal display 26 of the head-up display 24) an observation segment 28 of a spatially varying scene 30 that changes over time, the segment 28 corresponding to the orientation of the head-up display 24 exemplarily measured by an internal orientation sensor 32 (such as an inertial sensor of the head-up display 24). That is, the segment 28 presented to the user forms a segment of the spatially varying scene 30, the spatial position of which corresponds to the orientation of the head-up display 24. In Figure 1 the case where the spatially varying scene 30 that changes over time is depicted as an omnidirectional video or a spherical video, but Figure 1 the description and the embodiments explained subsequently can also be easily transferred to other examples, such as presenting a segment in a video, where the spatial position of the segment 28 is determined by the intersection of face access or eye access with a virtual or real projection wall or the like. Additionally, the sensor 32 and the display 26 can be included in different devices (such as a remote control and a corresponding television), respectively, or the sensor and the display can be part of a handheld device (such as a mobile device, such as a tablet computer or a mobile phone). Finally, it should be noted that some of the embodiments described later can also be applied to the situation where the area 28 presented to the user always covers the entire spatially varying scene 30 that changes over time, where the non-uniformity during the presentation of the spatially varying scene is related to, for example, an uneven distribution of quality in the spatial scene.
[0035] Additional details regarding the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are described in Figure 1 and are described below. However, these details should not be regarded as limiting the embodiments explained subsequently, but should rather serve as examples of how to implement any of the embodiments explained subsequently.
[0036] In particular, as Figure 1As shown, server 20 may include a memory 34 and a controller 36, such as a suitably programmed computer, an application specific integrated circuit, etc. The memory 34 has media segments stored thereon, and the media segments represent a spatially varying scene 30 that changes over time. Specific examples will be outlined in more detail below with respect to Figure 1 . The controller 36 answers requests sent by the client 10 by re - sending the requested media segments to the client 10, and the media presentation description may send information about itself to the client 10. Details about this are also stated below. The controller 36 may extract the requested media segments from the memory 34. Other information may also be stored in this memory, such as the media presentation description or parts thereof, which are sent from the server 20 to the client 10 in other signals.
[0037] As Figure 1 shown, the server 20 may optionally further include a stream modifier 38 that modifies the media segments sent from the server 20 to the client 10 in response to a request from the client 10 so as to produce a media data stream at the client 10 that forms a single media stream decodable by an associated decoder, but the media segments extracted in this way by the client 10 are actually aggregated from several media streams, for example. However, the presence of this stream modifier 38 is optional.
[0038] Figure 1 The client 10 is illustratively depicted as including a client device or controller 40 and one or more decoders 42 and a reprojection unit 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device (such as an FPGA or an application specific integrated circuit), etc. The client device 40 is responsible for selecting the segments to be extracted from the server 20 from among a plurality of 46 media segments provided at the server 20. For this purpose, the client device 40 first extracts a manifest or a media presentation description from the server 20. From the manifest or the media presentation description, the client device 40 obtains the calculation rules for calculating the addresses of the media segments corresponding to a specific desired spatial part of the spatial scene 30 among the plurality of 46 media segments. The client device 40 extracts the thus - selected media segments from the server 20 by sending corresponding requests to the server 20. These requests contain the calculated addresses.
[0039] The media segments extracted by the client device 40 in this way will be forwarded by the client device 40 to one or more decoders 42 for decoding. In Figure 1In the example of, the media segments thus extracted and decoded represent only the spatial segments 48 in the spatial scene 30 that vary over time for each time unit, but as indicated above, this can be different depending on (for example) the viewing segment 28 to be presented, which always covers other aspects of the entire scene. The reprojection unit 44 can optionally reproject the viewing segment 28 to be displayed to the user and cut out the viewing segment from the extracted and decoded scene content of the selected, extracted, and decoded media segments. For this purpose, as Figure 1 shown in, the client device 40 can (for example) continuously track the spatial position of the viewing segment 28 and update the spatial position in response to user orientation data from the sensor 32, and notify the reprojection unit 44 (for example) of this current spatial position of the viewing segment 28 and the reprojection mapping to be applied to the extracted and decoded media content in order to be mapped to the area forming the viewing segment 28. The reprojection unit 44 can accordingly apply the mapping and interpolation to (for example) a regular grid of pixels to be displayed on the display 26.
[0040] Figure 1 Illustrates the case where the spatial scene 30 has been mapped to the tiles 50 using cube mapping. The tiles are thus depicted as rectangular sub-regions of the cube onto which the scene 30 in the form of a sphere has been projected. The reprojection unit 44 reverses this projection. However, other examples can also be applied. For example, instead of cube projection, a projection onto a truncated cone or a non-truncated cone can be used. Furthermore, although Figure 1 the tiles are depicted as non-overlapping with respect to covering the spatial scene 30, the subdivision into tiles can involve mutual tile overlap. And as will be outlined in more detail below, it is also not mandatory for the scene 30 to be spatially subdivided into tiles 50 (as will be further explained below, each tile forms a representation).
[0041] Therefore, as Figure 1 depicted in, the entire spatial scene 30 is spatially subdivided into tiles 50. In Figure 1 the example of, each of the six faces of the cube is subdivided into 4 tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52, as Figure 1 depicted in. To be more precise, the server 20 provides more than one video 52 for each tile 50, and these videos have different qualities Q#. Even further, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# respectively form or are encoded into one of the media segments of a plurality of 46 media segments stored in the memory 34 of the server 20.
[0042] Even to emphasize again, Figure 1The examples of tile-based streamification described herein only form examples that may deviate significantly from it. For example, while Figure 1 may seem to indicate that a media segment of a higher-quality representation of scene 30 is for a tile that is consistent with the tile to which the media segment belongs, and the tile encodes scene 30 at quality Q1 therein, this consistency is not required and tiles of different qualities may even correspond to tiles of different projections of scene 30. Additionally, although not discussed so far, it is possible that Figure 1 the media segments corresponding to different quality levels depicted in
[0043] differ in terms of spatial resolution and / or signal-to-noise ratio and / or temporal resolution, etc.
[0044] Finally, different from the tile-based streamification concept (according to which the media segments that can be individually extracted by device 40 from server 20 are spatially subdivided into tiles 50 of scene 30), the media segments provided at server 20 may alternatively (for example) each encode scene 30 at a spatially varying sampling resolution in a spatially complete manner, where the sampling resolution reaches a maximum at different spatial positions in scene 30. For example, this situation can be achieved by providing at server 20 a sequence of segments 54 related to the projection of scene 30 onto a truncated cone, the frustum of the truncated cone being orientable in mutually different directions, resulting in resolution peaks in different orientations.
[0045] After having explained the system of server 20 and client 10 more generally, the functionality of client device 40 will be described in more detail with respect to an embodiment according to the first aspect of the present application. For this purpose, reference is made to Figure 2 which shows device 40 in more detail. As explained above, device 40 is used for streamifying media content of a spatial scene 30 that varies over time. As regarding Figure 1As explained, the apparatus 40 may be configured such that the streamed media content is spatially continuous with respect to the entire scene, or only with respect to a section 28 of the scene. In any case, the apparatus 40 includes: a selector 56 for selecting an appropriate media segment 58 from among a plurality of 46 media segments available on the server 20; and an extractor 60 for extracting the selected media segment from the server 20 by a corresponding request (such as an HTTP request). As described above, the selector 56 may use a media presentation description to calculate the addresses of the selected media segments, and the extractor 60 uses these addresses when extracting the selected media segment 58. For example, the calculation rules indicated in the media presentation description for calculating the addresses may depend on a quality parameter Q, a tile T, and a time segment t. For example, the address may be a URL.
[0046] As also discussed above, the selector 56 is configured to perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene encoded therein. The spatial section may cover the entire scene spatially continuously. Figure 2 An illustrative case of the apparatus 40 adapting a spatial section 62 of the scene 30 to overlap and surround the viewing section 28 is illustrated at 61. However, as mentioned above, this need not be the case, and the spatial section may cover the entire scene 30 continuously.
[0047] In addition, the selector 56 performs the selection such that the selected media segment has a section 62 encoded therein in a spatially non-uniform quality manner. More precisely, a first part 64 of the spatial section 62 (at Figure 2The middle part (indicated by the shaded line) is encoded to the selected media segment with a predetermined quality. This quality can be, for example, the highest quality provided by the server 20 or can be a "good" quality. For example, the device 42 moves or adjusts the first part 64 in a manner that spatially follows the observation segment 28 that changes over time. For example, the selector 56 selects the current time segment 54 of those tiles that inherit the current position of the observation segment 28. After such selection, as will be explained below with respect to other embodiments, the selector 56 can optionally keep the number of tiles that make up the first part 64 constant. In any case, the second part 66 of the segment 62 is encoded to the selected media segment 58 with another quality, such as a lower quality. For example, the selector 56 selects a media segment corresponding to the current time segment of tiles that are spatially adjacent to the part 64 and belong to the lower quality tiles. For example, in order to address the possible moment when the observation segment 28 moves too fast and leaves the part 64 before the end of the time interval corresponding to the current time segment and overlaps with the part 66 and the selector 56 will be able to reconfigure the part 64 spatially, the selector 56 mainly selects the media segment corresponding to the part 66. In this case, the part of the segment 28 that protrudes into the part 66 can still be presented to the user (i.e., with reduced quality).
[0048] The device 40 is not possible to evaluate for the user which negative quality degradation can be caused by presenting to the user in advance the scene content with reduced quality and the scene content within the higher quality part 64. In particular, a transition between these two qualities that can be clearly visible to the user is generated. At least, this transition depends on the current scene content within the segment 28 being visible. The severity of the negative impact of this transition within the user's field of view is determined by the characteristics of the scene content provided by the server 20 and may not be predicted by the device 40.
[0049] Therefore, according to Figure 2In an embodiment, device 40 includes a derivator 66 that derives a predetermined relationship that will be satisfied between the quality of portion 64 and the quality of portion 66. The derivator 66 derives this predetermined relationship from information that may be included in a media segment (such as within a delivery block in media segment 58) and / or included in a signaling from server 20 (such as within a media presentation description or a proprietary signal sent from server 20, such as a SAND message, etc.). Examples of what information 68 may look like are presented below. The predetermined relationship 70 derived by the derivator 66 based on information 68 is used by selector 56 to perform the selection appropriately. For example, compared to a completely independent selection of the quality of portions 64 and 66, the restrictions in selecting the quality of portions 64 and 66 affect the distribution of the available bandwidth for extracting media content of section 62 onto portions 64 and 66. In any case, selector 56 selects a media segment such that the quality used to encode portions 64 and 66 into the last extracted media segment satisfies the predetermined relationship. Examples of what the predetermined relationship may look like are also stated below.
[0050] The selected media segment, which is finally extracted by extractor 60, is finally forwarded to one or more decoders 42 for decoding.
[0051] For example, according to a first example, the signaling mechanism embodied by information 68 relates to information 68 that indicates to device 40 (which may be a DASH client) which quality combinations are acceptable for the provided video content. For example, information 68 may be a list of quality pairs that indicates to the user or device 40 the maximum quality (or resolution) difference by which different sections 64 and 66 can be mixed. Device 40 may be configured to inevitably use a particular quality level (such as the highest quality level provided at server 10) for portion 64 and derive from information 68 the quality level to be used to encode portion 66 into the selected media segment, where the information is included in the form of a list of quality levels for (e.g.) portion 68.
[0052] Information 68 may indicate a tolerable value for a measure of the difference between the quality of portion 68 and the quality of portion 64. The quality index of media segment 58 may be used as a "measure" of the quality difference, and media segments may be differentiated by the quality index in the media presentation description, and the address of the media segment is calculated using the calculation rules described in the media presentation description by the quality index. In MPEG-DASH, the corresponding attribute indicating quality may be (e.g.) @qualityRanking. Device 40 may consider the restrictions in the available quality level pairs that can be used to encode portions 64 and 66 into the selected media segment when performing the selection.
[0053] However, instead of this difference metric, the quality difference may alternatively be measured (e.g.) in terms of bitrate difference (i.e., the tolerable difference in the bitrates used to encode portions 64 and 66 into the corresponding media segments), assuming that the bitrate typically increases monotonically with increasing quality. Information 68 may indicate the allowed pairings of options for the quality used to encode portions 64 and 66 into the selected media segments. Alternatively, information 68 simply indicates the allowed quality for encoding portion 66, thereby indirectly indicating the allowed or tolerable quality difference, assuming that the main portion 64 is encoded using some default quality (such as the highest quality that may or is available). For example, information 68 may be a list of acceptable representation IDs or may indicate the minimum bitrate level for the media segment related to portion 66.
[0054] However, alternatively, a more gradual quality difference may be required, where instead of quality pairs, quality groups (more than two qualities) may be indicated, and depending on the distance from section 28 (viewport), the quality difference may increase. That is, in a manner depending on the distance from the viewing section 28, information 68 may indicate the tolerable values of the metric of the difference in quality between portions 64 and 66. This may be done by a list of pairs of the respective distances from the viewing section and the corresponding tolerable values of the metric of the quality difference for distances exceeding the respective distances. Below the respective distances, the quality difference must be lower. That is, each pair may indicate for the corresponding distance: a portion within portion 66 that is further from section 28 than the corresponding distance may have a quality difference from the quality of portion 64 that exceeds the corresponding tolerable value of this list entry.
[0055] The tolerable values may increase with the distance from the viewing section 28. The acceptability of the quality difference just discussed often depends on the time for which these different qualities are presented to the user. For example, content with a high quality difference may be acceptable if the content is presented for only 200 microseconds, while content with a lower quality difference may be acceptable if the content is presented for 500 microseconds. Thus, according to another example, in addition to (e.g.) the aforementioned quality combinations or in addition to the allowed quality differences, information 68 may also include a time interval during which the combination / quality difference may be acceptable. In other words, information 68 may indicate the tolerable or maximum allowed difference in quality between portions 66 and 64, and an indication of the maximum allowed time interval during which portion 66 may be presented simultaneously with portion 64 within the viewing section 28.
[0056] As previously mentioned, the acceptability of quality differences depends on the content itself. For example, the spatial position of different tiles 50 affects acceptability. Quality differences in a uniform background region with low-frequency signals are expected to be more acceptable than those in foreground objects. Additionally, due to content changes, the temporal position also affects acceptability. Thus, according to another example, the signal forming information 68 is sent intermittently (such as every representation or period in DASH) to device 40. That is, the predetermined relationship indicated by information 68 can be updated intermittently. Additionally and / or alternatively, the signaling mechanism implemented by information 68 can vary in space. That is, information 68 can be made spatially dependent, such as through the SRD parameter in DASH. That is, for different spatial regions of scene 30, different predetermined relationships can be indicated by information 68.
[0057] As regarding Figure 2 described, an embodiment of device 40 relates to the fact that device 40 desires to keep the quality degradation caused by pre-fetched portion 66 within the extracted portion 62 of video content 30 that is briefly visible in portion 28 as low as possible before being able to change the positions of portions 62 and 64 so as to adapt the portions to the position change caused by portion 28. That is, in Figure 2 portion 64 and 66 are different portions of portion 62, the quality of which is restricted until their possible combination is of interest to information 68, and the transition between the two portions 64 and 66 is continuously shifted or adjusted so as to track or overtake the moving viewing portion 28. According to Figure 3 the alternative embodiment shown in, device 40 uses information 68 to control the possible combination of the quality of portions 64 and 66, however, according to Figure 3 the embodiment of, portions 64 and 66 are defined as being different or distinct from each other in a manner defined, for example, in a media presentation description (i.e., in a manner independent of the position of viewing portion 28). The positions of portions 64 and 66 and the transition therebetween can be constant or vary in time. If varying in time, such variations are attributable to changes in the content of scene 30. For example, portion 64 can correspond to a region of interest worthy of consuming higher quality, while portion 66 is the portion for which quality degradation due to, for example, low bandwidth conditions should be considered before considering the quality degradation of portion 64.
[0058] In the following, another embodiment of an advantageous implementation of device 40 is described. In particular, Figure 4 device 40 is shown, which is structurally corresponding to Figure 2 and 3 but the operating mode is changed to correspond to the second aspect of the present application.
[0059] That is, the apparatus 40 includes a selector 56, an extractor 60, and a derivator 66. The selector 56 selects from among a plurality of media segments 58 provided by the server 20, and the extractor 60 extracts the selected media segment from the server. Figure 4 Assume that the apparatus 40 operates as depicted and illustrated with respect to Figure 2 and Figure 3 i.e., the selector 56 performs the selection such that the selected media segment 58 encodes a spatial segment 62 of the scene 30 in such a way that the spatial segment follows an observation segment 28 whose spatial position varies in time. However, a variant corresponding to the same aspect of the present application will subsequently be described with respect to Figure 5 wherein, for each time instant t, the selected and extracted media segment 58 encodes the entire scene or a constant spatial segment 62 therein.
[0060] In any case, similar to the description with respect to Figure 2 and 3 the selector 56 selects the media segment 58 such that a first part 64 within the segment 62 is encoded into the selected and extracted media segment at a predetermined quality, while a second part 66 of the segment 62 (which is spatially adjacent to the first part 64) is encoded into the selected media segment at a quality reduced with respect to the quality of the part 64. In Figure 6 a variant is described wherein the selector 56 restricts the selection and extraction of the media segment with respect to a moving template to tracking the position of the viewport 28, and wherein the media segment has a segment 62 encoded therein in its entirety at a predetermined quality such that the first part 64 completely covers the segment 62 while being surrounded by an uncoded part 72. In any case, the selector 56 performs the selection such that the first part 64 follows the observation segment 28 whose spatial position varies in time.
[0061] In this case, it is also not easy for the client 40 to predict how large the segment 62 or the part 64 should be. Depending on the scene content, most users are likely to perform similar actions when moving the observation segment 28 across the scene 30, and thus, the same actions apply to the intervals of the observation segment 28, which is likely to move across the scene 30 at approximately this speed. Therefore, according to Figures 4 to 6 an embodiment of, information 74 is provided by the server 20 to the apparatus 40 to assist the apparatus 40 in setting the size or size and / or position of the first part 64, or the size or size and / or position of the segment 62, respectively, depending on the information 74. Regarding the possibility of transmitting the information 74 from the server 20 to the apparatus 40, as described above with respect to Figure 2 and Figure 3The described situation applies. That is, the information may be included within media segment 58, such as within an event block of the media segment, or for this purpose, a media presentation description or a transmission within an exclusive message (such as a SAND message) sent from the server to device 40 may be used.
[0062] That is, according to Figures 4 to 6 an embodiment of, selector 56 is configured to set the size of first portion 64 depending on information 74 originating from server 20. In Figures 4 to 6 the embodiment illustrated, the size is set in units of tiles 50, but as described above with respect to Figure 1 it may be slightly different when using another concept of providing a scene 30 with spatially varying quality at server 20.
[0063] According to an example, information 70 may (for example) include the probability of a given movement speed of the viewport of viewing section 28. As indicated above, information 74 may cause a media presentation description to be available for client device 40 (which may be, for example, a DASH client), or some in-band mechanism may be used to convey information 74, such as an event block, that is, an EMSG or a SAND message in the case of DASH. Information 74 may also be included in any container format, such as the ISO file format or a transport format exceeding MPEG-DASH (such as MPEG-2TS). The information may also be conveyed in the video bitstream (such as, in an SEI message as described later). In other words, information 74 may indicate a predetermined value for a measure of the spatial speed for viewing section 28. In this way, information 74 indicates the size of portion 64, either in the form of a scaling relative to the size of viewing section 28 or in the form of an increment relative to the size of viewing section 28. That is, information 74 starts from a "base size" of portion 64 necessary to cover the size of section 28 and appropriately (such as incrementally or proportionally) increases this "base size". For example, the aforementioned movement speed of viewing section 28 may be used to correspondingly scale the perimeter of the current position of viewing section 28 in order to determine (for example) the furthest position of the perimeter of viewing section 28 in any spatial direction feasible after this time interval, for example, to determine the delay when adjusting the spatial position of portion 64, such as the duration of time segment 54 corresponding to the time length of media segment 58. The speed multiplied by this duration plus the perimeter of the current position of the omnidirectional viewport 28 may thus result in this worst-case perimeter and may be used to determine the magnification of portion 64 relative to a certain minimum expansion of portion 64 assuming a non-moving viewport 28.
[0064] The information 74 can even be about the evaluation of statistical data on user behavior. Subsequently, embodiments suitable for feeding such an evaluation program are described. For example, the information 74 can indicate the maximum speed of a certain percentage of users. For example, the information 74 can indicate that 90% of the users move at a speed lower than 0.2 radians per second and 98% of the users move at a speed lower than 0.5 radians per second. The information 74 or the message carrying the information can be defined such that a probability-speed pair is defined or the message can be defined to signal the maximum speed of a fixed percentage of users (e.g., always 99% of the users). The movement speed signaling 74 can additionally include direction information, i.e., an angle in 2D, or depth in 2D plus 3D (also referred to as light field applications). The information 74 can indicate different probability-speed pairs for different movement directions.
[0065] In other words, the information 74 can be applied to a given time span, such as the time length of a media segment. The information can consist of a trajectory-based (x percentage, average user path) or speed-based pair (x percentage, speed) or distance-based pair (x percentage, pore / diameter / preferred) or area-based pair (x percentage, recommended preferred area) or a single maximum boundary value of a path, speed, distance, or preferably area. Instead of associating the information with a percentage, a simple frequency grading can be made according to the fact that most users move at a particular speed, the second most users move at another speed, etc. Additionally or alternatively, the information 74 is not limited to indicating the speed of the observation segment 28, but can equally indicate the preferred areas to be observed separately in order to guide the attempt to track parts 62 and / or 64 of the observation segment 28, with or without an indication of the statistical significance of the indication (such as the percentage of users who have complied with that indication or an indication of whether the indication is consistent with the most frequently recorded user observation speed / observation segment), and with or without an indication of the time duration of the indication. The information 74 can indicate another measure of the speed of the observation segment 28, such as a measure of the travel distance of the observation segment 28 over a particular time period (such as within the time length of a media clip, or more specifically within the time length of the time segment 54). Alternatively, the information 74 can be notified in a way that differentiates between the particular movement directions in which the observation segment 28 can travel. This applies to both indicating the rate or speed of the observation segment 28 in a particular direction and indicating the travel distance of the observation segment 28 with respect to a particular movement direction. Additionally, the extension of part 64 can be signaled directly by the information 74 omnidirectionally or in a way that differentiates between different movement directions. Furthermore, all of the examples outlined just above can be modified where the information 74 indicates these values as well as the percentage of users for whom these values are sufficient to explain the statistical behavior when moving the observation segment 28. In this regard, it should be noted that the observation speed (i.e., the speed of the observation segment 28) can be quite large and is not limited to (for example) the speed value of the user's head. In fact, the observation segment 28 can move depending on (for example) the eye movement of the user, in which case the observation speed can be significantly larger. The observation segment 28 can also move according to the movement of another input device (such as according to the movement of a tablet computer, etc.). Since all these "input possibilities" that enable the user to move the segment 28 result in different expected speeds of the observation segment 28, the information 74 can even be designed such that the information differentiates between different concepts for controlling the movement of the observation segment 28. That is, the information 74 can indicate the size of part 64 in a way that indicates different magnitudes for different methods of controlling the movement of the observation segment 28, and the device 40 can use the size indicated by the information 74 for correct observation segment control.That is, the device 40 obtains knowledge of the way in which the viewing section 28 is controlled by the user, that is, checks whether the viewing section 28 is controlled by head movement, eye movement, or tablet computer movement or similar movement, and sets the size according to a part of the information 74 corresponding to such viewing section control.
[0066] Generally, the movement speed can be signaled according to content, time period, representation, segment, according to the SRD position, according to pixels, according to tiles (e.g., at any temporal or spatial granularity, etc.). The movement speed can also distinguish head movement and / or eye movement, as outlined just above. Additionally, the information 74 regarding the user movement probability can be conveyed as a recommendation regarding high-resolution prefetching (i.e., the video area outside the user's viewport, or the sphere coverage).
[0067] Figures 7a to 7c Briefly outline some of the options as explained regarding the information 74 in terms of its use by the device 40 to respectively modify the size and / or the position of the part 64 or the part 62. According to Figure 7a the option shown in, the device 40 magnifies the perimeter of the section 28 by a distance corresponding to the product of the signaled speed v and the duration Δt, where the duration can correspond to a time period corresponding to the time length of the time segment 54 encoded in the individual media segment 50a. Additionally and / or alternatively, the greater the speed, the farther the position of the part 62 and / or 64 can be placed away from the current position of the section 28, or the current position of the part 62 and / or 64 can be in the direction of the signaled speed or movement, as signaled by the information 74. The speed and direction can be derived from the recent development or change in measuring or extrapolating the recommended preferred area indicated by the information 74. Instead of applying v×Δt omnidirectionally, the speed can be signaled differently by the information 74 for different spatial directions. Figure 7b The alternative example depicted in shows that the information 74 can directly indicate the distance to magnify the perimeter of the viewing section 28, which is indicated by the parameter s in Figure 7b . Again, the magnification of the direction change of the section can be applied. Figure 7c shows that the magnification of the perimeter of the section 28 can be indicated by the information 74 through an area increase, such as in the form of the ratio of the area of the magnified section to the original area of the section 28. In any case, the perimeter of the area 28 after magnification (indicated by 76 in Figures 7a to 7c ) can be used by the selector 56 to dimension or set the size of the part 64 such that the part 64 covers at least the entire area within the magnified section 76 by a predetermined amount. Obviously, the larger the section 76, for example, the larger the number of tiles within the part 64. According to another alternative, the section 74 can directly indicate the size of the part 64, such as in the form of the number of tiles constituting the part 64.
[0068] In Figure 5A further possibility of signaling the size of signaling section 64 is depicted. Figure 5 The embodiment of Figure 4 can be modified in a manner similar to that by which the embodiment of Figure 6 is modified, i.e., the entire area of section 62 can be retrieved from server 20 in terms of the quality of section 64 by means of segment 58.
[0069] In any case, at the end of Figure 5 information 74 differentiates between different sizes of viewing section 28, i.e., different fields of view seen by viewing section 28. Information 74 simply indicates the size of section 64 depending on the size of viewing section 28 at which device 40 is currently aimed. This enables the service of server 20 to be used by devices having different fields of view or different sizes of viewing section 28 without a device such as device 40 having to cope with calculating or otherwise guessing the size of section 64 such that section 64 is sufficient to cover viewing section 28 regardless of any movement of section 28 (as discussed with respect to Figure 4 , Figure 6 and FIG. 7). As will become clear from the description of Figure 1 it is easy to evaluate, for example, which constant number of tiles may be sufficient to completely cover a particular size of viewing section 28 (i.e., a particular field of view) regardless of the orientation of viewing section 28 for spatial positioning 30. Here, information 74 alleviates this situation and device 40 can simply look up the value of the size of section 64 within information 74 for the size of viewing section 28 of device 40. That is, according to the embodiment of Figure 5 a media presentation description (such as an event chunk or a SAND message) available for use by a DASH client or some interested entity may include information 74 regarding the sphere coverage or field of view representing a set or a set of tiles respectively. An example could be providing M representations of tiles as depicted in Figure 1 . Information 74 may indicate a recommended number n < M of tiles (referred to as representations) to be downloaded for covering a given terminal device field of view. For example, in a cubical representation of tiles divided into 6×4 tiles as depicted in Figure 1 it is considered that 12 tiles are sufficient to cover a 90°×90° field of view. Due to the fact that the terminal device field of view may not always align perfectly with the tile boundaries, this recommendation cannot be trivially generated by device 40 itself. Device 40 may use information 74 by downloading, for example, at least N tiles, i.e., media segment 58 is associated with N tiles. Another way of using the information could be to focus on the quality of the N tiles closest to the current viewing center of the terminal device within section 62, i.e., using the N tiles to constitute section 64 of section 62.
[0070] Regarding Figure 8a an embodiment of another aspect of the present application is described. Here,Figure 8a The client device 10 and the server 20 are shown, and the two are based on the above Figure 1 7 to communicate with each other. That is, the devices 10 may communicate with each other according to the Figure 2 7, or may be implemented without these details as described above with respect to Figure 1 However, the device 10 is advantageously configured according to the above description. Figure 2 7 or any combination thereof, and further inherits the present Figure 8a In particular, the device 10 is understood internally as described above with respect to Figure 2 8, that is, the device 40 includes a selector 56, an extractor 60 and optionally a deriver 66. The selector 56 performs the selection for targeting unequal streaming, that is, selecting the media segments in a way that the media content is encoded into the selected and extracted media segments in a way that the quality varies spatially and / or there are unencoded parts. However, in addition to this, the device 40 also includes a log message sender 80 that sends log messages recorded in (for example) the following to the server 20 or the evaluation device 82:
[0071] measuring instantaneous measurements or statistics of the spatial position and / or movement of the first portion 64,
[0072] measuring instantaneous measurements or statistics of the quality of the time-varying spatial scene up to encoding into the selected media segment and up to being visible in the observation section 28, and / or
[0073] An instantaneous measurement or statistical value of the quality of the first portion or of the temporally varying spatial scene 30 up to the encoding into the selected media segment and up to the visible observation section 28 is measured.
[0074] The motivation is as follows.
[0075] In order to be able to derive statistics, such as the most interesting regions or speed-probability pairs, a reporting mechanism from the user is required as described previously. Additional DASH metrics to the statistics defined in Annex D of ISO / IEC 23009-1 are required.
[0076] One metric may be the client's field of view which is a DASH metric, where the DASH client sends characteristics of the terminal device regarding the field of view back to a metric server (which may be the same as the DASH server or another server).
[0077] Keywords Type Description EndDeviceFoVH Integer Horizontal field of view of the end device, in degrees EndDeviceFoVV Integer Vertical field of view of the end device, in degrees
[0078] A metric can be a ViewportList, where the DASH client sends back to the metric server (which can be the same as the DASH server or another server) in a timely manner the viewports seen by each client. The instantiation of this message can be as follows.
[0079]
[0080] For viewport (region of interest) messages, the DASH client may be required to report when a viewport change occurs, possibly with a given granularity (to avoid or not avoid reporting very small movements) or a given periodicity. This message can be included in the MPD as an attribute @reportViewPortPeriodicity or as an element or descriptor. This message can also be indicated out-of-band, such as using a SAND message or any other means.
[0081] Viewports can also be signaled at the tile granularity.
[0082] Additionally or alternatively, log messages can report on other current scene-related parameters that change in response to user input, such as any of the parameters discussed below Figure 10 such as the current user distance from the scene center and / or the current viewing depth.
[0083] Another metric can be a ViewportSpeedList, where the DASH client indicates the speed of movement of a given viewport when a movement occurs.
[0084]
[0085] This message can be sent only when the client performs a viewport movement. However, as in the previous case, the server can indicate that the message should be sent only when the movement is significant. This configuration can be somewhat similar to @minViewportDifferenceForReporting, for signaling the size in pixels or degrees or any other quantity that needs to be changed for the message being sent.
[0086] Another important aspect of the VR-DASH service (where the asymmetric quality as described above is provided) is to evaluate how quickly a user switches from an asymmetric representation or a set of unequal quality / resolution representations of a viewport to a more complete another representation or set of representations of another viewport. Using this metric, the server can derive statistics that help it understand the relevant factors affecting QoE. This metric can look as follows.
[0087]
[0088]
[0089] Alternatively, the previously described duration may be given as an average value.
[0090]
[0091] For other DASH metrics, all such metrics may additionally have a time measurement of when they were performed.
[0092] t Real-time Time at which the parameter is measured
[0093] In some cases, it may occur that if unequal quality content is downloaded and poor quality (or a mix of good and poor quality) is presented for a long enough time (which may only be a few seconds), the user will be unhappy and leave the session. Under the condition of leaving the session, the user may send a message with the quality presented in the most recent x time interval.
[0094]
[0095] Alternatively, the maximum quality difference may be reported or the maximum and minimum quality of the viewport may be reported.
[0096] As is clear from the above description regarding Figure 8a it is advantageous for the tile - based DASH streaming service operator to be able to derive statistics that exemplify the client reporting mechanism described above in order to set and optimize their service in a meaningful way (e.g., regarding resolution ratio, bitrate, and segment duration). Additional DASH metrics are stated below in addition to the metrics defined above and in addition to Appendix D of the document "ISO / IEC 23009 - 1:2014, Information technology--Dynamic adaptive streaming over HTTP (DASH)--Part 1: Media presentation description and segment formats".
[0097] It is envisioned that the tile - based streaming service uses a video with a cube projection as depicted in Figure 1 The reconstruction on the client side is in Figure 8bIt is described in [description], where the small circles 198 indicate the projection of the viewing directions that are equiangularly horizontally and vertically distributed within the client's viewport 28 onto the two-dimensional distribution over the image regions covered by the individual tiles 50. The tiles marked with hatching indicate high-resolution tiles, thus forming the high-resolution portion 64, and the tiles 50 shown without hatching represent low-resolution tiles, thus forming the low-resolution portion 66. It can be seen that as the viewport 28 changes, the user is partially presented with low-resolution tiles because the resolution of each tile on the cube determined by the most recent update for segment selection and download, and the projection plane or pixel array of the tile 50 encoded into the downloadable segment 58 falls on the cube.
[0098] Although the above description actually generally (especially) indicates feedback or log messages that indicate the quality of the video presented to the user in the viewport, hereinafter, more specific and advantageous metrics applicable in this regard will be outlined. The metrics now described can be reported back from the client side and are called the effective viewport resolution. It is speculated that the metric indicates to the service operator the effective resolution in the client's viewport. In the case where the reported effective viewport resolution indicates a resolution at which the user is only presented with the resolution towards the low-resolution tiles, the service operator can accordingly change the tiling configuration, resolution ratio, or segment length to achieve a higher effective viewport resolution.
[0099] One embodiment can be the average pixel count in the viewport 28 measured in the projection plane, where the pixel array of the tile 50 encoded into the segment 58 falls in the projection plane. The measurement can be distinguished with respect to the covered field of view (FoV) of the viewport 28 or be specific to the horizontal direction 204 and the vertical direction 206. The following table shows possible examples of the appropriate syntax and semantics that can be included in the log message to signal the outlined viewport quality metrics.
[0100]
[0101] The decomposition in the horizontal and vertical directions can be stopped by alternatively using a scalar value of the average pixel count. Together with an indication of the aperture or size of the viewport 28, also reportable to the recipient of the log message (i.e., the evaluator 82) is the average count that indicates the pixel density within the viewport.
[0102] Reducing the field of view considered for the metric to less than the field of view of the viewport actually presented to the user can be advantageous, thus excluding regions that are only for peripheral vision towards the boundaries of the viewport and therefore have no impact on the subjective quality perception. This alternative is illustrated by the dashed line 202, which encloses the pixels in this central section of the viewport 28. The reporting of the considered field of view 202 for the reported metric with respect to the total field of view of the viewport 28 can also be communicated to the log message recipient 82. The following table shows the corresponding extension of the previous example.
[0103]
[0104] According to another embodiment, instead of measuring the average pixel density by spatially averaging the mass uniformly in the projection plane (as was the case in the examples actually described so far for the examples containing EffectiveFoVResolutionH / V), it is measured in a way that non-uniformly weights this averaging for the pixels (i.e., the projection plane). The averaging can be performed in a spherically uniform manner. As an example, the averaging can be performed uniformly with respect to sample points distributed as in circle 198. In other words, the averaging can be performed by weighting the area density with weights that decrease quadratically with increasing local projection plane distance and increase according to the sine of the local tilt of the projection relative to the line connected to the viewport. The message can include an optional (flag-controlled) step size to adjust for the inherent oversampling in some of the available projections (such as the equirectangular projection), for example by using a uniform spherical sampling grid. Some projections do not have a large oversampling problem, and forcing the calculation to remove the oversampling can create unnecessary complexity issues. This must not be limited to the equirectangular projection. The report does not need to distinguish between horizontal and vertical resolutions, but can combine them. An example is given below.
[0105]
[0106]
[0107] In Figure 8e the application of equiangular uniformity in the averaging is illustrated by showing how points 302 (in terms of being within viewport 28) that are equiangularly horizontally and vertically distributed on a sphere 304 centered on viewport 306 project onto the projection plane 308 of the tile (here a cube) in order to perform the averaging of the pixel density 308 of the pixels arranged in an array (by rows and columns) in the projection area, in order to set the local weights for the pixel density according to the local density of the projections 198 of the points 302 onto the projection plane. Figure 8f A very similar method is depicted in Figure 8f . Here, the points 302 are equally spaced in the viewport plane perpendicular to the viewing direction 312, i.e., horizontally and vertically uniformly in rows and columns, and the projections onto the projection plane 308 define points 198, and the local density of the points controls the weights, with which the local density pixel density 308, which varies due to the high and low resolution tiles within viewport 28, contributes to the averaging. In examples such as the above table of the latest, Figure 8f an alternative to Figure 8e the example depicted in
[0108] In the following, embodiments of another classification of log messages are described, which relate to a DASH client 10 having a plurality of media buffers 300, as Figure 8a illustrated exemplarily in, i.e., the DASH client 10 forwards the downloaded segments 58 to subsequent decoding by one or more decoders 42 (compare Figure 1 ). The distribution of the segments 58 onto the buffers can be done in different ways. For example, the distribution can be made such that specific regions of the 360 video are downloaded separately from each other or cached after being downloaded into separate buffers. The following examples illustrate different distributions by making indications about: which Figure 1 tiles T, indexed #1 to #24 as shown in, (the enantiomers having a total of 25) are encoded to which individual downloadable representations R#1 to #P in qualities #1 to #M (1 being the best and M being the worst), and how these P representations R can be grouped into adaptive sets A, indexed #1 to #S in the MPD (optional), and how the segments 58 of the P representations R can be distributed onto the buffers of buffers B, indexed #1 to #N.
[0109]
[0110] Here, the representations can be provided at the server and announced in the MPD for download, each of these representations relating to one tile 50, i.e., a section of the scene. Representations that are related to one tile 50 but encode this tile 50 in different qualities can be summarized in optionally grouped adaptive sets, but precisely, this grouping is used to associate with the buffers. Thus, according to this example, for each tile 50, or in other words, for each viewport (observation section) encoded, there will be one buffer.
[0111] Another set of representations and distribution can be:
[0112]
[0113] According to this example, each representation can cover the entire area, but the high-quality area will be focused on one hemisphere, while the lower quality is used for the other hemisphere. Representations that only differ in the exact quality used in this way (i.e., evenly in the position of the higher-quality hemisphere) can be collected in one adaptive set and distributed onto the buffers (exemplarily six here) according to this characteristic.
[0114] Therefore, the following description assumes that this distribution to the buffers is applied according to different viewport encodings (video sub-regions, such as tiles) associated with an adaptive set or the like. Figure 8cDescribe the buffer fill levels over time of two separate buffers (e.g., tile 1 and tile 2) in a tile-based streaming scenario, which has been described in the last but not least table. Enabling the client to report the fill levels of all its buffers allows the service operator to correlate the data with other streaming parameters to understand the quality of experience (QoE) impact of its service settings.
[0115] Advantageously, the buffer fill levels of multiple media buffers on the client side can be reported with metrics and are identified and associated with buffer types. For example, the association types are as follows:
[0116] ● Tile
[0117] ● Viewport
[0118] ● Region
[0119] ● Adaptation set
[0120] ● Representation
[0121] ● Low-quality version of the full content
[0122] An embodiment of the present invention is given in Table 1, which defines metrics for reporting buffer level status events for each identified and associated buffer.
[0123] Table 1: List of buffer levels
[0124]
[0125]
[0126] Another embodiment using viewport-dependent encoding is as follows.
[0127] In a viewport-dependent streaming scenario, the DASH client downloads and pre-buffers several media segments related to a specific viewing orientation (viewport). If the amount of pre-buffered content is too high and the client changes its viewing orientation, the portion of the pre-buffered content that will be played after the viewport change is not presented and the corresponding media buffer is cleared. This scenario is depicted in Figure 8d in.
[0128] Another embodiment can be for a traditional video streaming scenario with multiple representations (quality / bitrate) of the same content and the quality used to encode the video content being spatially uniform.
[0129] The distribution can thus look as follows:
[0130]
[0131] That is, here, each representation can cover, for example, a complete scene that may not be a panoramic 360 scene with different qualities (i.e., spatially uniform qualities), and these representations can be individually distributed to the buffer. All examples stated in the last three tables should be considered not to limit the way the segments 58 of the representations provided at the server are distributed to the buffer. There are different methods, and the rules can be based on the membership of the segment 58 to the representation, the membership of the segment 58 to the adaptive set, the direction of the locally increased quality that encodes the spatial non-uniformity of the scene into the representation to which the corresponding segment belongs, the quality used to encode the scene into the corresponding segment to which it belongs, etc.
[0132] The client can maintain a buffer for each representation, and after experiencing an increase in available throughput, decide to clear the remaining low-quality / bitrate media buffer before playback and download high-quality media segments with a duration into the existing low-quality / bitrate buffer. Similar embodiments can be constructed for tile-based streaming and viewport-dependent encoding.
[0133] The service operator may not be interested in understanding how much and what kind of data is downloaded without presenting it, because this introduces a cost with no gain on the server side and reduces the quality on the client side. Therefore, the present invention should provide a reporting metric that correlates the two events "media download" and "media presentation" for easy interpretation. The present invention avoids analyzing the information about the download and playback status of each media segment reported at length, and only allows the efficient reporting of clear events. The present invention also includes the identification of the buffer as described above and the association to the type. Embodiments of the present invention are given in Table 2.
[0134] Table 2: List of clear events
[0135]
[0136]
[0137] Figure 9 Another embodiment showing how the device 40 can be advantageously implemented Figure 9 The device 40 can correspond to any of the examples stated above with respect to Figure 1 to FIG. 8. That is, the device may include a log transmitter as discussed above with respect to Figure 8a but not necessarily, and can use the information 68 as discussed above with respect to Figure 2 and Figure 3 or the information 74 as discussed above with respect to Figures 5 to 7c but not necessarily. However, different from the description of Figure 2 to FIG. 8, with respect to Figure 9, assume that the tile-based streaming method is truly applied. That is, the scene content 30 is provided at the server 20 in a tile-based manner, and the tile-based manner is discussed as the option for FIGS. Figure 2 to 8 above.
[0138] Although the internal structure of the device 40 may be different from Figure 9 the internal structure depicted in, the device 40 is illustratively shown as including the selector 56 and the extractor 60 discussed above with respect to FIGS. Figure 2 to 8, and optionally includes the inferencer 66. However, additionally, the device 40 includes a media presentation description analyzer 90 and a matcher 92. The MPD analyzer 90 is used to derive from the media presentation description obtained from the server 20: at least one version, for tile-based streaming, the time-varying spatial scene 30 is provided in the at least one version; and for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming of the corresponding version of the time-varying spatial scene. The meaning of "version" will become clear from the following description. In particular, the matcher 92 matches the beneficial requirements so obtained with the device capabilities of the device 40 or another device that interacts with the device 40 (such as the decoding capabilities of one or more decoders 42, the number of decoders 42, or the like). Figure 9The background or idea implicit in the concept is as follows. Imagine, assume a particular size of the viewing section 28. Based on the tile-based approach, a particular number of tiles need to be included by the section 62. Additionally, it can be assumed that the media segments belonging to a particular tile form a media stream or video stream, and the stream should be decoded by a separate decoding instance, separately from decoding the media segments belonging to another tile. Thus, the movement aggregation of a particular number of tiles within the section 62 where the corresponding media segments are selected by the selector 56 requires a particular decoding capability, such as the existence of corresponding decoding resources in the form of, for example, a corresponding number of decoding instances (i.e., a corresponding number of decoders 42). If this number of decoders does not exist, the service provided by the server 20 is not available to the client. Thus, the MPD provided by the server 20 can indicate a "beneficial requirement", i.e., the number of decoders required to use the provided service. However, the server 20 can provide MPDs for different versions. That is, different MPDs for different versions can be obtained by the server 20, or the MPD provided by the server 20 can be structured internally to distinguish different versions that can be serviced and used. For example, the versions can differ in the field of view (i.e., the size of the viewing section 28). Different sizes of the field of view manifest themselves as different numbers of tiles within the section 62 and can thus differ in beneficial requirements because, for example, these versions can require different numbers of decoders. Other examples can also be imagined. For example, although versions with different fields of view can involve the same number of media segments 46, according to another example, the differences between different versions of the scenario 30 provided for tile-based streaming at the server 20 can even lie in the number of 46 media segments involved according to the corresponding version. For example, the tile segmentation according to one version is coarser compared to the tile-hyphenation segmentation of the scenario according to another version, thus requiring, for example, a smaller number of decoders.
[0139] The matcher 92 matches the beneficial requirements and thus selects the corresponding version or rejects all versions completely.
[0140] However, the beneficial requirements can additionally focus on the profiles / layers that one or more decoders 42 must be able to handle. For example, the DASH MPD includes multiple locations that allow for indicating profiles. A typical profile describes the attributes, elements that can be present in the MPD, and the video or audio profiles for each representation of the media stream provided.
[0141] Other examples of beneficial requirements concern, for example, the ability to move the viewport 28 across scenes on the client side. The beneficial requirement may indicate a required viewport speed that should be available to the user to move the viewport so that the user can truly enjoy the provided scene content. The matcher may check, for example, whether this requirement is met, e.g., inserted in a user input device such as the HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in a directional sense, the set of "adequate types of input devices" may be indicated by the beneficial requirement.
[0142] In the tile streaming service of spherical videos, there are a plethora of configuration parameters that can be set dynamically, such as the number of qualities, the number of tiles. In the case where tiles are independent bitstreams that need to be decoded by separate decoders, if the number of tiles is too high, a hardware device with several decoders will not be able to decode all the bitstreams simultaneously. The possibility is to keep this as a degree of freedom, and the DASH device parses all possible representations and counts how many decoders are needed to decode all the representations or the given number of the field of view of the covering device, and thus determines whether the DASH client is likely to consume the content. However, a smarter solution for interoperability and capability negotiation is to use signaling in the MPD mapped to a profile, which is used as a commitment to the client: if the profile is supported, the provided VR content can be consumed. This signaling should be in the form of a URN (such as urn::dash-mpeg::vr::2016) that can be encapsulated at the MPD level or at the adaptation set. This parsing will mean that N decoders at profile X are sufficient to consume the content. Depending on the profile, the DASH client may ignore or accept the MPD or parts of the MPD (adaptation sets). Additionally, there are several mechanisms that do not include all the information, such as Xlink or MPD links, where little signaling for selection is available. In this case, the DASH client will not be able to determine whether it can consume the content. It is necessary to expose the decoding capabilities regarding the number of decoders and the profile / level of each decoder through this urn (or something similar) so that the DASH client can now make sense of whether to perform an Xlink or MPD link or a similar mechanism. The signaling may also imply different operating points, such as N decoders with profile X / level or Z decoders with profile Y / level.
[0143] Figure 10 Further illustration, regarding the Figures 1 to 9 Any of the above-described embodiments and descriptions presented for the client, device 40, server, etc. can be extended to the following scope: the provided service is extended to a scope where the spatio-temporal scene that changes over time not only changes over time but also depends on another parameter. For example, Figure 10 Illustration Figure 1variant, where multiple available media segments are obtained on the server to describe the scene content 30 for different positions of the viewing center 100. In Figure 10 In the schematic diagram shown in, the scene center is depicted as varying only along one direction X, but obviously, the viewing center can vary along more than one spatial direction (such as two-dimensionally or three-dimensionally). For example, this corresponds to a change in the user's position in a specific virtual environment. Depending on the user's position in the virtual environment, the available view changes, and correspondingly, the scene 30 changes. Therefore, in addition to describing that the scene 30 is subdivided into tiles, time segments, and media segments of different qualities, other media segments describe different contents of the scene 30 for different positions of the scene center 100. The device 40 or the selector 56 calculates the addresses of the media segments to be extracted in the selection process from among the multiple 46 media segments respectively depending on the viewing section position and at least one parameter (such as parameter X), and can then use the calculated addresses to extract these media segments from the server. For this purpose, the media presentation description can describe a function depending on the tile index, quality index, scene center position, and time t, and generate the corresponding addresses of the corresponding media segments. Therefore, according to Figure 10 the embodiment of, the media presentation description can include this calculation rule, in addition to the parameters described above with respect to Figures 1 to 9 which, the calculation rule also depends on one or more additional parameters. Parameter X can be quantized to any level in the hierarchy, for which level the corresponding scene representation is encoded by the corresponding media segment within the multiple 46 media segments in the server 20.
[0144] As an alternative, X can be a parameter that defines the viewing depth (i.e., the distance radially from the scene center 100). Although providing scenes with different X values in the viewing center part allows the user to "walk" through the scene, providing scenes with different viewing depths can allow the user to "radially zoom" back and forth in the scene.
[0145] For multiple non-concentric viewports, the MPD can thus be signaled using another signal of the position of the current viewport. The signaling can be done at the segment, representation, or period level or the like.
[0146] Non-concentric spheres: To enable user movement, the spatial relationship of different spheres should be signaled in the MPD. This signaling can be done by coordinates (x, y, z) in any unit relative to the sphere diameter. Additionally, the diameter of each sphere should be indicated. The spheres can be "good enough" for users at their centers and in the additional space for which the content will be good. If the user can move beyond the signaled diameter, another sphere should be used for presenting the content.
[0147] Exemplary signaling of the viewport can be relative to a predefined center point in space. Each viewport will be signaled relative to that center point. In MPEG-DASH, this can be signaled (for example) in the AdaptationSet element.
[0148]
[0149] Finally, Figure 11 Information that describes information such as or similar to that described above with respect to reference numeral 74 can reside in video bitstream 110 into which video 112 is encoded. Decoder 114 that decodes this video 110 can use information 74 to determine the size of the focus region 116 within video 112, and the decoding capabilities for decoding video 110 should be concentrated on that focus region. For example, information 74 can be conveyed within the SEI information of video bitstream 110. For example, the focus region can be decoded specifically, or decoder 114 can be configured to start decoding each image of the video at the focus region rather than (for example) at the upper left image corner, and / or decoder 114 can stop decoding each image of the video after the focus region 116 has been decoded. Additionally or alternatively, information 74 can be present in the data stream for use only in forwarding to subsequent renderers or viewport controls or streaming media devices of the client or segment selector for determining which segments to download or stream in order to cover a spatial section completely or with increased or predetermined quality. For example, as outlined above, information 74 indicates a recommended preferred region as a recommendation for placing viewing segment 62 or segment 66 to coincide with or cover or track this region. Information 74 can be used by the segment selector of the client. Just as is true with respect to the description of Figures 4 to 7c Information 74 can absolutely set the size of region 116 (such as in terms of the number of tiles), or can set, for example, the speed of region 116 used for region movement according to user input, etc., in order to follow the content of interest of the video spatiotemporally, thereby scaling region 116 to increase as the indication of speed increases.
[0150] A first aspect of the present application provides an apparatus for streaming media content regarding a spatially varying scene (30) over time.
[0151] (A1) The apparatus is configured to:
[0152] select media segments from among a plurality of media segments available on a server,
[0153] wherein the selected media segments are from the server, and wherein the apparatus is configured to:
[0154] Perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene with the time variation that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and
[0155] derive the predetermined relationship from information included in the selected media segment and / or from a signal received from the server.
[0156] (A2), The apparatus according to (A1), each of the plurality of media segments has an associated spatio-temporal part of the spatially varying scene with the time variation encoded therein at an associated quality level within a set of quality levels.
[0157] (A3), The apparatus according to (A1), wherein each of the spatio-temporal parts of the spatially varying scene (30) encoded into the plurality of media segments is a time segment of the spatially varying scene at a corresponding one of tiles (50) into which the spatially varying scene is spatially subdivided.
[0158] (A4), The apparatus according to any one of (A1)-(A3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
[0159] (A5), The apparatus according to (A4), wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the viewing section.
[0160] (A6), The apparatus according to (A4) or (A5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the viewing section by a list of pairs of the respective distances from the viewing section and the corresponding tolerable values for the measure of the difference beyond the respective distances.
[0161] (A7), The apparatus according to any one of (A4)-(A6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
[0162] (A8), The apparatus according to any one of (A4)-(A7), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval during which the second part can be together with the first part within the observation section.
[0163] (A9), The apparatus according to any one of (A8), wherein the information indicates another tolerable value for a measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval during which the second part can be together with the first part within the observation section.
[0164] (A10), The apparatus according to any one of (A4)-(A9), wherein the information is time-varying and / or spatially varying.
[0165] (A11), The apparatus according to any one of (A1)-(A10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0166] (A12), The apparatus according to any one of (A1)-(A11), wherein the apparatus is configured to perform the selection such that the first part follows a time-varying observation section of the time-varying spatial scene.
[0167] (A13), The apparatus according to (A12), wherein the apparatus is configured to change the spatial position of the time-varying observation section according to a user input.
[0168] (A14), The apparatus according to any one of (A1)-(A13), wherein the apparatus is configured to determine the first part to correspond to a region of interest.
[0169] (A15), The apparatus according to any one of (A14), wherein the apparatus is configured to extract information about the region of interest from the server.
[0170] The present invention also provides a streaming media server for media content of a time-varying spatial scene.
[0171] (B1) The streaming media server is configured to:
[0172] Making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the media segment having at least a spatial section of the spatially varying scene encoded therein with the time variation, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and
[0173] signaling information about a predetermined relationship in the media segment and / or by signaling to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
[0174] (B2), The streaming media server according to (B1), wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying scene encoded therein at an associated quality level within a set of quality levels.
[0175] (B3), The streaming media server according to (B2), wherein each of the spatio-temporal portions of the spatially varying scene encoded into the plurality of media segments is a time segment of the spatially varying scene at a corresponding one of the tiles into which the spatially varying scene is spatially subdivided.
[0176] (B4), The streaming media server according to any one of (B1)-(B3), wherein the information indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
[0177] (B5), The streaming media server according to (B4), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from an observation section.
[0178] (B6), The streaming media server according to (B4) or (B5), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section by a list of pairs of the respective distance from the observation section and the corresponding tolerable value for the measure of the difference beyond the respective distance.
[0179] (B7), The streaming media server according to any one of (B4)-(B6), wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.
[0180] (B8), a streaming media server according to any one of (B4)-(B7), wherein the information indicates a tolerable value for a measure of a difference between the other quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second portion can be together with the first portion within the observation section.
[0181] (B9), a streaming media server according to (B8), wherein the information indicates another tolerable value for a measure of a difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval within which the second portion can be together with the first portion within the observation section.
[0182] (B10), a streaming media server according to any one of (B4)-(B9), wherein the information is time-varying and / or spatially varying.
[0183] (B11), a streaming media server according to any one of (B4)-(B10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0184] (B12), a streaming media server according to any one of (B4)-(B11), wherein the server is configured to send information about an area of interest to the device.
[0185] This application also provides a media presentation description.
[0186] (C1), the media presentation description includes:
[0187] Information about calculating addresses of a plurality of media segments such that a device can use the information to select and extract media segments from the plurality of media segments, wherein at least a spatial section of the spatially varying scene over time is encoded into the media segments in such a way that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and
[0188] Information about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.
[0189] (C2), the media presentation description according to (C1), wherein each media segment of the plurality of media segments has an associated spatio-temporal portion of the spatially varying scene over time encoded therein at an associated quality level within a set of quality levels.
[0190] (C3), the media presentation description according to (C2), wherein each of the spatio-temporal parts of the spatially varying scene that varies over time and is encoded into the plurality of media segments is a temporal segment of the spatially varying scene that varies over time at a corresponding one of the tiles into which the spatially varying scene that varies over time is spatially subdivided.
[0191] (C4), the media presentation description according to any one of (C1)-(C3), wherein the information indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.
[0192] (C5), the media presentation description according to (C4), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner that depends on the distance from the viewing section.
[0193] (C6), the media presentation description according to (C4) or (C5), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner that depends on the distance from the viewing section by means of a list of pairs of the corresponding distance from the viewing section and the corresponding tolerable value for the measure of the difference beyond the corresponding distance.
[0194] (C7), the media presentation description according to any one of (C4)-(C6), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
[0195] (C8), the media presentation description according to any one of (C4)-(C7), wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.
[0196] (C9), the media presentation description according to (C8), wherein the information indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of another maximum allowable time interval during which the second part can be within the viewing section together with the first part.
[0197] (C10), the media presentation description according to any one of (C4)-(C9), wherein the information is time-varying and / or spatially varying.
[0198] (C11), a media presentation description according to any one of (C4)-(C10), wherein the information indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
[0199] (C12), a media presentation description according to any one of (C4)-(C11), wherein the media presentation description includes information about a region of interest.
[0200] This application also provides a device for streaming media content of a spatial scene changing over time.
[0201] (D1) The device is configured to:
[0202] select media segments from among a plurality of media segments available on a server,
[0203] wherein the device is configured to:
[0204] perform the selection such that the selected media segments have at least a spatial section of the spatial scene changing over time encoded therein, encoded in such a way that:
[0205] a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatial scene changing over time that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and
[0206] such that the first part follows an observation section of the time change of the spatial scene changing over time, and
[0207] set the size and / or position of the first part depending on information included in the selected media segment and / or a signal received from the server.
[0208] (D2), the device according to (D1), wherein the information indicates the size in the form of an increment relative to the size of the observation section of the time change or a scaling of the size of the observation section of the time change.
[0209] (D3), the device according to (D1) or (D2), wherein the information indicates a predetermined value for a measure of the spatial speed of the observation section.
[0210] (D4), the device according to (D3), wherein the information indicates the predetermined value for the measure of the spatial speed of the observation section for:
[0211] a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or
[0212] A percentage value indicating the percentage of users whose spatial speed of the viewing section does not exceed the predetermined value, and / or
[0213] A percentage value indicating the percentage of users for whom the viewing section is in a predetermined area, and / or
[0214] A hint indicating one or more types of user input for controlling the movement of the viewing section to which the predetermined value is applicable.
[0215] (D5), the apparatus according to (D3) or (D4), the apparatus being configured to perform the setting such that
[0216] The higher the predetermined value d of the measure of the spatial speed for the viewing section, the larger the size.
[0217] (D6), the apparatus according to any one of (D1) or (D5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement for the viewing section.
[0218] (D7), the apparatus according to (D6), the apparatus being configured to perform the setting such that
[0219] The higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.
[0220] (D8), the apparatus according to any one of (D1) or (D7), wherein the information indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.
[0221] (D9), the apparatus according to any one of (D1) or (D8), the apparatus being configured such that the first part coincides with the spatial section.
[0222] (D10), the apparatus according to any one of (D1) or (D9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.
[0223] (D11), the apparatus according to (D10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level within a set of quality levels.
[0224] (D12), the apparatus according to (D1), wherein the apparatus is configured to set the size in a manner independent of the size of the observation section that varies with time depending on the information.
[0225] (D13), the apparatus according to (D1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the apparatus uses the values included in the information for the size option that fits the actual size of the time-varying observation section.
[0226] (D14), the apparatus according to (D12) or (D13), wherein the information indicates the size in terms of the number of tiles.
[0227] This application also provides a streaming media server for media content of a spatial scene that varies over time.
[0228] (E1), the streaming media server is configured to:
[0229] Make a plurality of media segments available for extraction by a device, so that the device can select the media segments for extraction, and the media segments have at least a spatial section of the spatial scene that varies over time encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatial scene that varies over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment with another quality that is reduced relative to the predetermined quality, and such that the first part follows the time-varying observation section of the spatial scene that varies over time, and
[0230] Signal information in the media segment and / or by signaling to the device regarding how to set the size and / or position of the first part.
[0231] This application also provides a signal defining a media presentation description.
[0232] (F1) The signal includes:
[0233] Information regarding calculating addresses of a plurality of media segments such that a device can, using the information, select and extract a media segment from the plurality of media segments, the media segment having at least a spatial section of the spatially varying scene that varies over time encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that varies over time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first part follows an observation section that varies over time of the spatially varying scene, and
[0234] Information in the media segment and / or signaled to the device regarding how to set the size and / or position of the first part.
[0235] (F2), according to the signal of (F1), wherein the information indicates the size in the form of an increment relative to the size of the observation section that varies over time or a scaling of the size of the observation section that varies over time.
[0236] (F3), according to the signal of (F1) or (F2), wherein the information indicates a predetermined value of a measure of the spatial speed for the observation section.
[0237] (F4), according to the signal of (F3), wherein the information indicates the predetermined value of the measure of the spatial speed for the observation section for:
[0238] a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or
[0239] a percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or
[0240] a percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or
[0241] a hint of one or more types of user input for controlling movement of the observation section for which the predetermined value is applicable.
[0242] (F5), according to the signal of (F3) or (F4), the signal being configured to perform the setting such that
[0243] wherein the higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.
[0244] (F6), according to the signal of any one of (F1)-(F5), wherein the information indicates a predetermined value of a measure of the probability of the direction of movement for the observation section.
[0245] (F7), according to the signal of (F6), the signal is configured to perform the setting such that
[0246] The higher the probability of the corresponding moving direction, the more the first part extends into the corresponding direction.
[0247] (F8), according to the signals of (F1)-(F7), wherein the information indicates the predetermined spatial velocity in a time-varying and / or space-varying and / or direction-varying manner.
[0248] (F9), according to the signal of any one of (F1) or (F8), is configured such that the first part coincides with the spatial section.
[0249] (F10), according to the signal of any one of (F1) or (F9), wherein each of the spatio-temporal parts of the time-varying spatial scene encoded into the plurality of media segments is a time segment of the time-varying spatial scene at a corresponding one of the tiles into which the time-varying spatial scene is spatially subdivided.
[0250] (F11), according to the signal of (F10), wherein each of the plurality of media segments has an associated spatio-temporal part of the time-varying spatial scene encoded therein at an associated quality level among a set of quality levels.
[0251] (F12), according to the signal of (F11), wherein the device is configured to set the size in a manner independent of the size of the time-varying observation section depending on the information.
[0252] (F13), according to the signal of (F1), wherein the information includes different values of the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.
[0253] (F14), according to the signal of (F11) or (F12), wherein the information indicates the size in terms of the number of tiles.
[0254] This application also provides a video bitstream.
[0255] (G1) having video encoded therein, the video bitstream includes a signal function for one or more of the size of the focused area within the video to which the decoding capability for decoding the video should be concentrated and the recommended preferred observation section area of the video.
[0256] The present application also provides a decoder for decoding video from a video bitstream.
[0257] (H1) The decoder is configured to:
[0258] derive a signal indicative of the size of a focus region within the video from the video bitstream, and
[0259] concentrate decoding capabilities for decoding the video to the focus region.
[0260] (H2), the decoder according to (H1), the decoder is configured to specifically decode the focus region.
[0261] (H3), the decoder according to (H1), the decoder is configured to start decoding each image of the video at the focus region.
[0262] (H4), the decoder according to (H1), the decoder is configured to stop decoding each image of the video after decoding the focus region.
[0263] (H5), the decoder according to any one of (H1)-(H4), wherein the signal absolutely indicates the size, or the decoder is configured to scale the size of the focus region with a parameter included in the signal.
[0264] The present application also provides a device for streaming media content regarding a spatially varying scene over time.
[0265] (I1) The device is configured to:
[0266] derive from a media presentation description:
[0267] at least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version,
[0268] for each of the at least one version, an indication of the beneficial requirements for benefiting from the tile-based streaming the corresponding version of the spatially varying scene over time,
[0269] match the beneficial requirements of the at least one version with the device capabilities of the device or another device that interacts with the device regarding the tile-based streaming.
[0270] (I2), the device according to (I1), wherein the beneficial requirements and the device capabilities relate to decoding capabilities.
[0271] (I3), the apparatus according to (I1) or (I2), wherein the beneficial requirements and the apparatus capabilities are related to the number of available decoders.
[0272] (I4), the apparatus according to any one of (I1)-(I3), wherein the beneficial requirements and the apparatus capabilities are related to the layer and / or profile descriptor.
[0273] (I5), the apparatus according to any one of (I1)-(I4), wherein the beneficial requirements and the apparatus capabilities are related to the type of input device for moving the viewing section across the spatially varying scene over time, or to the speed of moving the viewing section across the spatially varying scene over time using the input device.
[0274] (I6), the apparatus according to any one of (I1)-(I5), wherein the apparatus is configured to:
[0275] select a media segment from a plurality of media segments available on the server, the selection being made by calculating the address of the selected media segment using a calculation rule included in the media presentation description,
[0276] extract the selected media segment from the server using the calculated address,
[0277] wherein the apparatus is configured to perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene varying over time encoded therein, encoded in such a way that:
[0278] a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene varying over time that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced with respect to the predetermined quality, and
[0279] such that the first part follows the viewing section of the spatially varying scene varying over time.
[0280] This application also provides a streaming media server for streaming media content regarding a spatially varying scene varying over time.
[0281] (J1) The streaming media server is configured to:
[0282] provide a media presentation description from which
[0283] at least one version can be obtained, for tile-based streaming, the spatially varying scene varying over time is provided in the at least one version,
[0284] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene that benefits from the tile-based streaming with respect to the time variation,
[0285] such that the apparatus for streaming the media content from the streaming server can
[0286] match the beneficial requirements of the at least one version with the apparatus capabilities of the apparatus or another apparatus that interacts with the apparatus with respect to the tile-based streaming.
[0287] This application also provides a media presentation description.
[0288] (K1) The media presentation description includes:
[0289] information about at least one version, for tile-based streaming, the spatially varying scene with time variation is provided in the at least one version,
[0290] For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene that benefits from the tile-based streaming with respect to the time variation.
[0291] (K2), the media presentation description according to (K1), wherein the beneficial requirements and the apparatus capabilities relate to decoding capabilities.
[0292] (K3), the media presentation description according to (K1) or (K2), wherein the beneficial requirements and the apparatus capabilities relate to the number of available decoders.
[0293] (K4), the media presentation description according to any one of (K1)-(K3), wherein the beneficial requirements and the apparatus capabilities relate to layer and / or profile descriptors.
[0294] (K5), the media presentation description according to any one of (K1)-(K4), wherein the beneficial requirements and the apparatus capabilities relate to the type of input device for moving the viewing section across the spatially varying scene with time variation, or to the speed of moving the viewing section across the spatially varying scene with time variation using the input device.
[0295] (K6), the media presentation description according to any one of (K1)-(K5), further includes a calculation rule, using which the apparatus can
[0296] select a media segment from among a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rule.
[0297] The present application also provides an apparatus for streaming media content of a spatial scene varying over time.
[0298] (L1) The apparatus is configured to:
[0299] calculate an address of a media segment depending on a spatial viewport position and at least one parameter, the media segment depicting a spatial scene varying in time and in the at least one parameter,
[0300] extract the media segment using the calculated address.
[0301] (L2) The apparatus according to claim (L1), wherein the at least one parameter includes one or more coordinates of a center of observation and / or an observation depth.
[0302] The present application also provides a media presentation description.
[0303] (M1) The media presentation description includes:
[0304] a calculation rule for calculating an address of a media segment depending on a spatial viewport position and at least one parameter, so as to extract the media segment using the calculated address, the media segment depicting a spatial scene varying in time and in the at least one parameter.
[0305] The present application also provides a streaming server for allowing an apparatus to stream media content of a spatial scene varying over time from a server.
[0306] (N1) The streaming server is configured to provide the media presentation description according to (M1).
[0307] The present application also provides an apparatus for streaming media content of a spatial scene varying over time.
[0308] (O1) The apparatus is configured to:
[0309] select a media segment from a plurality of media segments available on a server,
[0310] wherein the apparatus is configured to:
[0311] encode a first portion of the spatial scene varying over time into the selected media segment with increased quality compared to a spatial neighborhood of the first portion or in such a way that the spatial neighborhood of the first portion is not encoded into the selected media segment,
[0312] emit a log message that records:
[0313] an instantaneous measurement result of measuring a spatial position and / or movement of the first portion; and / or
[0314] Measure statistical values of the spatial position and / or movement of the first part, such as time average values; and / or
[0315] Measure instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or
[0316] Indication of the set of buffers of the device involved in buffering the selected media segment, description of the distribution rules applied in distributing the selected media segment into the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or
[0317] Measurement results of the amount of the selected media segment that has not been output from the buffer of the device for undergoing decoding; and / or
[0318] Measure statistical values of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or
[0319] Measure the instantaneous measurement results of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or
[0320] Measure statistical values of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or
[0321] The field of view covered by the observation section; and / or
[0322] Measure instantaneous measurement results of the user position or the observation depth relative to the scene center; and / or
[0323] Measure statistical values of the user position or the observation depth relative to the scene center, such as time average values.
[0324] (O2), The device according to (O1), wherein the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.
[0325] (O3), The device according to (O1) or (O2), configured to perform the selection such that the first part of the spatial scene of the time variation pins the observation section.
[0326] (O4), The apparatus according to any one of (O1)-(O3), configured to emit a log message that records the instantaneous measurement of the quality of the spatial scene that measures the time variation up to being encoded into the selected media segment and up to being visible in the observation section as one of the following:
[0327] The measurement result of the average density of the pixels falling within the observation section, where the spatial scene with the time variation is encoded into the selected media segment at the average density.
[0328] (O5), The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0329] (O6), The apparatus according to (O4), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0330] (O7), The apparatus according to (O4), configured to make the emitted log message indicate whether the measurement result measures the average density of the pixels by:
[0331] Averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment, or
[0332] Averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
[0333] (O8), The apparatus according to (O6) or (O7), wherein averaging the pixel density in a spatially non-uniform manner corresponds to
[0334] Averaging in a spherically uniform manner, or
[0335] Averaging uniformly in the viewport plane space, where the viewport plane is perpendicular to the central viewing direction of the observation section.
[0336] (O9), The apparatus according to any one of (O4)-(O8), configured to measure the average density of the pixels such that the measurement result averages the pixel density in a way that limits the averaging to the central sub-section of the observation section, or applies a higher averaging weight to the central sub-section compared to the edge part of the observation section around the central sub-section.
[0337] (O10) The apparatus according to any one of (O4)-(O9) is configured such that the measurement results measure the average density of the pixels separately along the horizontal observation section axis and the vertical observation section axis, respectively.
[0338] (O11) The apparatus according to any one of (O4)-(O10) is configured to intermittently emit log messages.
[0339] (O12) The apparatus according to any one of (O4)-(O11) is configured to emit log messages at a rate controlled by a manifest file, and the apparatus performs the selection of the media segments for download based on the manifest file.
[0340] (O13) The apparatus according to any one of (O4)-(O12), wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, and the representations differ in one or more of the following:
[0341] The scene segment of the time-varying spatial scene encoded into the media segment,
[0342] The quality used to encode the time-varying spatial scene into the media segment,
[0343] The spatial quality variation used to encode the time-varying spatial scene into the media segment,
[0344] wherein the apparatus is configured to emit log messages, and the log messages are recorded in the form of an association of each buffer with one or a combination of two or more of the following in the description of the distribution rule applied in distributing the selected media segments to the set of buffers:
[0345] The scene segment,
[0346] The quality,
[0347] The spatial quality distribution,
[0348] Representation.
[0349] (O14) The apparatus according to any one of (O4)-(O13), wherein the representations are grouped into adaptive sets according to one or more of the following:
[0350] The scene segment of the time-varying spatial scene encoded into the media segment,
[0351] The spatial quality variation used to encode the time-varying spatial scene into the media segment,
[0352] wherein the device is configured to emit a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, in the form of an association of each buffer with one of the adaptation sets, or in the form of an association of each buffer with one of the representations.
[0353] (O15), The device according to any one of (O4)-(O14), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in the form of a time measurement result.
[0354] (O16), The device according to (O15), wherein the device is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in the form of a time measurement result in time units less than the time length of the media segment and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.
[0355] (O17), The device according to any one of (O4)-(O16), wherein the device is configured to emit a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffer of the device and is to undergo decoding, in a format classified by one or more of the following:
[0356] the buffer of the decoder in which the corresponding media segment has been buffered,
[0357] the scene segment encoded into the corresponding media segment,
[0358] the quality used to encode the spatially varying scene over time into the corresponding media segment,
[0359] the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.
[0360] This application also provides a method for streaming media content regarding a spatially varying scene over time.
[0361] (P1) The method includes:
[0362] selecting a media segment from a plurality of media segments available on a server,
[0363] Extract the selected media segment from the server, perform the selection such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and
[0364] The method further includes deriving the predetermined relationship from information included in the selected media segment and / or from a signal obtained from the server.
[0365] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.
[0366] (Q1) The method includes:
[0367] Make a plurality of media segments available for extraction by a device, so that the device can select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is encoded into the selected media segment at another quality, and
[0368] Signal information regarding the predetermined relationship in the media segment and / or by signal action to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
[0369] This application also provides a method for streaming media content regarding a spatially varying scene with time variation.
[0370] (R1) The method includes:
[0371] Select a media segment from a plurality of media segments available on the server,
[0372] Extract the selected media segment from the server, where the selection is performed such that the selected media segment has at least a spatial section of the spatially varying scene with the time variation encoded therein, encoded in such a way that
[0373] A first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced with respect to the predetermined quality, and
[0374] An observation section that causes the first part to follow the temporal variations of the spatio-temporal scene that varies with time, and
[0375] The method further includes setting the size of the first part depending on information included in the selected media segment and / or a signal action obtained from the server.
[0376] This application also provides a method for streaming media content regarding a spatio-temporal scene that varies with time.
[0377] (S1) The method includes:
[0378] Making a plurality of media segments available for extraction by a device, so that the device can select a media segment for extraction, the selected media segment having at least a spatial section of the spatio-temporal scene that varies with time encoded therein, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part of the spatio-temporal scene that varies with time and is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and an observation section that causes the first part to follow the temporal variations of the spatio-temporal scene that varies with time, and
[0379] Signaling information in the media segment and / or by signal action to the device regarding how to set the size of the first part.
[0380] This application also provides a method for decoding video from a video bitstream.
[0381] (T1) The method includes:
[0382] Deriving a signal action of the size of a focused area within the video from the video bitstream, and
[0383] Concentrating the decoding capability for decoding the video on the focused area.
[0384] This application also provides a method for streaming media content regarding a spatio-temporal scene that varies with time, performed by a device.
[0385] (U1) The method includes:
[0386] Deriving from a media presentation description:
[0387] At least one version, for tile-based streaming, the spatio-temporal scene that varies with time is provided in the at least one version,
[0388] For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatial scene that benefits from the time variation for tile - based streaming
[0389] Match the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding tile - based streaming.
[0390] This application also provides a method for streaming media content of a spatial scene with time variation.
[0391] (V1) The method includes:
[0392] Select media segments from a plurality of media segments available on a server,
[0393] Extract the selected media segments from the server, where the selection is performed such that the selected media segments have an increased quality compared to the spatial neighborhood of a first part or are encoded in a way that the spatial neighborhood of the first part is not encoded into the selected media segments, for the first part of the spatial scene with time variation
[0394] The method further includes emitting a log message that records:
[0395] An instantaneous measurement of the spatial position and / or movement of the first part; and / or
[0396] A statistical value of the spatial position and / or movement of the first part, such as a time average; and / or
[0397] An instantaneous measurement of the quality of the spatial scene with time variation until it is encoded into the selected media segments and until it is visible in an observation section; and / or
[0398] An indication of the set of buffers of the device involved in buffering the selected media segments, a description of the distribution rules applied in distributing the selected media segments to the set of buffers, and the instantaneous buffer fullness of each of the set of buffers; and / or
[0399] A measurement of the amount of the selected media segments yet to be output from the device's buffer for undergoing decoding; and / or
[0400] A statistical value of the quality of the spatial scene with time variation until it is encoded into the selected media segments and until it is visible in an observation section, such as a time average; and / or
[0401] An instantaneous measurement of the quality of the first part or of the spatial scene up to the time variation visible in the observation section and encoded into the selected media segment; and / or
[0402] A statistical value of the quality of the first part or of the spatial scene up to the time variation visible in the observation section and encoded into the selected media segment, such as a time average; and / or
[0403] The field of view covered by the observation section; and / or
[0404] An instantaneous measurement of the user position or the observation depth relative to the scene center; and / or
[0405] A statistical value of the user position or the observation depth relative to the scene center, such as a time average.
[0406] This application also provides a method for streaming media content of a spatial scene varying over time.
[0407] (W1) The method includes:
[0408] Providing a media presentation description from which can be derived:
[0409] At least one version, for tile - based streaming, the spatial scene varying over time is provided in the at least one version,
[0410] For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatial scene varying over time for benefiting from the tile - based streaming,
[0411] Thereby enabling a device for streaming the media content from a streaming server to
[0412] Match the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile - based streaming.
[0413] This application also provides a method for allowing a device to stream media content of a spatial scene varying over time from a server.
[0414] (X1) The method includes providing the media presentation description as described in (M1).
[0415] This application also provides a computer program having program code for performing the method according to any one of claims (P1)-(X1) when the program is executed on a computer.
[0416] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, or an object, or a feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0417] The signals mentioned above (such as, a streamed signal, an MPD, or any other of the mentioned signals) may be stored on a digital storage medium, or may be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet.
[0418] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or in software. The implementation may be performed using a digital storage medium storing an electronically readable control signal, such as a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, the electronically readable control signal cooperating (or being capable of cooperating) with a programmable computer system to cause the execution of a corresponding method. Thus, the digital storage medium may be computer-readable.
[0419] Some embodiments according to the present invention include a data carrier having an electronically readable control signal capable of cooperating with a programmable computer system to cause the execution of one of the methods described herein.
[0420] In general, embodiments of the present invention may be implemented as a computer program product having program code that, when the computer program product is executed on a computer, is operable to execute one of the methods. The program code may, for example, be stored on a machine-readable carrier.
[0421] Other embodiments include a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0422] In other words, thus, an embodiment of the inventive method is a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0423] Thus, another embodiment of the method of the present invention is a data carrier (or a digital storage medium, or a computer-readable medium) that includes a computer program recorded thereon for executing one of the methods described herein. The data carrier, the digital storage medium, or the recorded medium is generally tangible and / or non-transitory.
[0424] Accordingly, another embodiment of the method of the present invention is a data stream or signal sequence that represents a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured (e.g.) to be transmitted via a data communication connection (e.g., via the Internet).
[0425] Another embodiment includes a processing component, such as a computer or a programmable logic device, that is configured or adapted to perform one of the methods described herein.
[0426] Another embodiment includes a computer on which is installed a computer program for performing one of the methods described herein.
[0427] Another embodiment according to the present invention includes an apparatus or system configured (e.g., electrically or optically) to transmit a computer program for performing one of the methods described herein to a receiver. The receiver can be (e.g.) a computer, a mobile device, a storage device, etc. The apparatus or system can include (e.g.) a file server for transmitting the computer program to the receiver.
[0428] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.
[0429] The apparatuses described herein can be implemented using hardware devices or using computers or using a combination of hardware devices and computers.
[0430] The apparatuses described herein or any components of the apparatuses described herein can be implemented at least in part in hardware and / or in software.
[0431] The methods described herein can be performed using hardware devices or using computers or using a combination of hardware devices and computers.
[0432] The methods described herein or any components of the apparatuses described herein can be performed at least in part by hardware and / or by software.
[0433] The above embodiments merely illustrate the principles of the present invention. It should be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is only intended to be limited by the scope of the appended patent claims, rather than by the specific details presented by the description and explanation of the embodiments herein.
Claims
1. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select (56) media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the selected media segments are from the server (20), and wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) varying over time encoded therein, encoded in such a way that a first part (64) of the spatial section is encoded into the selected media segment at a predetermined quality, and a second part (66) of the spatially varying scene varying over time that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and derive the predetermined relationship from information (68) included in the selected media segment and / or from a signal obtained from the server (20).
2. The apparatus according to claim 1, wherein each of the plurality (46) of media segments has an associated spatio-temporal part of the spatially varying scene (30) varying over time encoded therein at an associated quality level within a set of quality levels.
3. The apparatus according to claim 2, wherein the spatio-temporal part of the spatially varying scene (30) encoded into each of the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of tiles (50) into which the spatially varying scene (30) is spatially subdivided.
4. The apparatus according to any one of claims 1 to 3, wherein the information (68) indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
5. The apparatus according to claim 4, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from an observation section (28).
6. The apparatus according to claim 4 or 5, wherein the information indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality in a manner depending on the distance from the observation section (28) by a list of pairs of the respective distance from the observation section (28) and the corresponding tolerable value for the measure of the difference beyond the respective distance.
7. The apparatus according to any one of claims 4 to 6, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality such that the tolerable value increases as the distance from the observation section increases.
8. The apparatus according to any one of claims 4 to 7, wherein the information (68) indicates the tolerable value for the measure of the difference between the another quality and the predetermined quality, and an indication of a maximum allowable time interval within which the second part can be together with the first part within the observation section.
9. The apparatus according to claim 8, wherein the information (68) indicates another tolerable value for a measure of the difference between the another quality and the predetermined quality, and the second portion is capable of indicating another maximum allowable time interval within the observation section together with the first portion.
10. The apparatus according to any one of claims 4 to 9, wherein the information (68) is time-varying and / or spatially varying.
11. The apparatus according to any one of claims 1 to 10, wherein the information (68) indicates an allowable concurrent setting pair for the another quality and the predetermined quality.
12. The apparatus according to any one of claims 1 to 11, wherein the apparatus is configured to perform the selection such that the first portion (64) follows a time-varying observation section (28) of the time-varying spatial scene (30).
13. The apparatus according to claim 12, wherein the apparatus is configured such that a spatial position of the time-varying observation section (28) is changed according to a user input.
14. The apparatus according to any one of claims 1 to 13, wherein the apparatus is configured to determine the first portion (64) to correspond to a region of interest.
15. The apparatus according to claim 14, wherein the apparatus is configured to extract information about the region of interest from the server.
16. A streaming media server for media content of a time-varying spatial scene, configured to: make a plurality of media segments available for extraction by an apparatus, so that the apparatus can select media segments for extraction, the media segments having at least a spatial section of the time-varying spatial scene encoded therein, encoded in such a way that a first portion of the spatial section is encoded into a selected media segment at a predetermined quality, and a second portion of the time-varying spatial scene spatially adjacent to the first portion is encoded into the selected media segment at another quality, and signal information about a predetermined relationship in the media segment and / or by means of a signal to the apparatus, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
17. The streaming media server according to claim 16, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the time-varying spatial scene (30) encoded therein at an associated quality level within a set of quality levels.
18. The streaming media server according to claim 17, wherein each of the spatio-temporal portions of the time-varying spatial scene (30) encoded into the plurality (46) of media segments is a time segment (54) of the time-varying spatial scene (30) at a corresponding one of tiles (50) into which the time-varying spatial scene (30) is spatially subdivided.
19. The streaming media server according to any one of claims 16 to 18, wherein the information (68) indicates a tolerable value for a measure of the difference between the another quality and the predetermined quality.
20. The streaming media server according to claim 19, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).
21. The streaming media server according to claim 19 or 20, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by means of a list of pairs of the respective distances from the viewing section (28) and the corresponding tolerable values for the measure of the difference beyond the respective distances.
22. The streaming media server according to any one of claims 19 to 21, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
23. The streaming media server according to any one of claims 19 to 22, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.
24. The streaming media server according to claim 23, wherein the information (68) indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of the maximum allowable time interval during which the second part can be within the viewing section together with the first part.
25. The streaming media server according to any one of claims 19 to 24, wherein the information (68) is time-varying and / or spatially varying.
26. The streaming media server according to any one of claims 19 to 25, wherein the information (68) indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
27. The streaming media server according to any one of claims 19 to 26, wherein the server is configured to send information about the region of interest to the device.
28. A media presentation description, comprising: information about calculating the addresses of a plurality of media segments such that a device can use the information to select and extract media segments from the plurality of media segments, wherein at least a spatial section of the spatially varying scene that changes over time is encoded into the media segments in such a way that a first part of the spatial section is encoded into the selected media segment with a predetermined quality, and a second part of the spatially varying scene that changes over time and is spatially adjacent to the first part is encoded into the selected media segment with another quality, and information (68) about a predetermined relationship, wherein the other quality satisfies the predetermined relationship with respect to the predetermined quality.
29. The media presentation description according to claim 28, wherein each of the plurality (46) of media segments has an associated spatio-temporal portion of the spatially varying scene (30) with time variation encoded therein at an associated quality level within a set of quality levels.
30. The media presentation description according to claim 29, wherein each of the spatio-temporal portions of the spatially varying scene (30) with time variation encoded in the plurality (46) of media segments is a time segment (54) of the spatially varying scene (30) at a respective one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.
31. The media presentation description according to any one of claims 28 to 30, wherein the information (68) indicates a tolerable value for a measure of the difference between the other quality and the predetermined quality.
32. The media presentation description according to claim 31, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28).
33. The media presentation description according to claim 31 or 32, wherein the information indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality in a manner depending on the distance from the viewing section (28) by a list of pairs of the respective distances from the viewing section (28) and the corresponding tolerable values for the measure of the difference beyond the respective distances.
34. The media presentation description according to any one of claims 31 to 33, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality such that the tolerable value increases as the distance from the viewing section increases.
35. The media presentation description according to any one of claims 31 to 34, wherein the information (68) indicates the tolerable value for the measure of the difference between the other quality and the predetermined quality, and an indication of the maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
36. The media presentation description according to claim 35, wherein the information (68) indicates another tolerable value for the measure of the difference between the other quality and the predetermined quality, and another indication of the maximum allowable time interval within which the second portion can be together with the first portion within the viewing section.
37. The media presentation description according to any one of claims 31 to 36, wherein the information (68) is time-varying and / or spatially varying.
38. The media presentation description according to any one of claims 31 to 37, wherein the information (68) indicates an allowed concurrent setting pair for the other quality and the predetermined quality.
39. The media presentation description according to any one of claims 31 to 38, wherein the media presentation description includes information about a region of interest.
40. An apparatus for streaming media content of a spatially varying scene (30) over time, configured to: select media segments from among a plurality (46) of media segments (58) available on a server (20), wherein the apparatus is configured to: perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) encoded therein that varies over time, encoded such that: a first part (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second part (66; 72) of the spatially varying scene that is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and such that the first part (64) follows an observation section (28) of the spatially varying scene (30) that varies over time, and set the size and / or position of the first part (64) depending on information (74) included in the selected media segment and / or a signal received from the server.
41. The apparatus according to claim 40, wherein the information (74) indicates the size in the form of an increment relative to the size of the observation section that varies over time or a scaling of the size of the observation section that varies over time.
42. The apparatus according to claim 40 or 41, wherein the information (74) indicates a predetermined value of a measure of the spatial speed for the observation section.
43. The apparatus according to claim 42, wherein the information (74) indicates the predetermined value of the measure of the spatial speed for the observation section for: a default percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the measured spatial speed does not exceed the predetermined value, and / or a percentage value indicating the percentage of users for whom the observation section is in a predetermined region, and / or a hint indicating one or more types of user input for controlling the movement of the observation section to which the predetermined value is applicable.
44. The apparatus according to claim 42 or 43, the apparatus being configured to perform the setting such that the higher the predetermined value d of the measure of the spatial speed for the observation section, the larger the size.
45. The apparatus according to any one of claims 40 or 44, wherein the information (74) indicates a predetermined value of a measure of the probability of the direction of movement of the observation section.
46. The apparatus according to claim 45, the apparatus being configured to perform the setting such that the higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.
47. The apparatus according to any one of claims 40 or 46, wherein the information (74) indicates the predetermined spatial speed in a time-varying and / or spatially varying and / or direction-varying manner.
48. The apparatus according to any one of claims 40 or 47, configured such that the first portion (64) coincides with the spatial section (62).
49. The apparatus according to any one of claims 40 or 48, wherein each of the spatio-temporal portions of the spatially varying scene (30) encoding the temporal variation in the plurality of media segments is a temporal segment (54) of the spatially varying scene (30) at a respective one of the tiles (50) into which the spatially varying scene (30) is spatially subdivided.
50. The apparatus according to claim 49, wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying scene encoding the temporal variation at an associated quality level within a set of quality levels.
51. The apparatus according to claim 40, wherein the apparatus is configured to set the size in a manner independent of the size of the temporally varying observation section depending on the information (74).
52. The apparatus according to claim 40, wherein the information includes different values for the size for different size options of the time-varying observation section, and the apparatus uses the values included in the information for the size option suitable for the actual size of the time-varying observation section.
53. The apparatus according to claim 51 or 52, wherein the information indicates the size in terms of the number of tiles.
54. A streaming media server for media content of a spatially varying scene over time, configured to: make available a plurality of media segments for extraction by a device, such that the device can select media segments for extraction, the media segments having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality, and a second portion of the spatially varying scene over time spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first portion follows a temporally varying observation section of the spatially varying scene over time, and signal in the media segments and / or by signaling to the device information on how to set the size and / or position of the first portion.
55. A signal defining a media presentation description, comprising: Information regarding calculating addresses of multiple media segments such that a device can use the information to select and extract a media segment from the multiple media segments, the media segment having at least a spatial section of the spatially varying scene encoded therein with the time variation, encoded such that a first part of the spatial section is encoded into the selected media segment at a predetermined quality and a second part of the spatially varying scene that is spatially adjacent to the first part is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced relative to the predetermined quality, and such that the first part follows an observation section of the time variation of the spatially varying scene, and Information in the media segment and / or via a signal to the device on how to set the size and / or position of the first part.
56. The signal according to claim 55, wherein the information (74) indicates the size in the form of an increment relative to the size of the observation section of the time variation or a scaling of the size of the observation section of the time variation.
57. The signal according to claim 55 or 56, wherein the information (74) indicates a predetermined value of a measure of the spatial velocity for the observation section.
58. The signal according to claim 57, wherein the information (74) indicates the predetermined value of the measure of the spatial velocity for the observation section for: A default percentage of users for whom the measured spatial velocity does not exceed the predetermined value, and / or A percentage value indicating the percentage of users for whom the measured spatial velocity does not exceed the predetermined value, and / or A percentage value indicating the percentage of users for whom the observation section is in a predetermined area, and / or A hint of one or more types of user inputs controlling the movement of the observation section for which the predetermined value is applicable.
59. The signal according to claim 57 or 58, the signal being configured to perform the setting such that The higher the predetermined value d of the measure of the spatial velocity for the observation section, the larger the size.
60. The signal according to any one of claims 55 or 59, wherein the information (74) indicates a predetermined value of a measure of the probability of the direction of movement of the observation section.
61. The signal according to claim 60, the signal being configured to perform the setting such that The higher the probability of the corresponding direction of movement, the more the first part extends into the corresponding direction.
62. The signal according to any one of claims 55 or 61, wherein the information (74) indicates the predetermined spatial velocity in a time-varying and / or spatially varying and / or direction-varying manner.
63. The signal according to any one of claims 55 or 62, configured such that the first part (64) coincides with the spatial section (62).
64. The signal according to any one of claims 55 or 63, wherein each of the spatio-temporal portions of the spatially varying spatio-temporal scene (30) encoded into the plurality of media segments is a temporal segment (54) of the spatially varying spatio-temporal scene (30) at a respective one of the tiles (50) into which the spatially varying spatio-temporal scene (30) is spatially subdivided.
65. The signal according to claim 64, wherein each of the plurality of media segments has an associated spatio-temporal portion of the spatially varying spatio-temporal scene encoded therein at an associated quality level within a set of quality levels.
66. The signal according to claim 55, wherein the device is configured to set the size in a manner independent of the size of the temporally varying observation section depending on the information (74).
67. The signal according to claim 55, wherein the information includes different values for the size for different size options of the time-varying observation section, and the device uses the values included in the information for the size option that fits the actual size of the time-varying observation section.
68. The signal according to claim 65 or 66, wherein the information indicates the size in terms of the number of tiles.
69. A video bitstream having video encoded therein, the video bitstream including a signaling effect for one or more of the size of a focus area within the video to which the decoding capabilities for decoding the video should be concentrated and a recommended preferred observation section area of the video.
70. A decoder for decoding video from a video bitstream, configured to: derive a signaling effect (74) of the size of a focus area within the video from the video bitstream, and concentrate the decoding capabilities for decoding the video to the focus area (116).
71. The decoder according to claim 70, the decoder being configured to specifically decode the focus area.
72. The decoder according to claim 70, the decoder being configured to decode each image of the video starting at the focus area where decoding begins.
73. The decoder according to claim 70, the decoder being configured to stop decoding each image of the video after decoding the focus area.
74. The decoder according to any one of claims 70 to 73, wherein the signaling effect absolutely indicates the size, or the decoder is configured to scale the size of the focus area with a parameter included in the signaling effect.
75. A device for streaming media content regarding a spatially varying spatio-temporal scene (30), configured to: derive (90) from a media presentation description: at least one version, for tile-based streaming, the spatially varying spatio-temporal scene being provided in the at least one version, for each of the at least one version, an indication of the beneficial requirements for the respective version of the spatially varying spatio-temporal scene to benefit from the tile-based streaming, Match the beneficial requirements of the at least one version to the device capabilities of the device or another device interacting with the device regarding the tile-based streaming media (92).
76. The device according to claim 75, wherein the beneficial requirements and the device capabilities relate to decoding capabilities.
77. The device according to claim 75 or 76, wherein the beneficial requirements and the device capabilities relate to the number of available decoders.
78. The device according to any one of claims 75 to 77, wherein the beneficial requirements and the device capabilities relate to layer and / or profile descriptors.
79. The device according to any one of claims 75 to 78, wherein the beneficial requirements and the device capabilities relate to the type of input device for moving the viewing section across the spatial scene (30) that changes over time, or to the speed of using the input device to move the viewing section across the spatial scene (30) that changes over time.
80. The device according to any one of claims 75 to 79, wherein the device is configured to: Select media segments from a plurality (46) of media segments (58) available on the server (20), the selection being made by using the calculation rules contained in the media presentation description to calculate the addresses of the selected media segments, Extract the selected media segments from the server (20) using the calculated addresses, wherein the device is configured to perform the selection such that the selected media segments have at least a spatial section (62) of the spatially varying scene (30) that changes over time encoded therein in such a way that: A first part (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality, and a second part (66; 72) of the spatially varying scene that is spatially adjacent to the first part (64) is not encoded into the selected media segment or is encoded into the selected media segment at another quality reduced with respect to the predetermined quality, and such that the first part (64) follows the viewing section (28) of the spatially varying scene (30) that changes over time.
81. A streaming media server for streaming media content regarding a spatially varying scene (30) that changes over time, configured to: Provide (90) a media presentation description from which can be derived At least one version, for tile-based streaming, the spatially varying scene that changes over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements for benefiting from the corresponding version of the spatially varying scene that changes over time for tile-based streaming, Thereby enabling a device streaming the media content from the streaming media server to Match the beneficial requirements of the at least one version to the device capabilities of the device or another device interacting with the device regarding the tile-based streaming (92).
82. A media presentation description, comprising: Information about at least one version, for tile - based streaming, a spatio - temporal scene that changes over time is provided in the at least one version. For each of the at least one version, an indication of the beneficial requirements of the corresponding version of the spatio - temporal scene that changes over time for benefiting from the tile - based streaming.
83. The media presentation description according to claim 82, wherein the beneficial requirements and the device capabilities relate to decoding capabilities.
84. The media presentation description according to claim 82 or 83, wherein the beneficial requirements and the device capabilities relate to the number of available decoders.
85. The media presentation description according to any one of claims 82 to 84, wherein the beneficial requirements and the device capabilities relate to layer and / or profile descriptors.
86. The media presentation description according to any one of claims 82 to 85, wherein the beneficial requirements and the device capabilities relate to the type of input device for moving an observation section across the spatio - temporal scene (30) that changes over time, or to the speed of using the input device to move the observation section across the spatio - temporal scene (30).
87. The media presentation description according to any one of claims 82 to 86, further comprising calculation rules, using which the device can select a media segment from a plurality of media segments available on the server by calculating the address of the selected media segment using the included calculation rules.
88. A device for streaming media content regarding a spatio - temporal scene (30) that changes over time, configured to: calculate the address of a media segment depending on a spatial viewport position and at least one parameter, the media segment depicting the spatio - temporal scene (30) that changes over time and in the at least one parameter, use the calculated address to extract the media segment.
89. The device according to claim 88, wherein the at least one parameter includes one or more coordinates of an observation center and / or an observation depth.
90. A media presentation description, comprising: calculation rules for calculating the address of a media segment depending on a spatial viewport position and at least one parameter, in order to use the calculated address to extract the media segment, the media segment depicting the spatio - temporal scene (30) that changes over time and in the at least one parameter.
91. A streaming server for allowing a device to stream media content regarding a spatio - temporal scene that changes over time from a server, configured to provide the media presentation description according to claim 90.
92. A device for streaming media content regarding a spatio - temporal scene that changes over time, configured to: select a media segment from a plurality of media segments available on the server, wherein the device is configured to: encode a first part of the spatio - temporal scene that changes over time into the selected media segment with increased quality compared to the spatial neighborhood of the first part or in such a way that the spatial neighborhood of the first part is not encoded into the selected media segment, emit a log message that records: Instantaneous measurement results of the spatial position and / or movement of the first part; and / or Statistical values of the spatial position and / or movement of the first part, such as time average values; and / or Instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or Indications of the set of buffers (300) of the device involved in buffering the selected media segment, descriptions of the distribution rules applied in distributing the selected media segment to the set of buffers, and the instantaneous buffer fill levels of each of the set of buffers; and / or Measurement results of the amount of the selected media segment yet to be output from the buffer of the device for undergoing decoding (42); and / or Statistical values of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or Instantaneous measurement results of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section; and / or Statistical values of the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section, such as time average values; and / or The field of view covered by the observation section; and / or Instantaneous measurement results of the user position or the observation depth relative to the scene center (100); and / or Statistical values of the user position or the observation depth relative to the scene center (100), such as time average values.
93. The device according to claim 92, wherein the quality of the first part or the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section is measured as the duration during which the lower quality part and the higher quality part are visible in the observation section.
94. The device according to claim 92 or 93, configured to perform the selection such that the first part (64) of the spatial scene of the time variation pins the observation section (28).
95. The device according to any one of claims 92 to 94, configured to issue a log message that records the instantaneous measurement results of the quality of the spatial scene of the time variation until encoded into the selected media segment and until visible in the observation section as one of the following: Measurement results of the average density of the pixels falling within the observation section (28), with the spatial scene of the time variation encoded into the selected media segment at the average density.
96. The device according to claim 95, configured such that the measurement results measure the average density of the pixels by averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded into the selected media segment.
97. The apparatus according to claim 95, configured such that the measurement result measures the average density of pixels by averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded in the selected media segment.
98. The apparatus according to claim 95, configured such that the emitted log message indicates whether the measurement result measures the average density of pixels by: averaging the pixel density in a spatially uniform manner with respect to the pixel grid of the image encoded in the selected media segment, or averaging the pixel density in a spatially non-uniform manner with respect to the pixel grid of the image encoded in the selected media segment.
99. The apparatus according to claim 97 or 98, wherein averaging the pixel density in a spatially non-uniform manner corresponds to averaging in a spherically uniform manner, or averaging spatially uniformly with respect to a viewport plane (310) that is perpendicular to the central viewing direction (312) of the viewing section (28).
100. The apparatus according to any one of claims 95 to 99, configured such that the measurement result measures the average density of pixels by: averaging the pixel density in a manner that limits the averaging to a central sub-section of the viewing section (28), or applying a higher averaging weight to the central sub-section (202) compared to an edge portion (204) of the viewing section surrounding the central sub-section.
101. The apparatus according to any one of claims 95 to 100, configured such that the measurement result measures the average density of pixels separately along a horizontal viewing section axis (204) and a vertical viewing section axis (206).
102. The apparatus according to any one of claims 95 to 101, configured to intermittently emit log messages.
103. The apparatus according to any one of claims 95 to 102, configured to emit log messages at a rate controlled by a manifest file, the apparatus performing the selection of the media segment for download based on the manifest file.
104. The apparatus according to any one of claims 95 to 103, wherein each of the plurality of media segments available on the server belongs to one of a plurality of representations of the time-varying spatial scene, the representations differing in one or more of the following: a scene section (50) of the time-varying spatial scene encoded in the media segment, the quality used to encode the time-varying spatial scene in the media segment, the spatial quality variation used to encode the time-varying spatial scene in the media segment, wherein the apparatus is configured to emit a log message that records a description of the distribution rule applied in distributing the selected media segment to the set of buffers in the form of an association of each buffer with a combination of one or two or more of the following: the scene section, the quality, the spatial quality distribution, representation 105. The apparatus according to any one of claims 95 to 104, wherein the representations are grouped into adaptive sets according to one or more of the following: the scene segments of the spatially varying scene over time encoded into the media segment, the spatial quality variations used to encode the spatially varying scene over time into the media segment, wherein the apparatus is configured to issue a log message that records the description of the distribution rule applied in distributing the selected media segment to the set of buffers, in the form of an association of each buffer with one of the adaptive sets, or in the form of an association of each buffer with one of the representations.
106. The apparatus according to any one of claims 95 to 105, wherein the apparatus is configured to issue a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42), in the form of a time measurement result.
107. The apparatus according to claim 106, wherein the apparatus is configured to present the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42) in the form of a time measurement result in time units less than the time length of the media segment (58) and / or in a form defined independently of the time length of the media segment and / or in the form of milliseconds.
108. The apparatus according to any one of claims 95 to 107, wherein the apparatus is configured to issue a log message that records the measurement result of the amount of the selected media segment that has not been output from the buffers of the apparatus and is to undergo decoding (42) in a format classified by one or more of the following: the buffer of the decoder in which the corresponding media segment has been buffered, the scene segment encoded into the corresponding media segment, the quality used to encode the spatially varying scene over time into the corresponding media segment, the spatial quality distribution used to encode the spatially varying scene over time into the corresponding media segment.
109. A method for streaming media content regarding a spatially varying scene (30) over time, comprising: selecting (56) media segments from a plurality (46) of media segments (58) available on a server (20), extracting (60) the selected media segments from the server (20), performing the selection such that the selected media segments have at least a spatial segment (62) of the spatially varying scene (30) encoded therein, encoded such that a first part (64) of the spatial segment is encoded into the selected media segment at a predetermined quality, and a second part (66) of the spatially varying scene that is spatially adjacent to the first part (64) is encoded into the selected media segment at another quality, the another quality satisfying a predetermined relationship with respect to the predetermined quality, and The method further includes deriving the predetermined relationship from information (68) included in the selected media segment and / or from a signal received from the server (20).
110. A method for streaming media content of a spatially varying scene over time, comprising: making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is encoded into the selected media segment at another quality, and signaling information about a predetermined relationship in the media segment and / or via a signal to the device, the another quality satisfying the predetermined relationship with respect to the predetermined quality.
111. A method for streaming media content of a spatially varying scene (30) over time, comprising: selecting a media segment from a plurality (46) of media segments (58) available on a server (20), extracting the selected media segment from the server (20), wherein the selection is performed such that the selected media segment has at least a spatial section (62) of the spatially varying scene (30) over time encoded therein, encoded such that: a first portion (64) of the spatial section (62) is encoded into the selected media segment at a predetermined quality and a second portion (66; 72) of the spatially varying scene over time that is spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and causing the first portion (64) to follow a time-varying viewing section (28) of the spatially varying scene (30), and the method further includes setting the size of the first portion (64) depending on information (74) included in the selected media segment and / or on a signal received from the server.
112. A method for streaming media content of a spatially varying scene over time, comprising: making a plurality of media segments available for extraction by a device, enabling the device to select a media segment for extraction, the selected media segment having at least a spatial section of the spatially varying scene over time encoded therein, encoded such that a first portion of the spatial section is encoded into the selected media segment at a predetermined quality and a second portion of the spatially varying scene over time that is spatially adjacent to the first portion is not encoded into the selected media segment or is encoded into the selected media segment at another quality that is reduced relative to the predetermined quality, and causing the first portion to follow a time-varying viewing section of the spatially varying scene, and Signal information on how to size the first part in the media segment and / or by signaling to the device.
113. A method for decoding video from a video bitstream, comprising: Deriving a signal for the size of a focus area within the video from the video bitstream, and Concentrating decoding capabilities for decoding the video on the focus area.
114. A method for streaming media content regarding a spatially varying scene (30) over time, performed by a device, comprising: Deriving (90) from a media presentation description: At least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of the beneficial requirements for the corresponding version of the spatially varying scene over time to benefit from the tile-based streaming, Matching (92) the beneficial requirements of the at least one version with the device capabilities of the device or another device interacting with the device regarding the tile-based streaming.
115. A method for streaming media content regarding a spatially varying scene over time, comprising: Selecting a media segment from a plurality of media segments available on a server, Extracting the selected media segment from the server, where the selection is performed such that the selected media segment has a first part of the spatially varying scene encoded therein with increased quality compared to the spatial neighborhood of the first part or in a manner such that the spatial neighborhood of the first part is not encoded into the selected media segment, The method further comprises emitting a log message that records: An instantaneous measurement of the spatial position and / or movement of the first part; and / or A statistical value of the spatial position and / or movement of the first part, such as a time average; and / or An instantaneous measurement of the quality of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section; and / or An indication of the set of buffers of the device involved in buffering the selected media segment, a description of the distribution rules applied in distributing the selected media segment into the set of buffers, and the instantaneous buffer fullness of each of the set of buffers; and / or A measurement of the amount of the selected media segment yet to be output from the device's buffer for undergoing decoding (42); and / or A statistical value of the quality of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section, such as a time average; and / or An instantaneous measurement of the quality of the first part or of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section; and / or A statistical value of the quality of the first part or of the spatially varying scene until encoded into the selected media segment and until visible in the viewing section, such as a time average; and / or The field of view covered by the viewing section; and / or Instantaneous measurements of a user position or viewing depth relative to a scene center (100); and / or Statistical values, such as time averages, of a user position or viewing depth relative to a scene center (100).
116. A method for streaming media content of a spatially varying scene (30) over time, comprising: Providing (90) a media presentation description from which can be derived: At least one version, for tile-based streaming, the spatially varying scene over time is provided in the at least one version, For each of the at least one version, an indication of beneficial requirements of a corresponding version of the spatially varying scene over time for benefiting from the tile-based streaming, Thereby enabling an apparatus for streaming the media content from a streaming server to Match (92) the beneficial requirements of the at least one version with the apparatus capabilities of the apparatus or another apparatus interacting with the apparatus regarding the tile-based streaming.
117. A method for allowing an apparatus to stream media content of a spatially varying scene (30) over time from a server, comprising providing the media presentation description as claimed in claim 90.
118. A computer program having program code for performing the method according to any one of claims 109 to 117 when the program is executed on a computer.
Citation Information
Patent Citations
Spatially unequal streaming
CN115037917A