Spatial Unequal Streaming
By strategically selecting media segments based on server-provided signaling, the VR streaming method addresses the challenges of bandwidth and computational resource management, resulting in improved visual quality and efficiency in VR streaming.
Patent Information
- Application Number
- JP2022162680
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-07-08
- Filing Date
- 2022-10-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2037-10-11
AI Technical Summary
Current VR streaming technologies face challenges in efficiently managing bandwidth and computational resources while maintaining high visual quality, especially when streaming spatially varying scenes with high resolution requirements.
The proposed solution involves streaming media content related to spatially varying scenes by selecting media segments that prioritize quality in specific portions of the scene, based on signaling from the server that indicates the quality relationships between different parts of the scene. This approach allows for adaptive streaming that improves visual quality and reduces processing complexity without increasing bandwidth consumption.
This method enhances the visual quality of VR streaming by ensuring that the most critical parts of the scene are streamed at the highest quality, while reducing the computational resources and bandwidth required, thereby improving the overall streaming experience.
Smart Images

Figure 0007684263000017 
Figure 0007684263000018 
Figure 0007684263000019
Abstract
Description
Technical Field
[0001] The present invention relates to spatially unequal streaming, such as occurs in virtual reality (VR) streaming.
Background Art
[0002] VR streaming typically involves the transmission of very high-resolution video. The resolution of the human fovea is about 60 pixels per degree. Considering a 360°×180° global transmission, it would be necessary to transmit a resolution of about 22k×11k pixels. Transmitting such a high resolution results in very high bandwidth requirements, so another solution is to transmit only the viewport shown on a head-mounted display (HMD) having a 90°×90° field of view (FoV). In this case, the video will be about 6K×6K pixels. The trade-off between transmitting the entire video at the highest resolution and transmitting only the viewport is to transmit the viewport at a high resolution and some adjacent data (or the rest of the spherical video) at a low resolution or low quality.
[0003] In a DASH scenario, an omnidirectional video (also known as a spherical video) can be provided such that the aforementioned mixed-resolution or mixed-quality video is controlled by a DASH client. The DASH client only needs to know the information explaining how the content is provided.
[0004] One example can be to provide different representations using different projections with asymmetric characteristics such as different qualities and distortions for different parts of the video. Each representation corresponds to a given viewport, and that viewport is encoded at a higher quality / resolution than other content. Knowing the direction information (the direction of the viewport where the content is encoded at a higher quality / resolution), the DASH client can dynamically select one or the other representation to match the user's line of sight direction at any time.
[0005] A more flexible option for a DASH client to select such asymmetric characteristics for omnidirectional video would be when the video is divided into several spatial regions and each region is available at a different resolution or quality. One option is that it can be divided into rectangular regions (also known as tiles) based on a grid, but other options are also conceivable. In such a case, the DASH client requires some signaling about the different qualities provided by different regions and can download different regions at different qualities so that the viewport displayed to the user is of higher quality than other non-displayed content.
[0006] In any of the above cases, when a user operation occurs and the viewport is changed, the DASH client will take some time to download the content in a way that matches the new viewport in response to the user's movement. Between the time the user moves and the DASH client adjusts its request to match the new viewport, the user will see some regions of high and low quality in the viewport at the same time. The acceptable quality / resolution differences depend on the content, but the quality the user sees will decrease anyway.
[0007] Therefore, for the partial presentation of spatially scene content streamed by adaptive streaming, it is preferable to have in hand the concept of reducing, more efficient rendering, or even improving the visual quality for the user. Summary of the Invention Problems to be Solved by the Invention
[0008] Accordingly, an object of the present invention is to provide a concept for streaming spatial scene content in a non-uniform spatial manner to improve the visual quality of a user, or to reduce the complexity of processing or the required bandwidth in a streaming search site, or to provide a concept for streaming spatial scene content in a way that expands its applicability to further application scenarios.
[0009] This object is achieved by the subject matter of the independent claims in question.
Means for Solving the Problem
[0010] A first aspect of the present invention is based on the discovery that streaming media content related to a spatially varying scene over time, such as video, can be improved in terms of visible quality and / or computational complexity at comparable bandwidth consumption at the streaming receiving site when a selected and retrieved media segment and / or signaling obtained from a server has a hint about a predetermined relationship that different portions of the spatially varying scene that change over time should conform to the quality encoded in the selected and retrieved media segments. Otherwise, the search device may not be able to know in advance about the adverse effects of the juxtaposition of portions encoded with different qualities in the selected media segment on the overall visible quality experienced by the user. The search device can appropriately select from among the media segments provided by the server based on the information contained in the media segments and / or signaling obtained from the server, such as in a manifest file (media presentation description), or additional streaming-related control messages from the server to the client, such as SAND messages. In this way, virtual reality streaming or partial streaming of video content can be made more robust against quality degradation, which might otherwise occur due to an inappropriate distribution of the available bandwidth over this spatial section of the temporally varying spatial scene presented to the user.
[0011] A further aspect of the present invention relates to streaming of media content for a temporally varying spatial scene such as video in a spatially non-uniform manner, using a first quality in a first portion and a second lower quality in a second portion, or leaving the second portion non-streaming, which may improve the visual quality and / or reduce the complexity in terms of bandwidth consumption and / or computational complexity on the streaming search side, depending on the information included in the media segment and / or the signaling obtained from the server by determining the size and / or position of the first portion. For example, assume that a temporally varying spatial scene is provided to the server in a tile-based manner for tile-based streaming. That is, assume that the media segment represents a temporal portion of the spectrum of the temporally varying spatial scene, each of which is a temporal segment of the spatial scene within the corresponding tile of the distribution of tiles into which the spatial scene is sub-divided. In such a case, it is up to the search device (client) to determine how to allocate the available bandwidth and / or computational power across the spatial scene, i.e., at the tile granularity. The search device performs the selection of the media segment such that the first portion of the subsequent spatial scene tracks each of the temporally varying view sections of the spatial scene and is selected and searched for in the media segment encoded to a predetermined quality that can be the highest quality achievable with the current bandwidth and / or computational power conditions. The second spatially adjacent portion of the spatial scene may, for example, not be encoded in the selected and searched media segment or may be encoded there with a reduced further quality for a predetermined quality. In such a situation, regardless of the orientation of the view section, its Calculating the number / sum of adjacent tiles that completely cover a view section that changes over time is a computationally complex problem or even infeasible. Depending on the projection selected to map the spatial scene onto individual tiles, the angular scene coverage per tile can vary due to the fact that this scene and the individual tiles can overlap with each other, and calculating the number of adjacent tiles sufficient to cover the view section spatially becomes more difficult regardless of the orientation of the view section. Thus, in such situations, the aforementioned information can indicate the number N of tiles or the size of the first part as the number of tiles respectively. By this means, the device can track a view section that changes over time by selecting those media segments that have an aggregate placed at the same location of the N tiles encoded there with a given quality. The fact that this set of N tiles sufficiently covers the view section can be guaranteed by the information indicating N. Another example is the information contained in the media segment and / or the signaling obtained from the server, which indicates the size of the first part relative to the size of the view section itself. For example, this information can set up some form of "safety zone" or prefetch zone around the actual view section in order to account for the movement of the view section that changes over time. The faster the view section that changes over time moves across the spatial scene, the larger the safety zone. Thus, the aforementioned information can indicate the size of the first part in a way relative to the size of the view section that changes over time, such as in an incremental or scaling manner. A retrieval device that sets the size of the first part according to such information would be able to avoid the quality degradation that could occur because otherwise an unretrieved or low-quality part of the spatial scene could be visible in the view section. Here, it is irrelevant whether this scene is provided in a tile-based method or in some other way.
[0012] In connection with the situation immediately preceding the present application, if signaling of the size of the focus area is provided in a video for which the decoding capabilities of the video bitstream for decoding the video should be focused, the video bitstream having the video encoded therein can be decoded with higher quality. By this means, a decoder that decodes video from the bitstream can concentrate, or even limit, its decoding capabilities on the portion of the video having the size of the focus area signaled in the video bitstream. For example, it can be seen that the portion decoded in this way is decodable by the available decoding capabilities and spatially covers the necessary portion of the video. For example, the size of the focus area signaled in this way can be selected to be large enough to cover the size of the view section and the motion of this view section, taking into account the decoding latency when decoding the video. Alternatively, in other words, signaling of the recommended preferred view section area of the video included in the video bitstream enables the decoder to handle this area in a preferred manner, whereby the decoder can concentrate the decoding power accordingly. Irrespective of performing area-specific decoding power focusing, area signaling can be transferred to the stage of selecting where to download, i.e., where to place, on which media segment and how to determine the dimensions of the quality-improved portion.
[0013] The first and second aspects of this application are closely related to the third aspect of this application and utilize the fact that a vast number of search devices stream media content from a server, and then set the size of the first portion, or the size and / or position, and / or appropriately set a predetermined relationship between the first and second qualities. It is to obtain information that can be used to appropriately set information of the foregoing type. Thus, according to this aspect of the present application, the search device (client) measures the momentaneous measurement, or statistical values that measure the spatial position and / or movement of the first portion, the momentaneous measurement, or as long as it is encoded in the selected media segment and as long as it is visible in the view section, statistical values that measure the quality of the spatially varying scene over time, and encoded in the selected media segment and as long as it is visible in the view section, send a log message that collects one of the momentaneous measurement or statistical values that measure the quality of the first portion or the spatially varying scene over time. The momentaneous measurement and / or statistical value can provide time information regarding the time at which each momentaneous measurement or statistical value was obtained. The log message may be sent to the server that provides the media segment, or sent to another device that evaluates the received log message, and based on that, update the current setting of the foregoing information used to set the size of the first portion, or the size and / or position, and / or derive a predetermined relationship based on that.
[0014] According to a further aspect of the present application, streaming media content related to a spatially varying scene that changes over time, such as a video, and in particular streaming media content in a tile-based method, is provided with a media presentation description that includes at least one version of the spatially varying scene provided for tile-based streaming and, for each of the at least one version, a display of the benefits of the requirements for each version of the spatially varying scene to benefit from tile-based streaming. By this means, the search device can match the beneficial requirements of at least one version with the device capabilities of the search device itself or other devices that interact with the search device with respect to tile-based streaming. For example, the requirements for obtaining benefits may be related to the decoding capability requirements. That is, if the decoding capability for decoding the streamed / search media content is not sufficient to decode all the media segments necessary to cover the view section of the spatially varying scene over time, attempting to stream and present the media content is wasteful in terms of time, bandwidth, and computing power, and accordingly, it is more effective not to attempt it anyway. For example, if the media segments for a certain tile form a media stream, such as a video stream, separately from the media segments for other tiles, the decoding capability requirements can indicate, for example, some decoder instantiations required for each version. The decoding capability requirements may also be related to further information, such as a specific part of the decoder instantiation required to conform to a given decoding profile and / or level, or may indicate a specific minimum capability of the user input device to move the viewport / section at a speed sufficient for the user to view the scene through it. Depending on the content of the scene, if the ability to move is low, the user may not be able to view the interesting parts of the scene sufficiently.
[0015] A further aspect of the present invention relates to the extension of the streaming of media content related to a spatially varying scene over time. In particular, the concept according to this aspect is that not only does the spatial scene actually vary over time, but it can also vary with respect to at least one further parameter that suggests, for example, a view and position, view depth or some other physical parameter. The retrieval device can use adaptive streaming in this context by calculating the address of a media segment according to the viewport direction and at least one further parameter, the media segment describes a spatially varying scene over time and at least one parameter, and the media segment is retrieved using the address calculated from the server.
[0016] The aspects outlined above and their advantageous embodiments, which are the subject of the dependent claims, can be combined individually or all together.
[0017] Preferred embodiments of the present application are described below with reference to the drawings.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7a
Figure 7b
Figure 7c
Figure 8a
Figure 8b
Figure 8c
Figure 8d
Figure 8e
Figure 8f
Figure 9
Figure 10
Figure 11
[0019] To facilitate understanding of the description of the embodiments of the present application regarding various aspects of the present application, FIG. 1 shows an example of an environment to which the embodiments described below of the present application can be applied and advantageously used. In particular, FIG. 1 shows a system composed of a client 10 and a server 20 that communicate via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) can be used for communication 22 between the client 10 and the server 20. However, the embodiments outlined later should not be construed as being limited to the use of DASH. Similarly, terms such as Media Presentation Description (MPD) should be understood in a broad sense to also cover manifest files defined differently from DASH.
[0020] Figure 1 shows a system configured to implement a virtual reality application. That is, the system presents a view section 28 to a user wearing a head-up display 24, i.e., via an internal display 26 of the head-up display 24, from a spatially varying scene 30 that varies in time corresponding to the orientation of the head-up display 24, as exemplarily measured by an internal orientation sensor 32 such as an inertial sensor of the head-up display 24. That is, the section 28 presented to the user forms a section of the spatial scene 30 whose spatial position corresponds to the orientation of the head-up display 24. In the case of Figure 1, the spatially varying scene 30 that varies in time is depicted as an omnidirectional video or a spherical video, but the description of Figure 1 and the embodiments described later can be easily transferred to other examples as well. For example, presenting a section from a video having the spatial position of the section 28 is determined by the intersection of face access or eye access with a virtual or physical projector wall or the like. Further, the sensor 32 and the display 26 may be constituted by different devices such as a remote control device and a corresponding television, respectively, or may be part of a handheld device such as a mobile device such as a tablet or a mobile phone. Finally, it should be noted that some of the embodiments described later can also be applied to scenarios where the region 28 presented to the user always covers the spatially varying scene 30 that varies in time as a whole with unevenness when presenting the spatially varying scene 30, for example, scenarios regarding an uneven distribution of quality across the spatial scene.
[0021] Further details regarding the server 20, the client 10, and the method by which the spatial content 30 is provided at the server 20 are shown in Figure 1 and described below. However, these details should not be treated as limiting the embodiments described later, but rather should serve as examples of how to implement any of the embodiments described later.
[0022] In particular, as shown in FIG. 1, the server 20 can include a storage device 34 and a controller 36 such as a computer appropriately programmed or an application-specific integrated circuit. The storage device 34 stores media segments representing the spatially varying scene 30 that changes over time. Specific examples are outlined in more detail below with respect to the figures in FIG. 1. The controller 36 can respond to requests sent by the client 10 by resending the requested media segment, media presentation description to the client 10, and can also send additional information to the client 10 by itself. Details regarding this are also described below. The controller 36 can retrieve the requested media segment from the storage device 34. Other information such as a media presentation description or a part thereof can also be stored in this storage device for other signals sent from the server 20 to the client 10.
[0023] As shown in FIG. 1, the server 20 may optionally further include a stream modification unit 38 that modifies media segments transmitted from the server 20 to the client 10 in response to requests from the latter so as to form a media data stream at the client 10. For example, the media segments retrieved by the client 10 in this way are actually aggregated from several media streams, but a single media stream can be decoded by a single associated decoder. However, the presence of such a stream modification unit 38 is optional.
[0024] Client 10 in FIG. 1 is illustratively depicted as including client device or controller 40, or one or more decoders 42 and reprojectors (also referred to as reprojection units) 44. Client device 40 can be a properly programmed computer, a programmed hardware device such as a microprocessor, FPGA, or an application-specific integrated circuit, etc. Client device 40 undertakes to select the segments to be retrieved from server 20 from among a plurality of 46 media segments provided by server 20. For this purpose, client device 40 first retrieves a manifest or media presentation description from server 20. Similarly, client device 40 obtains calculation rules for calculating the addresses of media segments from among a plurality 46 corresponding to a required spatial portion of spatial scene 30. The media segments thus selected are retrieved from server 20 by client device 40 by sending respective requests to server 20. These requests contain the calculated addresses.
[0025] In this way, the media segments retrieved by the client device 40 are transferred by the latter to one or more decoders 42 for decoding. In the example of FIG. 1, the media segments retrieved and decoded in this way represent only the spatial section 48 of the spatially varying spatial scene 30 for each temporal time unit. However, as already mentioned above, this may be different according to other aspects. For example, the presented view section 28 may always cover the entire scene. The reprojector 44 can arbitrarily reproject and cut out the view section 28 from the retrieved and decoded scene content of the selected, retrieved and decoded media segment so as to be displayed to the user. For this purpose, as shown in FIG. 1, the client device 40 continuously tracks and updates the spatial position of the view section 28 in response to, for example, user orientation data from the sensor 32, and notifies the reprojector 44, for example, of this current spatial position of the scene section 28 and the reprojection mapping applied to the retrieved and decoded media content so as to be mapped to the region-forming view section 28. The reprojector 44 can accordingly apply mapping and interpolation, for example, to a regular grid of pixels displayed on the display 26.
[0026] FIG. 1 shows the case where cube mapping is used to map the spatial scene 30 to the tiles 50. Thus, the tiles are depicted as rectangular small regions of the cube onto which the spherical-shaped scene 30 is projected. The reprojector 44 reverses this projection. However, other examples may equally apply. However, other examples may equally apply. For example, instead of cube projection, projection onto a frustum or a pyramid without frustum can be used. Further, the tiles in FIG. 1 are depicted so as not to overlap with respect to the coverage of the spatial scene 30, but the subdivision of the tiles may include mutual tile overlap. And as will be outlined in more detail below, it is also not essential to spatially subdivide the scene 30 into tiles 50, with each tile forming one representation as will be further described below.
[0027] Thus, as shown in FIG. 1, the entire spatial scene 30 is spatially subdivided into tiles 50. In the example of FIG. 1, each of the six faces of the cube is subdivided into four tiles. For the sake of explanation, the tiles are enumerated. For each tile 50, the server 20 provides a video 52 as shown in FIG. 1. More precisely, the server 20 may even provide a plurality of videos 5 2 for each tile 50, and these videos have different qualities Q#. Further, the video 52 is temporally subdivided into temporal segments 54. The temporal segments 54 of all the videos 52 of all the tiles T# form or are encoded into one of the plurality of 46 media segments stored in the storage device 34 of the server 20.
[0028] It is again emphasized that even the tile - based streaming example shown in FIG. 1 forms only an example from which many variations are possible. For example, FIG. 1 seems to suggest that media segments related to a higher - quality representation of the scene 30 coincide with the tiles that coincide with the tiles in which the media segment encoding the scene 30 at quality Q1 is located. This coincidence is not necessarily required, and tiles of different qualities may even correspond to tiles of different projections of the scene 30. Further, although not discussed so far, the media segments corresponding to the different quality levels shown in FIG. 1 may have different spatial resolutions and / or signal - to - noise ratios and / or temporal resolutions, etc.
[0029] Finally, unlike the tile-based streaming concept for the tiles 50 into which the scene 30 is spatially subdivided, where media segments that can be individually retrieved from the server 20 by the apparatus 40 are concerned, the media segments provided at the server 20 are alternatively, for example, each of which encodes the scene 30 therein in a spatially complete manner with a spatially varying sampling resolution, but has the maximum sampling resolution at different spatial positions within the scene 30. For example, it can be achieved by providing the server 20 with a series of segments 54 regarding the projection of the scene 30 onto frustum pyramids whose frustum tips are oriented in different directions from each other, thereby resulting in resolution peaks of different orientations.
[0030] Furthermore, regarding presenting the stream modification unit 38 as necessary, within the network device where the client 10 and the server 20 exchange the signals described herein, the same can be part of the client 10 or can be arranged therebetween.
[0031] After fairly generally describing the system of the server 20 and the client 10, the functions of the client device 40 according to the first aspect of the present application will be described in more detail. For this purpose, refer to FIG. 2 which shows the device 40 in more detail. As already described above, the device 40 is for streaming media content related to a spatially changing scene 30 over time. As described with respect to FIG. 1, the device 40 may be configured such that the streamed media content is spatially continuously related to the entire scene, or may be configured to be limited to its section 28 only. In any case, the device 40 includes a selection unit 56 for selecting an appropriate media segment 58 from a plurality 46 of media segments available on the server 20, and a search unit 60 for searching for the media segment selected from the server 20 by each request such as an HTTP request. As described above, the selection unit 56 can use the media presentation description to calculate the address of the media segment selected by the search unit 60 using these addresses when searching for the selected media segment 58. For example, the calculation rule for calculating the address shown in the media presentation description may depend on the quality parameter Q, the tile T, and several time segments t. The address may be, for example, a URL.
[0032] Also as described above, the selection unit 56 is configured to perform the selection such that the selected media segment has at least one spatial section of the spatially changing scene over time and is encoded therein. The spatial section may completely cover the scene spatially continuously. FIG. 2 shows, at 61, an exemplary case where the device 40 is adapted to overlap the spatial section 62 of the scene 30 and surround the view section 28. However, as already described above, this is not necessarily the case, and the spatial portion may continuously cover the entire scene 30.
[0033] Furthermore, the selection unit 56 performs the selection such that the selected media segment has a section 62 encoded therein in a spatially non-uniform quality mode. More precisely, a first portion 64 of the spatial section 62, which is hatched in FIG. 2, is encoded in the media segment selected at a given quality. This quality can be, for example, the highest quality provided by the server 20 or can be a "good" quality. The device 42 moves or adapts the first portion 64, for example, so as to spatially follow a temporally varying view section 28. For example, the selection unit 56 selects the current time segment 54 of those tiles that inherit the current position of the view section 28. In so doing, the selection unit 56 can optionally keep the number of tiles constituting the first portion 64 constant, as will be explained with respect to a further embodiment below. In any case, a second portion 66 of the section 62 is encoded in a media segment 58 selected at another quality, such as a low quality. For example, the selection unit 56 selects a media segment corresponding to the time segment of the current tile that is spatially adjacent to the tiles of the portion 64 and belongs to a lower quality. For example, the selection unit 56 mainly selects the media segment corresponding to the portion 66 to address the possibility that the view section 28 may move too fast before the time interval corresponding to the current time segment ends and the portions 64 and the overlapping portion 66 are left behind, and the selection unit 56 will be able to newly spatially arrange the portion 64. In this situation, nevertheless, the portion of the section 28 protruding into the portion 66 may be presented to the user in a degraded quality state.
[0034] It is impossible for the apparatus 40 to evaluate what kind of quality degradation can occur by presenting the user with reduced-quality scene content together with high-quality scene content within the high-quality portion 64 in advance. In particular, a transition occurs between these two qualities, which may be clearly visible to the user. At least, such a transition may be visible depending on the current scene content within the section 28. The severity of the adverse effects of such a transition within the user's field of view is a characteristic of the scene content as provided by the server 20 and may not be predicted by the apparatus 40.
[0035] Therefore, according to the embodiment of FIG. 2, the apparatus 40 includes a derivation unit 66 that derives a predetermined relationship to be satisfied between the quality of the portion 64 and the quality of the portion 66. The derivation unit 66 may be included within a media segment, such as within a transport box in the media segment 58, and / or may be included in a unique signal transmitted from the server 20, such as a signalized or SAND message obtained from the server 20, such as within a media presentation description. This predetermined relationship is derived from information that may be included in the media segment 58 or the unique signal transmitted from the server 20. An example of how the information 68 looks is presented below. The predetermined relationship 70 derived by the derivation unit 66 based on the information 68 is used by the selection unit 56 to appropriately execute the selection. For example, the restrictions in the selection of the quality of the portions 64 and 66 compared to a completely independent selection of the quality of the portions 64 and 66 affect the distribution of the bandwidth available for extracting the media content regarding the portion 62 onto the portions 64 and 66. In any case, the selection unit 56 selects a media segment such that the quality encoded in the media segment finally retrieved for the portions 64 and 66 satisfies a predetermined relationship. An example of what the predetermined relationship looks like is also shown below.
[0036] The media segment selected by the search unit 60 and finally retrieved is transferred to one or more decoders 42 for final decoding.
[0037] According to a first example, for example, the signaling mechanism embodied by information 68 is , includes information 68 indicating to device 40, which may be a DASH client, which quality combinations are acceptable for the video content provided. For example, information 68 may be a list of quality pairs indicating to a user or device 40 that different regions 64 and 66 may be mixed with a maximum quality (or resolution) difference. Device 40 may be configured to necessarily use a particular quality level for portion 64, such as the highest quality level provided by server 10, and to derive the quality level that portion 66 may encode into the selected media segment from information 68 that portion 68 includes, for example in the form of a list of quality levels for portion 68.
[0038] Information 68 may indicate a tolerance for a measure of difference between the quality of part 68 and the quality of part 64. As a "measure" of quality difference the quality indicator of media segment 58 may be used, by which the same are differentiated in the media presentation description and by which its address is calculated using the calculation rules described in the media presentation description. In MPEG-DASH a corresponding attribute indicating quality would be for example @qualityRanking. When performing the selection, device 40 may take into account limitations on selectable quality level pairs at which parts 64 and 66 can be encoded into the selected media segment.
[0039] However, instead of this measure of difference, assuming that the bitrate usually increases monotonically as the quality goes up, the quality difference can also be measured, for example, by the bitrate difference, i.e., the acceptable difference in the bitrates at which parts 64 and 66 are respectively encoded in the corresponding media segments. Information 68 can indicate pairs of acceptable options for the quality encoded in the selected media segments for parts 64 and 66. Information 68 can indicate pairs of acceptable options for the quality encoded in the selected media segments for parts 64 and 66. Alternatively, information 68 can simply indicate the acceptable quality of encoded part 66, thereby indirectly indicating the acceptable or durable quality difference, assuming that main part 64 is encoded using some default quality, such as the highest possible or available quality. For example, information 68 can be a list of acceptable presentation IDs or can indicate the minimum bitrate level for the media segment regarding part 66.
[0040] However, instead, a more gradual quality difference may be desired, in which case, instead of quality pairs, quality groups (two or more qualities) may be indicated, where the quality difference increases depending on the distance to section 28, i.e., the viewport. That is, information 68 can indicate acceptable values for the measure of the difference between the qualities of parts 64 and 66 in a way that depends on the distance to view section 28. This can be done by a list of pairs of each distance to the view section and the corresponding acceptable value for measuring the quality difference beyond each distance. Below each distance, the quality difference must be smaller. That is, each pair indicates for the corresponding distance that a part within part 66 that is farther from the corresponding distance within section 28 may have a quality difference to the quality of part 64 that exceeds the corresponding acceptable value of this list item.
[0041] The acceptable values may increase as the distance to view section 28 increases. The acceptance of the quality differences described now is for these different productThe quality depends on the time it is displayed to the user. For example, content with a large quality difference may be acceptable if it is only displayed for 200 microseconds, while content with a small quality difference may be acceptable if it is displayed for 500 microseconds. Thus, according to a further example, information 68 may include, for example, in addition to the above-described quality combinations, or in addition to the acceptable quality differences, the time intervals during which the combination / quality differences may be acceptable. In other words, information 68 may indicate the acceptable or maximum allowable difference between the qualities of parts 66 and 64, along with the display of the maximum allowable time interval during which part 66 may be shown within view section 28 simultaneously with part 64.
[0042] As already mentioned, the acceptance of quality differences varies depending on the content itself. For example, the spatial position of different tiles 50 affects acceptance. Quality differences in a uniform background area containing low-frequency signals are expected to be more acceptable than those in foreground objects. Additionally, the temporal position also affects the acceptance rate due to content changes. Thus, according to another example, the signal forming information 68 is intermittently transmitted to device 40, for example, in the form of DASH representations or per period. That is, the predetermined relationship indicated by information 68 may be updated intermittently. Additionally and / or alternatively, the signaling mechanism implemented by information 68 may vary spatially. That is, information 68 can be made spatially dependent, for example, by SRD parameters within DASH. That is, different predetermined relationships can be indicated by information 68 regarding different spatial regions of scene 30.
[0043] As described with respect to FIG. 2, an embodiment of apparatus 40 desires to be able to temporarily view in section 28 before being able to vary the positions of sections 62 and 64 such that the apparatus 40 minimizes quality degradation due to prefetched portion 66 within the retrieved portion 62 of video content 30 and adapts to changes in position due to section 28. That is, in FIG. 2, portions 64 and 66, whose quality is limited as long as their possible combinations are concerned by information 68, are different portions of portion 62, and the transition between both portions 64 and 66 is continuously shifted or adapted to track or travel the moving view section 28. According to an alternative embodiment shown in FIG. 3, apparatus 40 uses information 68 to control the possible combinations of the quality of portions 64 and 66, but according to the embodiment of FIG. 3, they are defined as portions that are distinguished from each other or distinguished from each other in a manner defined, for example, in the media presentation description, i.e., in a manner independent of the position of view section 28. The positions of portions 64 and 66 and the transitions between them may be constant or may vary over time. When varying over time, the variation is due to a change in the content of scene 30. For example, portion 64 corresponds to an area of interest where a higher quality expenditure is worthwhile, while portion 66 is, for example, a portion where a quality degradation due to low bandwidth conditions should be considered before considering a quality degradation for portion 64.
[0044] In the following, further embodiments for an advantageous implementation of apparatus 40 are described. In particular, FIG. 4 shows apparatus 40 in a manner that structurally corresponds to FIGS. 2-3. In FIGS. 2 and 3, the operating mode has been changed such that it corresponds to the second aspect of the present invention.
[0045] That is, the apparatus 40 includes a selection unit 56, a search unit 60, and a derivation unit 66. The selection unit 56 selects from a plurality 46 of media segments 58 provided by the server 20, and the search unit 60 searches for the media segments selected from the server. FIG. 4 assumes that the apparatus 40 operates as illustrated with respect to FIGS. 2-3. That is, the selection unit 56 executes the selection such that the selected media segment 58 has a spatial portion 62 of the scene 30 encoded such that this spatial portion follows a view portion 28 whose spatial position changes over time. However, a variation corresponding to the same aspect of the present application will be described later with respect to FIG. 5, and for each time t, the selected and searched media segment 58 has the entire scene or a certain spatial section 62 encoded therein.
[0046] In any case, similar to the description regarding FIGS. 2 and 3, the selection unit 56 selects the media segment 58, and while a first portion 64 within the section 62 is encoded in the selected and searched media segment 58 with a predetermined quality, a second portion 66 of the section 62 that is spatially adjacent to the first portion 64 is encoded in the selected media segment with a quality lower than the predetermined quality of the portion 64. The first portion 64 completely covers the section 62 while being surrounded by an unencoded portion 72, so that the selection unit 56 restricts the selection and search to a media segment related to a movement template that tracks the position of the viewport 2 8, and a variant in which the media segment is completely encoded in the section 62 with a predetermined quality is shown in FIG. 6. In any case, the selection unit 56 executes the selection such that the first portion 64 follows a view section 28 whose spatial position changes over time.
[0047] In such a situation, it is not easy for the client 40 to predict how large the section 62 or the portion 64 should be. Depending on the content of the scene, most users can act in the same way in moving the view section 28 across the scene 30, and thus the same applies to the interval of the view section 28 at the speed at which the view section 28 is estimated to move across the scene 30. Thus, according to the embodiments of FIGS. 4 to 6, the information 74 is provided to the device 40 by the server 20 to assist the device 40 in setting the size of the first portion 64, or the size and / or position, or the size of the section 62, or the size and / or position, respectively, depending on the information 74. Regarding the possibility of transmitting the information 74 from the server 20 to the device 40, the same applies as described above with respect to FIGS. 2 and 3. That is, the information may be included within the media segment 58, such as within its event box, or a transmission within the media presentation description, or a proprietary message transmitted from the server 40 to the device 40, such as an SAND message, may be used for this purpose.
[0048] That is, according to the embodiments of FIGS. 4-6, the selection unit 56 is configured to set the size of the first portion 64 according to the information 74 originating from the server 20. In the embodiments shown in FIGS. 4-6, the size is set in units of tiles 50, but as already described above with respect to FIG. 1, the situation may be slightly different when using another concept of providing the scene 30 at the server 20 with spatially varying quality.
[0049] According to one example, the information 70 can include, for example, the probability for a given movement speed of the viewport of the view section 28. The information 74 can occur within a media presentation description made available to the client device 40, which can be, for example, a DASH client as already described above, or some in-band mechanisms can be used to convey the information 74, such as in an event box, i.e., an EMSG or SAND message in the case of DASH. The information 74 can also be included in any container format, such as a transport format beyond MPEG-DASH like ISO file format or MPEG-2TS. It can also be transmitted within a video bitstream, such as within an SEI message, as will be described later. In other words, the information 74 can indicate a predetermined value for a measure of the spatial speed of the view section 28. In this way, the information 74 indicates the size of the portion 64 in the form of scaling or in the form of an increment with respect to the size of the view section 28. That is, the information 74 starts from a certain "base size" for the portion 64 required to cover the size of the section 28 and is appropriately increased by incrementally or scaling this "base size". For example, the aforementioned movement speed of the view section 28 can be used to scale the area around the current position of the view section 28 after this time interval, for example, to determine the farthest position around the view section 28 along any possible spatial direction. For example, a waiting time is determined when adjusting the spatial position of the portion 64, such as the duration of the time segment 54 corresponding to the temporal length of the media segment 58. The speed time added to the area around the current position of the viewport 28, which is all-directional, can be such a worst-case area, and can be used to determine the expansion of the portion 64 with respect to a certain minimum expansion of the portion 64 assuming a non-moving viewport 28.
[0050] The information 74 can also be related to the evaluation of user behavior statistics. Later, such an evaluation process Embodiments suitable for performing are described. For example, the information 74 can indicate the maximum speed for a certain percentage of users. For example, the information 74 can indicate that 90% of users move at a speed slower than 0.2 rad / s and 98% of users move at a speed lower than 0.5 rad / s. The information 74 or the message carrying it can be defined such that a probability - speed pair is defined, or such that a message informing the maximum speed for a certain percentage of users, for example always 99% of users, is defined. The movement speed signaling 74 can further include direction information, that is, what is also known as an angle in 2D or 2D plus depth and light field applications in 3D. The information 74 will show different probability - speed pairs for different movement directions.
[0051] In other words, the information 74 can be applied to a given period, such as the temporal length of a media segment. It includes a trajectory - based (X - percentile, average user path) or speed - based pair (X - percentile, speed) or distance - based pair (X - percentile, aperture / diameter / recommendation), or an area - based pair (X - percentile, recommended priority area) or one maximum boundary value of path, speed, distance, or priority area. Instead of associating the information with percentiles, a simple frequency ranking can be done, such that most users move at a certain speed and then the next most users move at a faster speed. further, also, or , the information 74 indicates the speed of the view section 28 not only , Similarly, in order to direct portions 62 and / or 64 that are required to track view section 28, the respective preferred regions to view, According to the display —ed The ratio of users —ed or the display is most frequently recorded The viewing speed / display section of the user —n and A display of whether they match with or without the display of statistical superiority of such displays as , and Of the display time Target eternal Persistence with or withoutcan be shown. Information 74 can show another measure of the speed of view section 28, such as a measure of the moving distance of view section 28 within the time length of the media segment, or more specifically within the time length of time segment 54. Alternatively, information 74 can be shown to distinguish a specific moving direction in which view section 28 can move. This relates to both the indication of the speed or velocity of view section 28 in a specific direction and the indication of the moving distance of view section 28 with respect to a specific moving direction. Further, the expansion of portion 64 can be directly shown by information 74 either omnidirectionally or in a way that distinguishes different moving directions. Further, information 74 can modify all of the examples outlined above in that it shows these values along with the percentage of users for which these values are sufficient to explain the statistical behavior in the moving view section 28. In this regard, it should be noted that the view speed, i.e., view speed section 28, can be quite significant and is not limited, for example, to the speed value of the user's head. Rather, view section 28 can be moved, for example, in response to the user's eye movements, in which case the view speed can be quite large. View section 28 can also be moved in accordance with the movement of other input devices, such as by the movement of a tablet. All of these "input possibilities" that enable the user to move section 28 result in different expected speeds of view section 28, so information 74 can be designed to distinguish different concepts for controlling the movement of view section 28. That is, information 74 can indicate or be shown to indicate the size of portion 64 as having different sizes for different ways of controlling the movement of view section 28, and device 40 will use the size shown by information 74 for proper view section control. That is, view section 28 is controlled by the user, i.e., device 40 obtains knowledge about how view section 28 is controlled, whether by head movement, eye movement, tablet movement, etc., and sets the size according to the portion of information 74 corresponding to such view portion control.
[0052] Generally, the movement speed can be signaled for each content, period, representation, segment, SRD position, pixel, tile, for example, at any temporal or spatial granularity, etc. As outlined, the movement speed can also be distinguished by head movement and / or eye movement. Further, information 74 regarding the user's movement probability may be conveyed as recommendations regarding high-resolution prefetching, i.e., video regions outside the user's viewport, or spherical ranges.
[0053] Figures 7a through 7c briefly summarize some of the options described with respect to information 74 in a method used by apparatus 40 to modify the size and / or the position of portion 64 or portion 62. According to the option shown in Figure 7a, apparatus 40 expands the circumference of section 28 by a distance corresponding to the product of signal speed v and a time length Δt corresponding to the time length of the time segments encoded in individual media segment 50a. Additionally and / or alternatively, the current position of portion 62 and / or 64 can be placed further from the current position of section 28 or from the current position of signals 62 and / or 64 in the signal speed or direction of motion as indicated by information 74, the greater the speed. The speed and direction can be derived from examining or estimating changes in the display of the latest developments or recommended preferred regions according to information 74. Instead of applying v×Δt in all directions, the speed may be indicated by different information 74 for different spatial directions. The alternative shown in Figure 7d shows that information 74 can indicate the distance by which to expand the circumference of view section 28, this distance being indicated by parameter s in Figure 7b. Also, an expansion of the cross-section varying by direction may be applied. Figure 7c shows that the expansion around section 28 can be indicated by information 74 by an increase in area, such as in the form of the ratio of the area of the expanded section compared to the original area of section 28. In any case, the circumference of the expanded region 28 is indicated at 76 in Figures 7a through 7c. The enlarged views of Figures 7a - 7c can be used by selector 56 to dimension or set the dimensions of portion 64 such that portion 64 covers at least its predetermined amount of the entire area within the enlarged portion 76. Clearly, for a larger section 76, a greater number of tiles are, for example, within portion 64. According to a further alternative, section 74 can directly indicate the size of portion 64, such as in the form of the number of tiles making up portion 64.
[0054] The latter possibility of signaling the size of portion 64 is shown in FIG. 5. The embodiment of FIG. 5 can be modified in the same way that the embodiment of FIG. 4 was modified by the embodiment of FIG. 6, i.e., the entire area of section 62 can be retrieved from server 20 by segment 58 with the quality of portion 64.
[0055] In any case, at the end of FIG. 5, information 74 distinguishes between view sections 28 of different sizes, i.e., between different fields of view seen by view section 28. Information 74 simply indicates the size of portion 64 according to the size of view section 28 that device 40 is currently targeting. This way, without the need for a device such as device 40 to calculate or otherwise infer the size of portion 64, as described with respect to FIGS. 4, 6, and 7, regardless of any movement of section 28, different fields of view or devices with view sections 28 of different sizes can use the service of server 20 as long as portion 64 covers view section 28. As is apparent from the description of FIG. 1, for example, it is mostly easy to evaluate how many tiles of a certain number are sufficient to completely cover a field of view of a certain size, i.e., a view section 28 of a certain size, regardless of the orientation of view section 28 for spatial positioning 30. Here, information 74 alleviates this situation, and device 40 can simply look up the value of the size of portion 64 to be used for the size of view section 28 applied to device 40 within information 74. That is, according to the embodiment of FIG. 5, a media presentation description made available to a DASH client, or some convoluted mechanism such as an event box or SAND message, can include information 74 regarding a spherical range or a set of representations or a set of tiles for each field of view. An example is shown in FIG. 1 It can be an offering of tile laying with M expressions. The information 74 can indicate a recommended number n < M (referred to as an expression) to be downloaded to cover the field of view of a given end device from among cubic expressions arranged in a tile-like manner on a 6×4 tile as shown in FIG. 1, for example, and it is considered sufficient to cover a 90°×90° field of view with 12 tiles. Since the field of view of the end device does not necessarily exactly align with the boundaries of the tiles, this recommendation cannot be trivially generated by the device 40 alone. The device 40 can use the information 74, for example, by downloading at least N tiles, that is, media segments 58 regarding N tiles. Another way to utilize the information is to emphasize the quality of the N tiles within the section 62 closest to the current view center of the end device, that is, to use the N tiles to constitute the portion 64 of the section 62.
[0056] With respect to FIG. 8a, an embodiment regarding a further aspect of the invention of the present application is described. Here, FIG. 8a shows a client device 10 and a server 20 that communicate with each other according to any of the possibilities described above with respect to FIGS. 1 to 7. That is, the device 10 can be embodied according to any of the embodiments described with respect to FIGS. 2-7. Alternatively, as described above with respect to FIG. 1, it can simply operate without these details. However, preferably, the device 10 is embodied according to any of the embodiments described above with respect to FIGS. 2-7 or a combination thereof, and further inherits the operation mode now described with respect to FIG. 8a. In particular, the device 10 is internally configured as described above with respect to FIGS. 2-8, that is, the device 40 includes a selection unit 56, a search unit 60, and an optional derivation unit 66. The selection unit 56 makes a selection for the purpose of unequal streaming, that is, the media content is selected and encoded into the searched media segment so that the quality varies spatially and / or there are unencoded parts. However, in addition to this, the device 40 constitutes, for example, a log message transmitter 80 that sends a logged-in log message to the server 20 and the evaluation device 82. The instantaneous magnitude or statistical value that measures the spatial position and / or movement of the first part 64 The instantaneous magnitude or statistical value that measures the quality of the spatially varying scene over time, encoded in the selected media segment and visible in the view section 28, and / or The instantaneous magnitude or statistical value that measures the quality of the first part or the spatially varying scene 30 over time, encoded in the selected media segment and visible in the view section 28
[0057] The motivation is as follows.
[0058] As described above, a reporting mechanism from the user is required to derive statistics such as the most interesting regions and pairs of speed and probability. Additional DASH metrics are required in addition to those defined in Annex D of ISO / IEC 23009-1.
[0059] The DASH client will return the characteristics of the terminal device in terms of FoV as a DASH metric to the metrics server (which can be the same as the DASH server or another one), and the FoV of the client as a DASH metric will be one indicator. TIFF0007684263000001.tif22161
[0060] One metric is the ViewportList. The DASH client returns to the metrics server (which can be the same as the DASH server or another server) the viewports that each client has monitored within a period of time. The instantiation of such a message is as follows. TIFF0007684263000002.tif59159
[0061] Regarding the viewport (region of interest) message, the DASH client can be requested to report each time a viewport change occurs, potentially using a given granularity (regardless of whether to avoid reporting very small movements) or a given periodicity. Such a message can be included in the MPD as the attribute @reportViewPortPeriodicity or as an element or descriptor. It may also be indicated outside the band in the SAND message or other ways.
[0062] The viewport can also be notified at the tile granularity.
[0063] Additionally or alternatively, the log message can report on other current scene-related parameters that change in response to user input, such as the current user distance from the scene center and / or the current view depth, among other parameters discussed below with respect to Figure 10.
[0064] Another metric is the ViewportSpeedList. The DASH client displays the movement speed for a particular viewport when movement occurs. TIFF0007684263000003.tif86159
[0065] This message is sent only when the client performs a viewport movement. However, the server can instruct, as in the previous case, that the message be sent only when the movement is large. Such a setting can indicate the size in pixels, angle, or magnitude that needs to be changed to send the message, such as @ minViewportDifferenceForReporting.
[0066] Another important aspect in the VR-DASH service with asymmetric quality provided as above is to evaluate how quickly the user can switch from an asymmetric representation of the viewport or a set of unequal quality / resolution representations to another more appropriate representation or set of representations for other viewports. With such a metric, the server can derive statistics that help understand the relevant factors affecting QoE. Such a metric is as follows. TIFF0007684263000004.tif88159
[0067] Alternatively, the aforementioned period can be given as an average value. TIFF0007684263000005.tif63164
[0068] Regarding other DASH metrics, for all of these metrics, the time at which the measurement was performed can be additionally included. TIFF0007684263000006.tif20160
[0069] In some cases, if content of uneven quality is downloaded and content of poor quality (or a mixture of good and poor quality) is displayed for a sufficiently long time (it might be just two or three seconds), the user may feel dissatisfied and end the session. On condition of ending the session, the user can send a quality message displayed in the most recent X-hour period. TIFF0007684263000007.tif49160
[0070] Alternatively, the difference in maximum quality, or the max_quality and min_quality of the viewport, can also be reported.
[0071] As is apparent from the above description regarding FIG. 8a, for a tile-based DASH streaming service provider to configure and optimize its service in a meaningful way (e.g., regarding resolution ratio, bitrate, and segment duration), it is advantageous if the service provider can derive statistics that require a client reporting mechanism, examples of which have been described above. In addition to Appendix D of Non-Patent Document 1 ([A1]), additional DASH metrics for those defined above are shown below.
[0072] As shown in FIG. 1, imagine a tile-based streaming service using a cube-projected video. The reconstruction on the client side is shown in FIG. 8b, where small Circle 198 shows the projection onto the image area covered by the individual tiles 50 of a two-dimensional distribution of viewing directions that are equally angularly distributed in the horizontal and vertical directions within the client's viewport 28. The hatched tiles indicate high-resolution tiles and thus form the high-resolution portion 64, while the tiles 50 shown without hatching indicate low-resolution tiles and thus form the low-resolution portion 66. When the viewport 28 is changed after the last update of segment selection and download, the partially low-resolution tiles are presented to the user, from which it can be seen that the resolution of each tile is determined on a cube where there is a projection plane or pixel array of the tiles 50 encoded in the downloadable segment 58.
[0073] The above description generally shows feedback or log messages indicating the quality of the video presented to the user in the viewport. In the following, more specific and advantageous measurement criteria applicable in this regard will be outlined. The measurement criteria described here are reported from the client side and may be called effective viewport resolution. It is assumed to show the service operator the effective resolution in the client's viewport. If the reported effective viewport resolution indicates that the user has only been presented with the resolution towards the resolution of the low-resolution tiles, the service operator can accordingly change the tiling configuration, resolution ratio, or segment length to achieve a higher effective viewport resolution.
[0074] One embodiment would be the average number of pixels within the viewport 28, measured on the projection drawing where there is a pixel array of the tiles 50 encoded in the segment 58. The measurement can distinguish or be specified with respect to the horizontal direction 204 and the vertical direction 206 in relation to the covered field of view (FoV) of the viewport 28. The following table shows possible examples for the appropriate syntax and semantics that may be included in the log message. —ly Shows possible examples for the appropriate syntax and semantics that may be included in the log message. TIFF0007684263000008.tif34140
[0075] The decomposition into horizontal and vertical directions can be interrupted by using scalar values for the average pixel count instead. Along with the indication of the aperture or size of the viewport 28 that can also be reported to the recipient of the log message, i.e., the evaluation unit 82, the average count indicates the pixel density within the viewport.
[0076] It is advantageous to make the FoV considered for the metric smaller than the FoV of the viewport actually presented to the user, thereby excluding the regions towards the boundaries of the viewport that are only used for peripheral vision and thus do not affect the subjective quality perception. This alternative is shown by the dashed line 202 surrounding the pixels in such a central part of the viewport 28. Reporting the FoV 202 considered for the metric reported for the overall FoV of the viewport 28 can also be shown to the log message recipient 82. The following table shows the corresponding extensions for the previous example. TIFF0007684263000009.tif65131
[0077] According to a further embodiment, the average pixel density is not measured by averaging the quality in a spatially uniform manner within the projection plane as was effective in the cases of the examples described so far with respect to the examples including EffectiveFoVResolutionH / V, but this averaging weights the pixels, i.e., unevenly across the projection plane. The averaging may be performed in a spherically uniform manner. As an example, the averaging may be performed uniformly with respect to sample points distributed as the circle 198 is. In other words, the averaging can be performed by weighting the local density with a weight that decreases quadratically with the increase in the local projection plane distance and increases according to the sine of the local slope of the projection with respect to the line connecting to the viewpoint. The message includes an optional (flag-controlled) step, and some of the available projections (such as Equirectangular Projection) can be adjusted, for example, by using a uniform spherical sampling grid. Some projections do not have a large oversampling problem, and forcing the computational removal of oversampling can lead to unnecessary complexity issues. This should not be limited to Equirectangular Projection. In reporting, it is not necessary to distinguish between the horizontal and vertical resolutions, but it is possible to combine them. One embodiment is shown below. TIFF0007684263000010.tif134131
[0078] When applying equiangular uniformity to averaging, in FIG. 8e, points 302 that are horizontally and vertically distributed on sphere 304 centered at viewpoint 306 are shown to be projected onto projection plane 308 of the tile, here onto a cube, as long as they are within viewpoint 28. Thereby, according to the local density of the projection of points 302 onto projection plane 198, the averaging of the pixel density of pixels 308 arranged in rows and columns within the projection plane is performed so as to set the local weight of the pixel density. An exactly similar approach is depicted in FIG. 8f. Here, points 302 are evenly spaced, i.e., evenly distributed vertically and horizontally, within a viewport plane perpendicular to view direction 312. The projection onto projection plane 308 defines point 198, and its local density controls the weight by which local density pixel density 308 (which varies for high-resolution and low-resolution tiles within viewport 28) contributes to the average. In the above example such as the latest table, an alternative of FIG. 8f can be used instead of that shown in FIG. 8e.
[0079] Hereinafter, embodiments regarding a further type of log message related to DASH client 10 (see FIG. 1) having a plurality of media buffers 300, as exemplarily shown in FIG. 8a, i.e., a DASH client 10 that transfers downloaded segment 58 for subsequent decoding by one or more decoders 58, will be described. The distribution of segment 58 onto the buffer can be done in different ways. For example, the distribution can be performed such that specific regions of 360 video are downloaded separately from each other, or buffered in separate buffers after download. The following examples show different distributions by indicating which tiles T (having a total logarithm of 25) indexed #1 to #24, as shown in FIG. 1, are encoded into which representations R#1 to #P of which quality Q of #1 to #M quality (1 being the highest, M being the lowest) can be downloaded individually, and how these P representations R are (optionally) grouped into sets A and how the segments 58 of representation R can be distributed into buffers indexed #1 to #N. It shows whether it can be grouped into set A and how the segments 58 of representation R can be distributed into buffers indexed #1 to #N. TIFF0007684263000011.tif118147
[0080] Here, the representations are provided by the server and notified to the MPD for download, and each of them relates to one tile 50, i.e., one section of the scene. Although it is a representation related to one tile 50, encoding this tile 50 with different qualities, although the grouping is arbitrary, will precisely be summarized in the adaptation set used for association with the buffer. Thus, according to this example, there will be one buffer for each tile 50, in other words, for each viewport (view section) encoding. Another set of representations and distributions is as follows. TIFF0007684263000012.tif119154
[0081] According to this example, each representation will cover the entire area, but the high-quality area focuses on one hemisphere, while the lower quality is used for the other hemisphere. Representations with different exact qualities used in this way, i.e., with the same position of the high-quality hemisphere, are grouped into one adaptation set and, according to this characteristic, are distributed here as an example to six buffers.
[0082] Thus, in the following description, it is assumed that such distribution to the buffer by different viewport encodings, video sub-regions such as tiles, etc., associated with an adaptation set, etc., is applied. Figure 8c shows the buffer fullness level over time for tiles 1 and 2 in a tile-based streaming scenario for two separate buffers, for example, as shown in the last but smallest table. Enabling the client to report the fullness level of all its buffers enables the service operator to correlate other streaming parameters and data to understand the impact on the quality of experience (QoE) of his service settings.
[0083] The advantage therefrom is that the buffer fullness of multiple media buffers on the client side can be reported together with metrics, identified, and associated with the type of buffer. The associated types are, for example ·Tile ·Viewport ·Region ·AdaptationSet ·Representation ·Low quality version of the whole content
[0084] One embodiment of the present invention is given in Table 1, and Table 1 defines metrics for reporting buffer level status events for each buffer together with identification and association are defined. Table 1: List of buffer levels TIFF0007684263000013.tif133170
[0085] A further embodiment using viewport-dependent encoding is as follows.
[0086] In a streaming scenario that depends on the viewport, the DASH client downloads and pre-buffers several media segments related to a specific viewing direction (viewport). If the amount of pre-buffered content is too large and the client changes its display direction, the portion of the pre-buffered content that is played back after the viewport change is not displayed, and each media buffer is purged. This scenario is depicted in Figure 8d.
[0087] Another embodiment relates to multiple representations (quality / bitrate) of the same content, and in some cases, a conventional video streaming scenario where the video content has spatially uniform quality in which it is encoded. In that case, the distribution will be as follows: TIFF0007684263000014.tif81151
[0088] That is, here, each representation may cover a scene of different quality, i.e., spatially uniform quality, although it may not be a panorama 360 scene for example, and these representations will be individually distributed on the buffer. All examples shown in the last three tables should be treated as not limiting the way the segment 58 of the representation provided by the server is distributed to the buffer. Different methods exist, and the rules can be obtained based on the membership of segment 58 in the representation, the membership of segment 58 in the adaptation set, the direction of locally improved quality of the spatially non-uniform coding of the scene into the representation to which each segment belongs, and the orientation. The quality with which the scene is encoded into each segment belongs as follows.
[0089] The client can maintain a buffer for each representation and, when the available throughput increases, decide to erase the remaining low-quality / bitrate media buffer before playback and download the high-quality media segments of the period within the existing low-quality / bitrate buffer. Similar embodiments can be constructed for streaming based on tile and viewport-dependent encoding.
[0090] Since it introduces costs without benefits on the server side and reduces quality on the client side, the service operator may be interested in understanding the amount and type of data downloaded without being presented. Therefore, the present invention is to provide a reporting metric that correlates the two events of "media download" and "media presentation" so that they can be easily interpreted. The present invention avoids the analysis of redundant information reported about the download and playback status of each media segment and enables efficient reporting of only the purge event. The present invention also includes the identification of the buffer as described above and the association with the type. Table 2 shows an embodiment of the present invention. Table 2: List of deletion events TIFF0007684263000015.tif106163
[0091] FIG. 9 shows a further embodiment of how to advantageously implement the apparatus 40. The apparatus 40 of FIG. 9 can correspond to any of the examples described above with respect to FIGS. 1-8. That is, it may perhaps include a lock messenger as described above with respect to FIG. 8a, but need not have, as described above with respect to FIGS. 5 through 7c, or need not have information 68 as described above with respect to FIGS. 2 and 3 or information 74, and may be used. However, with respect to FIG. 9, unlike the description of FIGS. 2 through 8, a tile-based streaming approach is assumed to be actually applied. That is, the scene content 30 is provided to the server 20 in a tile-based manner as described as an option above with respect to FIGS. 2 through 8.
[0092] The internal structure of apparatus 40 may be different from that shown in FIG. 9, but apparatus 40 is illustratively shown to include the selection unit 56 and the search unit 60 already discussed above with respect to FIGS. 2 through 8, and optionally the derivation unit 66. However, further, apparatus 40 includes a media presentation description analyzer 90 and an adaptation unit 92. The MPD analysis unit 90 derives from the media presentation description obtained from the server 20 at least one version in which the spatially varying scene 30 over time is provided for tile-based streaming, and for each of the at least one version, derives an indication of the requirements for benefiting from the spatially varying scene of each version from tile-based streaming. The meaning of "version" will become clear from the following explanation. In particular, the adaptation unit 92 matches the benefit requirements thus obtained to the capabilities of the apparatus 40 or another apparatus that interacts with the apparatus 40, such as the decoding capabilities of one or more decoders 42, the number of decoders 42, etc. The background or concept underlying the FIG. 9 concept is as follows. In a tile-based approach, assuming a view section 28 of a certain size, imagine that a certain number of tiles need to be included in section 62. Further, assume that a media segment belonging to a certain tile forms one media stream or video stream that should be decoded by a separate decoding instantiation, separately from the decoding of media segments belonging to another tile. Thus, the corresponding media segment is selected by the selector 56 The movement aggregation of a number of tiles within section 62 requires, for example, the existence of a corresponding number of decoding resources in the form of respective decoders, i.e., a certain decoding ability such as a corresponding number of decoders 42. If such a number of decoders does not exist, the service provided by server 20 may not be useful to the client. Thus, the MPD provided by server 20 can indicate a "benefit requirement", i.e., the number of decoders necessary to use the provided service. However, server 20 may provide MPDs for different versions. That is, different MPDs for different versions may be available by server 20, or the MPD provided by server 20 may be configured internally to distinguish between different versions in which the service can be used. For example, the versions may differ in the field of view, i.e., the size of the field of view section 28. Different sizes of the field of view are revealed by different numbers of tiles within section 62, and thus the benefit requirements may differ in that, for example, different numbers of decoders may be required for these versions. Other examples are conceivable. For example, versions with different fields of view may include the same plurality of media segments 46, but according to another example, different versions in which scene 30 is provided for tile streaming in server 20 may even differ in the plurality 46 of media segments included according to the corresponding version. For example, the tile splitting according to one version is coarser compared to the tile hyphenation of the scene according to another version, and thus, for example, requires a smaller number of decoders.
[0093] The adaptation unit 92 conforms to the benefit requirement and thus selects the corresponding version or completely rejects all versions.
[0094] However, the beneficial requirements may further relate to profiles / levels that one or more decoders 42 must be able to handle. For example, a DASH MPD contains multiple locations that allow it to indicate a profile. A typical profile describes attributes, elements that may be present in the MPD, and the video or audio profile of the streams provided for each representation.
[0095] A further example of a beneficial requirement relates to, for example, the client-side ability to move the viewport 28 across the entire scene. The beneficial requirement can indicate the necessary viewport speed that should be available for the user to move the viewport in order to be able to truly enjoy the provided scene content. The adaptation section will check, for example, whether this requirement is met, such as whether it is a plugged-in user input device like the HMD 26. Alternatively, assuming that different types of input devices for moving the viewport are associated with typical movement speeds in terms of direction, a set of "sufficient types of input devices" can be indicated as a beneficial requirement.
[0096] In a tile streaming service for spherical video, there are too many configuration parameters that can be dynamically set, such as the number of qualities, the number of tiles, etc. If the tiles are independent bitstreams that need to be decoded by separate decoders, if the number of tiles is too large, it may be impossible for a hardware device with a small number of decoders to decode all the bitstreams simultaneously. The possibility is left as this degree of freedom, for the DASH device to analyze all possible representations and count the number of decoders required to decode a given number that covers the FoV of all the representations or the device, leading to whether the DASH client can consume the content or not. However, a smarter solution for interoperability and function negotiation is to use signaling in the MPD mapped to a kind of profile that is used as a promise to the client that the provided VR content may be consumed if the profile is supported. Such signaling should be done in the form of a URN such as urn :: dash-mpeg :: vr :: 2016, which can be packed either at the MPD level or in an adaptation set This profiling means that it is sufficient to consume the content with N decoders of the X profile. Depending on the profile, the DASH client can ignore or accept a part of the MPD or MPD (adaptation set). Furthermore, there are also some mechanisms that do not contain all information, such as little signaling for selection is available, such as Xlink or MPD chaining. In such a situation, the DASH client cannot derive whether it can consume the content. Regardless of whether it makes sense to execute Xlink or MPD Chaining or a similar mechanism, it is necessary to use such urn (or similar) to expose the decoding capabilities regarding the number of decoders and the profile / level of each decoder so that the DASH client can do so. The signaling can also mean different operating points, such as N decoders with X profiles / levels, or Z decoders with Y profiles / levels.
[0097] FIG. 10 further shows any of the above-described embodiments and descriptions with respect to FIGS. 1-9, and the embodiments of FIGS. 1-9 for clients, apparatus 40, servers, etc. can be extended to the extent that the provided services are extended not only in that the spatio-temporal scene varies temporally but also depends on other parameters. For example, FIG. 10 shows a modification of FIG. 1 depicting scene content 30 for different positions of view center 100 with a plurality of available media segments provided on the server. In the schematic diagram shown in FIG. 10, the scene center is depicted as changing simply along one direction X, but it is clear that the view center can change along two or more spatial directions such as two-dimensional or three-dimensional. This corresponds, for example, to changes in the user at the user's position in a virtual environment. Depending on the position of the user within the virtual environment, the available views change, and accordingly the scene 30 changes. Thus, in addition to tiles and time segments, and media segments describing scene 30 subdivided into different qualities, additional media segments describe different contents of scene 30 for different positions of scene center 100. Apparatus 40 or selection unit 56 respectively calculates the addresses of media segments among a plurality 46 that will be searched within the selection process according to at least one parameter such as the view section position and parameter X, and then these media segments are retrieved from the server using the calculated addresses. For this purpose, the media presentation description may describe a function that depends on the tile index, quality index, and scene center position, similar to time t, resulting in the corresponding addresses of the respective media segments. Thus, according to the embodiment of FIG. 10, the media presentation description will include such calculation rules that depend on one or more additional parameters in addition to the parameters described above with respect to FIGS. 1-9. Parameter X can be quantized to any of the levels at which each scene representation is encoded by the corresponding media segments within the plurality 46 within server 20.
[0098] Alternatively, X can be a parameter that defines the view depth, i.e., the distance from the scene center 100 in the radial direction. Providing scenes in different versions with different view center parts X enables the user to "walk" through the scene, while providing scenes in different versions with different view depths enables the user to "radially zoom" the scene back and forth.
[0099] Therefore, in the case of multiple non-concentric circular viewports, the MPD can use additional signaling regarding the position of the current viewport. The signaling can be done at the segment, representation, or period level, etc.
[0100] Non-concentric spheres: For the user to move, the spatial relationship of different spheres should be signaled within the MPD. This can be done using coordinates (x, y, z) in any unit with respect to the diameter of the sphere. Also, the diameter of the sphere needs to be displayed for each sphere. The sphere should be "sufficiently large" to be used as adding additional space so that there is no problem with the content for the user at the center. If the user moves beyond the indicated diameter, another sphere needs to be used to display the content.
[0101] Exemplary signaling of viewports can be done with respect to a predetermined center point in space. Each viewport is shown based on that center point. In MPEG-DASH, this can be shown, for example, in the AdaptationSet element. TIFF0007684263000016.tif121130
[0102] Finally, FIG. 11 shows that information such as that described above with respect to reference numeral 74 or similar thereto may be present in video bitstream 110 in which video 112 is encoded. A decoder 114 that decodes such a video 110 may use information 74 to determine the size of focus region 116 within video 112 where decoding power for decoding video 110 should be focused. Information 74 may be conveyed, for example, within an SEI message of video bitstream 110. For example, the focus area can be decoded exclusively, or decoder 114 can be configured to start decoding each picture of the video in the focus area rather than, for example, the upper left picture corner, and / or decoder 114 can abort decoding of each picture of the video when focus region 116 has been decoded. Additionally or alternatively, information 74 may simply be present in the data stream for transfer to a subsequent renderer or viewport control or client streaming device or segment selector to determine which spatial sections to cover or which segments to download or stream to cover with improved or predetermined quality. Information 74 may indicate, for example, a preferred area recommended to be positioned to coincide with, or cover, or track this section, such as view section 62 or section 66, as outlined above. It may be used by a client's segment selector. As was the case with the description regarding FIGS. 4-7c, information 74 can set the dimensions of region 116 absolutely, such as the number of tiles, or can set the speed of moving region 116, moving it, for example, in accordance with user input or the like to track interesting content of the video spatiotemporally, thereby scaling area 116 to increase with an increase in the indication of speed.
[0103] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent corresponding method descriptions, and that blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or items or features of corresponding apparatuses. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0104] Signals generated above, such as a streaming signal, an MPD, or any other of the signals described above, can be stored on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0105] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. It can be implemented using a digital storage medium storing electronically readable control signals, such as a floppy disk (registered trademark), a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory. They can cooperate (or be capable of cooperating) with a programmable computer system such that their respective methods are executed. Thus, the digital storage medium can be computer-readable.
[0106] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system such that one of the methods described herein is executed.
[0107] In general, embodiments of the present invention can be implemented as a computer program product having program code, which is operable to execute one of the methods when the computer program product operates on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0108] Other embodiments include a computer program for executing one of the methods described herein, stored on a machine-readable carrier.
[0109] In other words, one embodiment of the method of the present invention is thus a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0110] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0111] Accordingly, a further embodiment of the method of the present invention is a data stream or a series of signals representing a computer program for executing one of the methods described herein. The data stream or series of signals may be configured to be transferred via a data communication connection such as the Internet.
[0112] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.
[0113] Further embodiments include a computer having installed thereon a computer program for executing one of the methods described herein.
[0114] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0115] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to execute some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to execute one of the methods described herein. Generally, the method is preferably executed by any hardware device.
[0116] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0117] The apparatuses described herein, or any component of the apparatuses described herein, can be implemented at least partially in hardware and / or software.
[0118] The methods described herein can be executed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0119] The methods described herein, or any component of the apparatuses described herein, can be executed at least partially by hardware and / or software.
[0120] The above embodiments are merely examples for explaining the principles of the present invention. It is understood that modifications and variations of the configurations and details described herein will be apparent to other persons skilled in the art. Therefore, it is intended to be limited only by the scope of the impending claims and not by the specific details presented for the description and explanation of the embodiments herein.
Prior Art Documents
Non-Patent Documents
[0121]
Non-Patent Document 1
Claims
1. An encoder for encoding a video into a video bitstream such that the video shows a spatial scene and the spatial scene is mapped onto the image of the video using a stereoscopic projection, the encoder comprising a microprocessor or an electronic circuit configured to supply signaling of the size and position of a recommended view section area of the video to an SEI message of the video bitstream, or a computer programmed to supply signaling of the size and position of the recommended view section area of the video to the SEI message of the video bitstream, the recommended view section area forms a viewport section from the spatial scene, the encoder is configured to provide the signaling such that the signaling ranks and shows the size and position multiple times according to a frequency rank obtained from statistics of user actions in conjunction with a display of temporal persistence of the signaling, an omnidirectional or spherical video is encoded in the video bitstream, and the encoder provides signaling of the size and the position of the recommended view section area of the video as a recommendation for spatiotemporal continuation in a viewport to the recommended view section area to the SEI message of the video bitstream, the encoder is configured to encode the video in the video bitstream in units of tiles spatially subdivided from the video, Encoder.
2. The encoder according to claim 1, configured to provide the signaling such that the signaling distinguishes and shows the size and position between operation control methods of different view sections.
3. The encoder according to claim 1, configured to provide the signaling such that the signaling distinguishes and shows the size and position by at least two of view section controls by head movement, eye movement, and tablet movement.
4. A decoder for decoding a video bitstream in which video is encoded, the video showing a spatial scene, the video being encoded in the video bitstream with the spatial scene mapped onto an image of the video using a stereoscopic projection, the decoder comprising a microprocessor or electronic circuit configured to derive signaling of the size and position of a recommended view section region of the video from an SEI message of the video bitstream, or a computer programmed to derive signaling of the size and position of the recommended view section region of the video from an SEI message of the video bitstream, the recommended view section region forming a viewport section from the spatial scene, the signaling showing the size and position multiple times, ranked according to a frequency rank obtained from statistics of user actions, in conjunction with a display of the temporal persistence of the signaling, the video bitstream encoding an omnidirectional or spherical video, the decoder being configured to derive the signaling of the size and the position of the recommended view section region of the video from the SEI message of the video bitstream as a recommendation for spatiotemporally following the recommended view section region with a viewport, the video bitstream encoding the video in units of tiles spatially subdivided by the video, Decoder. **Claim 5** The decoder according to claim 4, wherein the signaling distinguishes and shows the size and position between operation control methods of different view sections. **Claim 6** The decoder according to claim 4, wherein the signaling distinguishes and shows the size and position by at least two of view section control by head movement, eye movement, and tablet movement. **Claim 7** The decoder according to claim 4, configured to transfer the signaling or the size and position information to a renderer or a viewport control or a streaming device.
Citation Information
Patent Citations
IEC23009-1
Video Data Stream Concept
JP2015526006A
Top region of interest in the image
JP2019519981A
Omnidirectional video transmission method, omnidirectional video reception method, omnidirectional video transmission device, and omnidirectional video reception device
JP2019525675A
Spatially Non-Uniform Streaming
JP2019537338A