Method and apparatus for viewport-adaptive 360-degree video distribution
Viewport-adaptive 360° video distribution using layer-based overlay and projection methods optimizes bandwidth usage by encoding and delivering user-specific viewport areas, enhancing the user experience with efficient streaming.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-04-07
AI Technical Summary
The large video size of 360° video leads to high transmission bandwidth requirements, with significant bandwidth consumption during streaming, as the entire video is encoded in high quality, despite users often viewing only a small portion of the image.
Viewport-adaptive 360° video distribution is achieved through layer-based viewport overlay and signaling, using projection methods like equirectangular, cube-map, cylindrical, and pyramidal projections, and bitstream enhancement to distribute 360° video based on user-specific viewports, allowing flexible encoding and decoding.
This approach reduces bandwidth consumption by encoding and delivering only the necessary viewport areas at high quality, enhancing user experience with efficient and adaptive video streaming.
Smart Images

Figure 0007842148000031 
Figure 0007842148000032 
Figure 0007842148000033
Abstract
Description
[Background technology]
[0001] Cross-reference of related applications This application claims priority and interest of U.S. Provisional Patent Application No. 62 / 342158, filed on 26 May 2016, which is incorporated herein by reference.
[0002] 360° video is a rapidly developing format in the media industry. This is made possible by the increased availability of virtual reality (VR) devices. 360° video can offer viewers a new sense of immersion. Compared to rectilinear video (e.g., 2D or 3D), 360° video can present significant technical challenges in video processing and / or distribution. High video quality and / or very low latency may be required to enable a comfortable and / or immersive user experience. The large video size of 360° video can be an obstacle to distributing it on a large scale with high quality.
[0003] 360° video applications and / or services may encode the entire 360° video into a standards-compliant stream for progressive download and / or adaptive streaming. Delivering the entire 360° video to the client can enable low-latency rendering (for example, the client can access the entire 360° video content and / or choose to render only the parts they want to see without further restrictions). From the server's perspective, the same stream can support multiple users, possibly using different viewports. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] The video size can become very large, which can lead to high transmission bandwidth requirements when the video is streamed (for example, because the entire 360° video needs to be encoded in high quality, such as 4K at 60fps per eye or 6K at 90fps). For example, users may only see a small portion of the entire image (e.g., the viewport), so the high bandwidth consumption during streaming can be wasted. [Means for solving the problem]
[0005] A system, method, and means for viewport-adaptive 360° video distribution are disclosed. Viewport enhancement-based 360° video may be distributed and / or signaled. 360° video may be distributed using a layer-based viewport overlay. Signaling of 360° video mapping may be provided.
[0006] A first viewport of the 360-degree video may be determined. The 360-degree video may include one or more of the following mappings: equirectangular, cube-map, cylindrical, pyramidal, and / or spherical projection. The first viewport may be associated with a spatial region of the 360-degree video. Adjacent areas extending around the spatial region may be determined. A second viewport of the 360-degree video may be determined. A bitstream associated with the 360-degree video may be received. The bitstream may include one or more enhanced regions. One or more enhanced regions may correspond to the first viewport and / or the second viewport. A high encoding bitrate may be associated with the first viewport and / or the second viewport. Signaling indicating one or more viewport properties associated with the 360-degree video distribution may be received.
[0007] A WTRU for processing 360-degree video may include a processor configured (for example, by executable instructions stored in memory) to perform one or more of the following: (i) receive a media presentation description (MPD) associated with 360-degree video, which includes an essential property element indicating a face-packing layout for a multi-face geometric projection format of the media segment; (ii) receive a media segment; (iii) determine at least one face-packing layout for the received media segment from a set of face-packing layouts based on the essential property element; and (iv) construct the received media segment based on the determined at least one face-packing layout.
[0008] The set of surface packing layouts includes plate carree, poles on the side half height, poles on the side full height, single row, 2x3, and 180 degrees. Essential property elements can be either at the adaptation level or the representation level.
[0009] The WTRU processor may be configured (for example, by executable instructions stored in memory) to determine the media representation associated with the MPD in order to request future media segments and to send requests for the determined media representation.
[0010] The MPD may include a video type selected from a set of video types for a media segment. The set of video types may include linear, panoramic, spherical, and lightfield. The WTRU processor may be configured (for example, by executable instructions stored in memory) to determine the video type for an received media segment and / or construct the received media segment using the determined video type.
[0011] The MPD may include at least one projection format used to project 360-degree video from an omnidirectional format onto a linear video frame. The projection format may include one of the following: equirectangular projection, cube, offset cube, squished sphere, pyramid, and cylinder. The WTRU processor may be configured (e.g., by executable instructions stored in memory) to determine the projection format for receiving the video file and / or send a request for the determined projection format. The 360-degree video may include an Omnidirectional Media Application Format (OMAF) file.
[0012] A method of using a WTRU to process 360-degree video may include one or more of the following steps: (i) receiving a Media Presentation Description (MPD) associated with the 360-degree video, which includes an Essential Property element indicating a face packing layout for a multi-face geometric projection format of the media segment; (ii) receiving a media segment; (iii) determining at least one face packing layout for the received media segment from a set of face packing layouts based on the Essential Property element; and (iv) constructing the received media segment based on the determined at least one face packing layout.
[0013] The above method using WTRU may include the steps of determining a media representation associated with the MPD to request future received video files, and / or sending a request for the determined media representation. The above method may include the steps of determining a video type for a received media segment, and / or constructing the received media segment using the determined video type. The above method may include the steps of determining a projection format for a video file, and / or sending a request for the determined projection format.
[0014] A WTRU for processing 360-degree video may include a processor configured (for example, by executable instructions stored in memory) to perform one or more of the following: receiving a media presentation description (MPD) associated with 360-degree video, the MPD comprising a first essential property element indicating a first face packing layout for a polyhedral geometric projection format of a media segment and a second essential property element indicating a first second packing layout for a polyhedral geometric projection format of a media segment; determining whether to use the first or second face packing layout for the media segment; requesting at least the determined first or second face packing layout; receiving a media segment; and reconstructing the 360-degree video associated with the received media segment based on the requested face packing layout. [Effects of the Invention]
[0015] This invention provides a system, method, and means for novel viewport-adaptive 360° video distribution. [Brief explanation of the drawing]
[0016] [Figure 1] A diagram showing an exemplary portion of a 360° video displayed on a head-mounted device (HMD). [Figure 2] A diagram showing an exemplary orthographic cylindrical projection of a 360° video. [Figure 3] A diagram showing an exemplary 360° video mapping. [Figure 4] A diagram showing an exemplary media presentation description (MPD) hierarchical data model. [Figure 5] A diagram showing an exemplary dynamic adaptive streaming over HTTP (DASH) spatial relationship description (SRD) for video. [Figure 6] A diagram showing an exemplary tile-based video segmentation. [Figure 7] A diagram showing an exemplary tile set with constrained temporal motion. [Figure 8] A diagram showing an exemplary 360° video streaming quality degradation. [Figure 9] A diagram showing an exemplary viewport area with associated adjacent areas. [Figure 10A] A diagram showing an exemplary cube map layout. [Figure 10B] A diagram showing an exemplary cube map layout. [Figure 10C] A diagram showing an exemplary cube map layout. [Figure 10D] A diagram showing an exemplary cube map layout. [Figure 11A] A diagram showing an exemplary orthographic cylindrical coordinate viewport mapping. [Figure 11B] A diagram showing an exemplary cube map coordinate viewport mapping. [Figure 12] A diagram showing an exemplary spherical coordinate viewport mapping. [Figure 13A]This is an example of a viewport enhancement representation. [Figure 13B] This is an example of a viewport enhancement representation. [Figure 14] This figure shows an example of a layer-based 360° video overlay. [Figure 15] This diagram illustrates an example of a layer-based 360° video representation. [Figure 16] This figure shows an example of a layer-based 360° video overlay. [Figure 17] This is an illustrative flowchart of a layer-based 360° video overlay. [Figure 18] This figure shows an exemplary layer-based 360° video overlay with multiple viewports. [Figure 19] This figure shows an exemplary equirectangular projection representation with half-height poles on the sides. [Figure 20] This figure shows an exemplary equirectangular projection representation with full-height poles on the sides. [Figure 21] This figure shows an example of a single-column layout cube representation. [Figure 22] This is a diagram illustrating an exemplary 2x3 layout cube representation. [Figure 23] This figure shows an exemplary 180° cubemap layout. [Figure 24A] This is a diagram of an exemplary communication system in which one or more disclosed embodiments may be implemented. [Figure 24B] Figure 24A is a system diagram of an exemplary wireless transceiver unit (WTRU) that may be used in the communication system shown. [Figure 24C] Figure 24A is a system diagram of an exemplary radio access network and core network that may be used within the communication system shown. [Figure 24D]Figure 24A is a system diagram of another exemplary radio access network and core network that may be used within the communication system shown in Figure 24A. [Figure 24E] Figure 24A is a system diagram of another exemplary radio access network and core network that may be used within the communication system shown in Figure 24A. [Modes for carrying out the invention]
[0017] Herein, a detailed description of exemplary embodiments is provided with reference to various figures. This description provides detailed examples of possible implementations, but it should be noted that the details are illustrative and not intended to limit the scope of this application in any way.
[0018] Figure 1 shows an exemplary portion of a 360° image displayed on a head-mounted device (HMD). When viewing a 360° image, the user may be presented with a portion of the image, as shown in Figure 1, for example. The portion of the image may change as the user looks around and / or zooms in on the image. The portion of the image may change based on feedback provided by the HMD or other types of user interfaces (e.g., a wireless transceiver unit (WTRU)). The viewport may be the spatial region of the entire 360° image, or may include such a spatial region. The viewport may be presented to the user entirely or partially. The viewport may have one or more qualities that differ from other parts of the 360° image.
[0019] 360° video can be captured and / or rendered on a sphere (for example, to give the user the ability to choose any viewport). Spherical video formats can be delivered directly using conventional video codecs. 360° video (e.g., spherical video) can be compressed by projecting the spherical video onto a 2D plane using a projection method. The projected 2D video can be encoded (for example, using conventional video codecs). Examples of projection methods include equirectangular projection. Figure 2 shows an exemplary equirectangular projection of 360° video. For example, an equirectangular projection can map a first point P having coordinates (θ,φ) on the sphere to a second point P having coordinates (u,v) on the 2D plane using one or more of the following equations: u = φ / (2*pi) + 0.5 (1) v = 0.5 - θ / (pi) (2)
[0020] Figure 3 shows an exemplary 360° video mapping. For example, one or more other projection methods (e.g., mappings) can be used to convert 360° video to a 2D planar video (for example, to reduce bandwidth requirements). For example, one or more other projection methods may include pyramidal maps, cube maps, and / or offset cube maps. One or more other projection methods may be used to represent spherical video with less data.
[0021] Viewport-specific representations may be used. One or more of the projection methods shown in Figure 3, such as cubemap and / or pyramidal projection, may provide representations of non-uniform quality for different viewports (for example, some viewports may be represented with higher quality than others). Multiple versions of the same image with different target viewports may be generated and / or stored on the server side (for example, to support all viewports of a spherical image). For example, Facebook's implementation of VR video delivery may use the offset cubemap format shown in Figure 3. An offset cubemap can provide the highest resolution (e.g., highest quality) for the front viewport, the lowest resolution (e.g., lowest quality) for the back view, and intermediate resolution (e.g., intermediate quality) for one or more side views. The server may store different versions of the same content (for example, to accept client requests for different viewports of the same content). For example, a total of 150 different versions of the same content (e.g., 30 viewports × 5 resolutions per viewport). During delivery (e.g., streaming), a client can request a specific version corresponding to its current viewport. This specific version may be delivered by the server.
[0022] To describe various projections that may be used to represent 360-degree video and / or other unconventional video formats (e.g., cylindrical video that may be used in panoramic video representations), ISO / IEC / MPEG may define the Omnidirectional Media Application Format (OMAF). OMAF file format metadata for the projections described herein may include support for projection metadata regarding video onto spheres, cubes, cylinders, and / or pyramids. Table 1 may show example OMAF syntax for supporting projections such as flattened spheres, cylinders, and pyramids.
[0023] [Table 1]
[0024] HTTP streaming has become the dominant method in commercial deployments. For example, streaming platforms such as Apple's HTTP Live Streaming (HLS), Microsoft's Smooth Streaming (SS), and / or Adobe's HTTP Dynamic Streaming (HDS) can use HTTP streaming as the underlying delivery method. The HTTP streaming standard for multimedia content can enable standards-based clients to stream content from any standards-based server (for example, thereby enabling interoperability between servers and clients from different vendors). MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH) can be a general-purpose delivery format that provides the best possible video experience to end users by dynamically adapting to changing network conditions. DASH can be built on top of the HTTP / TCP / IP stack. DASH can define manifest formats, Media Presentation Descriptions (MPDs), and segment formats for ISO-based media file formats and MPEG-2 transport streams.
[0025] Dynamic HTTP streaming can be associated with various bitrate options for multimedia content made available on the server. Multimedia content can include several media components (e.g., audio, video, text), each of which may have different characteristics. In MPEG-DASH, these characteristics can be described by a Media Presentation Description (MPD).
[0026] An MPD can be an XML document containing metadata necessary for a DASH client to construct appropriate HTTP-URLs to access video segments in an adaptive manner during a streaming session (for example, as described herein). Figure 4 shows an exemplary Media Presentation Description (MPD) hierarchical data model. An MPD can describe a sequence of Periods, during which a consistent set of encoded versions of media content components remains unchanged. A Period can have a start time and a duration. A Period can consist of one or more adaptation sets (e.g., an AdaptationSet).
[0027] An AdaptationSet can represent a set of encoded versions of one or more media content components that share one or more identical properties (e.g., language, media type, picture aspect ratio, role, accessibility, viewpoint, and / or rating properties). For example, a first AdaptationSet could contain different bitrates for the video component of the same multimedia content. A second AdaptationSet could contain different bitrates for the audio component of the same multimedia content (e.g., lower quality stereo and / or higher quality surround sound). An AdaptationSet can contain multiple Representations.
[0028] A Representation can describe a deliverable encoded version of one or more media components that differs from other representations by bitrate, resolution, number of channels, and / or other characteristics. A Representation can contain one or more segments. One or more attributes of the Representation element (e.g., @id, @bandwidth, @quality Ranking, and @dependencyId) may be used to specify one or more properties of the associated Representation.
[0029] A segment can be the largest unit of data that can be retrieved in a single HTTP request. A segment can have a URL (for example, an addressable location on a server). Segments can be downloaded using HTTP GET or HTTP GET with a byte range.
[0030] A DASH client can parse an MPD XML document. A DASH client can select a set of AdaptationSets appropriate for its environment based on information provided, for example, within an AdaptationSet element. Within an AdaptationSet, the client can select a Representation. The client can select a Representation based on the value of the @bandwidth attribute, client decoding capabilities, and / or client rendering capabilities. The client can download an initialization segment of the selected Representation. The client can access the content (for example, by requesting the entire Segment or a byte range of the Segment). Once the presentation begins, the client can continue consuming media content. For example, the client can request (e.g., continuously request) a Media Segment and / or a portion of a Media Segment during the presentation. The client can play the content according to the media presentation timeline. The client can switch from a first Representation to a second based on updated information from the client's environment. The client can play the content continuously over two or more Periods. When the client is consuming the media contained in the Segment toward the end of the published media in the Representation, the Media Presentation may be terminated, the Period may be started, and / or the MPD may be re-fetched.
[0031] MPD descriptor elements, or Descriptors, may be provided to an application (for example, to instantiate one or more Descriptors with appropriate scheme information). One or more Descriptors (e.g., content protection, role, accessibility, rating, viewpoint, framepacking, and / or UTC timing descriptors) may include an @schemeIdUri attribute to identify the relative scheme.
[0032] SupplementalProperty can contain metadata that can be used by DASH clients to optimize processing.
[0033] An EssentialProperty can contain metadata for processing the elements it contains.
[0034] The Role MPD element can use the @schemeIdUri attribute to identify the role scheme used to identify the role of a media content component. One or more Roles can define and / or describe one or more characteristics and / or structural functions of a media content component. An Adaptation Set and / or media content component can have multiple designated roles (even within the same scheme).
[0035] MPEG-DASH can provide a Spatial Relation Description (SRD) scheme. The SRD scheme can represent the spatial relationship of images, where two MPD elements (e.g., AdaptationSet and SubRepresentation) represent spatial portions of another full-frame image. SupplementalProperty and / or EssentialProperty descriptors with @schemeIdURI equal to "urn:mpeg:dash:srd:2014" can be used to provide spatial relationship information associated with AdaptationSet and / or SubRepresentation. The attribute @value of the SupplementalProperty and / or EssentialProperty elements can provide one or more values for SRD parameters such as source_id, object_x, object_y, object_width, object_height, total_width, total_height, and / or spatial_set_id. The values and semantics of SRD parameters can be defined as shown in Table 2.
[0036] [Table 2-1]
[0037] [Table 2-2]
[0038] Figure 5 shows an exemplary DASH SRD video. An SRD can represent that a video stream represents a spatial portion of a full-frame video. The spatial portion may be a tile and / or region of interest (ROI) of the full-frame video. An SRD can describe a video stream in terms of the position (object_x, object_y) and / or size (object_width, object_height) of the spatial portion relative to the full-frame video (total_width, total_height). SRD descriptions can provide flexibility for clients with respect to adaptation. An SRD-aware DASH client can use one or more SRD annotations to select a full-frame representation and / or a spatial portion of a full-frame representation. By using one or more SRDs to select a full-frame representation or a spatial portion, bandwidth and / or client-side computation can be saved, for example, by avoiding full-frame fetching, decoding, and / or cropping. By using one or more SRD annotations to decide which representation to select, the quality of a given spatial portion of a full-frame video (e.g., region of interest i.e., ROI) can be increased, for example, after zooming. For example, a client can request a first video stream corresponding to a higher-quality ROI space portion, and a second video stream that does not correspond to a lower-quality ROI, without increasing the overall bitrate.
[0039] Table 3 shows an example of an MPD supporting SRD for the scenario shown in Figure 5, where each tile has a resolution of 1920×1080, and the entire frame has a resolution of 5760×3240 with 9 tiles.
[0040] [Table 3-1]
[0041] [Table 3-2]
[0042] [Table 3-3]
[0043] Figure 6 shows an exemplary tile-based video segmentation. A 2D frame can be segmented into one or more tiles as shown in Figure 6. Given MPEG-DASH SRD support, tile-based adaptive streaming (TAS) can be used to support features such as zooming and panning in large panoramas, spatial resolution enhancement, and / or server-based mosaic services. An SRD-aware DASH client can use one or more SRD annotations to select between a full-frame representation and a tile representation. By deciding whether to select a full-frame representation of a tile representation using one or more SRD annotations, bandwidth and / or client-side computations can be saved (e.g., by avoiding full-frame fetching, decoding, and / or cropping).
[0044] Figure 7 shows an exemplary time-motion-constrained tileset. A frame can be encoded into a bitstream containing several time-motion-constrained tilesets specified in HEVC. Each of these time-motion-constrained tilesets can be decoded independently. As shown in Figure 7, two or more left tiles can form a time-motion-constrained tileset that can be decoded independently (for example, without decoding the entire picture).
[0045] 360° video content can be delivered primarily via HTTP-based streaming solutions such as progressive download or DASH adaptive streaming. Dynamic streaming for 360° video, based on UDP instead of HTTP, has been proposed (for example, by Facebook) to reduce latency.
[0046] To represent 360-degree video of sufficient quality, 2D projections may need to have high resolution. When a full 2D layout is encoded at high quality, the resulting bandwidth may be too large for effective delivery. The amount of data can be reduced by using several mappings and / or projections to allow different parts of the 360-degree video to be represented at different quality levels. For example, the front view (e.g., a viewport) may be represented at high quality, while the back view (e.g., the opposite) may be represented at lower quality. One or more other views may be represented at one or more other intermediate quality levels. The offset cubemap and pyramidal map shown in Figure 3 may be examples of mappings and / or projections that represent different parts of the 360-degree video at different quality levels. Pyramidal maps can reduce the number of pixels and / or save bitrate for each viewport, while multiple viewport versions can handle different viewing positions that the client may request. Viewport-specific representations and / or deliveries may have greater latency to adapt as the user changes their viewing position compared to delivering the entire 360-degree video. For example, using an offset cubemap can result in a decrease in image quality when the user's head position rotates 180 degrees.
[0047] Figure 8 illustrates an example of 360 video streaming quality degradation. DASH segment length and / or client buffer size can affect display quality. Longer segment lengths can result in higher encoding efficiency. Longer segment lengths may also make it difficult to quickly adapt to viewport changes. As shown in Figure 8, a user may have three possible viewports (e.g., A, B, and C) of 360 video. One or more (e.g., three) segment types, i.e., S A S B and S C However, it can be associated with 360-degree video. Each of one or more segment types can be responsible for a corresponding viewport of higher quality and other viewports of lower quality. The user can have segment S responsible for higher quality video for viewport A and lower quality video for viewports B and C. A During playback, the user can pan from viewport A to viewport B at time t1. The user can then select the next segment (S) which will be responsible for the higher quality video for viewport B and the lower quality video for viewports A and C. B Before switching to viewport B, the user may have to watch a lower-quality viewport B. Such a negative user experience can be resolved with shorter segment lengths. Shorter segment lengths may reduce encoding efficiency. When the user's streaming client logic pre-downloads too many segments based on the previous viewport, the user may have to watch lower-quality video. To prevent the user's streaming client from pre-downloading too many segments based on the previous viewport, the streaming buffer size can be reduced. A smaller streaming buffer size may affect streaming quality adaptation, for example, by causing buffer underflows more frequently.
[0048] Viewport-adaptive 360° video streaming can include viewport-enhanced based delivery and / or layer-based delivery.
[0049] Efficient 360 video streaming can take into account both bitrate adaptation and viewport adaptation. Viewport-enhanced based 360° video delivery can include encoding one or more identified viewports of a high-quality 360 video frame. For example, viewport-based bit allocation can be performed during encoding. Viewport-based bit allocation can allocate a larger portion of bits to one or more viewports and / or allocate a reduced amount of bits to other areas accordingly. FIG. 9 shows an exemplary viewport area having associated adjacent areas. A viewport area, adjacent areas, and other areas can be determined for a 360° video frame.
[0050] A bitrate weight for the viewport area can be defined as α. A bitrate weight for the adjacent areas can be defined as β. A bitrate weight for the other areas can be defined as γ. One or more of the following equations can be used to determine the target bitrate for each area. α + β + γ = 1 (3) BR HQ = R × α (4) BR MQ = R × β (5) BR LQ = R × γ (6) Where R can represent a constant encoding bitrate for the entire 360° video, BR HQ can represent the encoding bitrate for the target viewport area, BR MQ can represent the encoding bitrate for the viewport adjacent areas, and / or BR LQ can represent the encoding bitrate for the other areas.
[0051] In equation (3), the values of α, β, and γ can sum to 1, which can mean that the overall bitrate is kept the same. For example, bits may simply be redistributed between different areas (e.g., the viewport, adjacent areas, and other areas). At the overall bitrate of R, the same video may be encoded into different versions. Each of the different versions may be associated with a different viewport quality level. Each of the different viewport quality levels may correspond to a different value of α.
[0052] The overall bitrate R may not be maintained constant. For example, the viewport area, adjacent areas, and other areas may be encoded to the target quality level. In this case, the representation bitrate for each area may differ.
[0053] The projection methods described herein (e.g., offset cubemaps, pyramidal maps, etc.) can enhance the quality of one or more target viewports and / or reduce the quality of other areas of the image.
[0054] On the server side, an AdaptationSet can contain multiple Representation elements. A Representation of a 360 video stream can be encoded to a specific resolution and / or bitrate. A Representation can be associated with one or more viewports with specific quality enhancements. A Viewport Relationship Description (VRD) can specify one or more corresponding viewport spatial coordinate relationships. A SupplementalProperty and / or EssentialProperty descriptor whose @schemeIdUri is equal to "urn:mpeg:dash:viewport:2d:2016" may be used to provide VRDs associated with AdaptationSet, Representation, and / or Sub-Representation elements.
[0055] The @value of SupplementalProperty and / or EssentialProperty elements using the VRD scheme can be a comma-separated list of viewport description parameter values. Each AdaptationSet, Representation, and / or Sub-Representation can contain one or more VRDs to represent one or more enhanced viewports. The properties of an enhanced viewport can be described by VRD parameters as shown in Table 4.
[0056] [Table 4]
[0057] Table 5 provides an MPD example for 4K 360 video. The AdaptationSet can be annotated by two SupplementalProperty descriptors having the VRD scheme identifier "urn:mpeg:dash:viewport:2d:2016". The first descriptor can specify an enhanced 320x640 viewport #1 at (150,150) with quality level 3. The second descriptor can specify an enhanced 640x960 viewport #2 at (1000,1000) with quality level 5. Both viewports can represent the spatial portion of a 4096x2048 full-frame 360° video. There may be two Representations, the first Representation may be at full resolution and the second Representation may be at half resolution. The viewport position and / or size of a Representation can be identified based on the values of the VRD attribute @full_width and / or @full_height, and / or the Representation attribute @width and / or @height. For example, viewport #1 in a full-resolution Representation (e.g., @width=4096 and @height=2048) could be (150,150) at a size of 320x640, while viewport #1 in a half-resolution Representation (e.g., @width=2048 and @height=1024) could be (75,75) at a size of 160x320. Depending on one or more of the user's WTRU capabilities, the half-resolution may be extended to full-resolution.
[0058] [Table 5]
[0059] Projections (e.g., equirectangular projection, cubemap, cylinder, and pyramidal projection) can be used to map the surface of a spherical image onto a flat image for processing. One or more layouts may be available for a particular projection. For example, different layouts may be used for equirectangular projection and / or cubemap projection formats. Figures 10A–10D show exemplary cubemap layouts. Cubemap layouts can include cube layouts, 2x3 layouts, side-pole layouts, and / or single-row layouts.
[0060] The VRDs in Table 4 may be specified for a specific projection format, such as equirectangular projection. The VRDs in Table 4 may also be specified for a specific projection and layout combination, such as cubemap projection and a 2x3 layout (e.g., layout B shown in Figure 10B). The VRDs in Table 4 may be extended to support various projection formats and / or combinations of projection and layout formats simultaneously, as shown in Table 6. The server may support a set of common projections and / or projection and layout formats using, for example, the signaling syntax shown in Table 6. Figure 11 shows exemplary viewport coordinates in equirectangular and cubemap projections. One or more (e.g., two) viewports of a 360 image may be identified. As illustrated, viewport coordinate values may differ between the equirectangular projection format and the cubemap projection format.
[0061] A VRD can specify one or more viewports associated with a Representation element. Table 7 shows an MPD example where both equirectangular and cubemap representations are provided, with corresponding VRDs signaled accordingly. One or more first VRDs (e.g., VRD@viewport_id=0,1) can specify one or more properties of a viewport in one or more first Representations (e.g., Representation@id=0,1) in an equirectangular projection format. One or more second VRDs (e.g., VRD@viewport_id=2,3) can specify one or more properties of a viewport in one or more second Representations (e.g., Representation@id=2,3) in a cubemap projection format.
[0062] A VRD can specify one or more viewport coordinates in one or more (e.g., all) common projection / layout formats (even if the associated Representation is a single projection / layout format). Table 8 shows an MPD example where a Representation for an equirectangular projection is provided, and VRDs for corresponding viewports in both equirectangular and cubemap are provided. For example, projection and layout formats can be signaled at the Representation level as described herein.
[0063] The server can specify viewports in different formats and / or give the client flexibility to choose an appropriate viewport (for example, depending on the client's capabilities and / or technical specifications). If the client cannot find a preferred format (for example, from the methods specified in Table 4 or the set of methods specified in the table), the client can convert one or more viewports from one of the specified formats to the format the client wishes to use. For example, one or more viewports may be specified in cubemap format, but the client may want to use equirectangular projection format. Based on the projection and / or layout description available in the MPD, the client can derive a user orientation position from its gyroscope, accelerometer, and / or magnetometer tracking information. The client can convert the orientation position to a corresponding 2D position on a particular projection layout. The client can request a Representation with identified viewports based on the values of one or more VRD parameters such as @viewport_x, @viewport_y, @viewport_width, and / or @viewport_height.
[0064] [Table 6]
[0065] [Table 7-1]
[0066] [Table 7-2]
[0067] [Table 8-1]
[0068] [Table 8-2]
[0069] A generic viewport descriptor may be provided as shown in Table 9. The generic viewport descriptor can describe a viewport position using spherical coordinates (θ,φ) as shown in Figure 2, where θ can represent the tilt angle or polar angle, φ can represent the azimuth angle, and / or the normalization radius may be 1.
[0070] A SupplementalProperty and / or EssentialProperty descriptor where @schemeIdUri is equal to "urn:mpeg:dash:viewport:sphere:2016" may be used to provide a spherical coordinate-based VRD associated with AdaptationSet, Representation, and / or Sub-Representation elements. Regions specified in spherical coordinates may not correspond to a rectangular region on the projected 2D plane. If a region does not correspond to a rectangular region on the projected 2D plane, the bounding rectangle of the signaled region may be derived and / or used to specify the viewport.
[0071] The @value of SupplementalProperty and / or EssentialProperty elements that use VRDs can be a comma-separated list of values for one or more viewport description parameters. Each AdaptationSet, Representation, and / or Sub-Representation can contain one or more VRDs to represent one or more enhanced viewports. The properties of an enhanced viewport can be described by parameters as shown in Table 9.
[0072] [Table 9]
[0073] Figure 12 shows an exemplary spherical coordinate viewport. With respect to the viewport (shown as V in Figure 12, for example), viewport_inc can specify the polar angle θ, viewport_az can specify the azimuth angle (φ), viewport_delta_inc can specify dθ, and / or viewport_delta_az can specify dφ.
[0074] To distinguish between VRDs that use 2D coordinates (e.g., Table 4 or Table 6) and VRDs that use spherical coordinates (e.g., Table 9), different @schemeIdUri values may be used in each case. For example, for a 2D viewport descriptor, the @schemeIdUri value may be urn:mpeg:dash:viewport:2d:2016. For a viewport based on spherical coordinates, the @schemeIdUri value may be "urn:mpeg:dash:viewport:sphere:2016".
[0075] Using spherical coordinate-based viewport descriptors reduces signaling costs because a viewport can only be specified for one coordinate system (e.g., a spherical coordinate system). Using spherical coordinate-based viewport descriptors simplifies the transformation process on the client side, as each client only needs to implement a predefined transformation process. For example, each client can transform between the projection format they choose to use (e.g., equirectangular projection) and the spherical representation. When a client aligns viewport coordinates, they can use similar logic to determine which representation to request.
[0076] A VRD can be signaled by one or more PeriodSupplementalProperty elements. Each VRD can list the available (e.g., all available) enhanced viewports. A Representation can signal one or more viewport indices using the attribute @viewportId to identify which enhanced viewports are associated with the current Representation or Sub-Representation. Using such descriptor referencing techniques can avoid redundant signaling of VRDs within a Representation. One or more Representations with different associated viewports, projection formats, and / or layout formats can be assigned within a single AdaptationSet. Table 10 shows exemplary semantics for the Representation element attributes @viewportId and @viewport_quality.
[0077] As shown in Table 10, the @viewport_quality attribute can be signaled at the Representation level (for example, instead of the Period level) as part of the VRD. The @viewport_quality attribute allows the client to select an appropriate quality level for one or more viewports of interest. For example, if the user frequently navigates the 360 view (e.g., constantly turning their head to look around), the client can select a Representation with a balanced quality between viewports and non-viewports. If the user is focused on a viewport, the client can select a Representation with high viewport quality (but relatively reduced quality in non-viewport areas, for example). The @viewport_quality attribute can be signaled at the Representation level.
[0078] [Table 10]
[0079] Table 11 shows an example of an MPD using the Representation attributes @viewportId and @viewport_quality shown in Table 10 to specify the associated viewport. Two VRDs may be specified in Period.SupplementalProperty. The first Representation (e.g., @id=0) may contain one or more viewports (e.g., @viewportId=0) where viewport #0 has a quality level of 2. The second Representation (e.g., @id=1) may contain the same one or more viewports as the first Representation (e.g., @viewportId=0), but viewport #0 has a quality level of 4. The third Representation (e.g., @id=2) may contain one or more enhanced viewports (e.g., @viewportId=1) with a quality level of 5. The fourth Representation (e.g., @id=3) may contain one or more enhanced viewports (e.g., @viewportId=0,1) with the highest quality level (e.g., @viewport_quality=5,5).
[0080] [Table 11-1]
[0081] [Table 11-2]
[0082] The enhanced viewport area can be selected to be larger than the actual display resolution, thereby mitigating quality degradation when the user changes the viewport (e.g., slightly). The enhanced viewport area can be selected to cover the area around which the target object of interest moves during a segment period (e.g., so that the user can focus on the target object as it moves around with the same quality). For example, one or more of the most focused viewports may be enhanced in a single bitstream to reduce the total number of Representations and / or corresponding media streams on the origin server or CDN. The viewport descriptors specified in Tables 4, 6, 9, and / or 10 can support multiple quality-enhanced viewports within a single bitstream.
[0083] Figure 13 shows an exemplary viewport enhancement representation. A 360° video may have two identified most viewed viewports (e.g., viewport #1 and viewport #2). One or more DASH representations can correspond to the two identified most viewed viewports. At high encoding bitrates (e.g., 20 mbps), both viewports may be enhanced in high quality. When both viewports are enhanced in high quality, both enhanced viewports may be provided in a single representation (e.g., to facilitate fast viewport changes and / or to save storage costs). At medium bitrates, three representations may be provided with quality enhancements in both viewports or in individual viewports. Clients can request different representations based on fast viewport changes and / or preference for high quality in either viewport. At lower bitrates, enhancements for each viewport may be provided individually in two representations so that clients can request appropriate viewport quality based on their viewing direction. Representation selection can allow for different trade-offs between low-latency viewport switching, lower storage costs, and / or adequate display quality.
[0084] During an adaptive 360° video streaming session, a client can request a specific representation based on available bandwidth, the viewport the user is focusing on, and / or changes in viewing direction (e.g., how quickly and / or frequently the viewport changes). The client WTRU can locally analyze one or more user habits using one or more gyroscope, accelerometer, and / or magnetometer tracking parameters to determine which representation to request. For example, if the client WTRU detects that the user is not and / or does not frequently change their viewing direction, it may request a representation with a single enhanced viewport. To ensure low-latency rendering and / or sufficient viewport quality, the client WTRU may request a representation with multiple enhanced viewports if it detects that the user continues to and / or tends to change their viewing direction.
[0085] Layer-based 360 video streaming can be a viewport-adaptive technique for 360° video streaming. Layer-based 360 video streaming can separate the viewport area from the entire frame. Layer-based 360 video streaming can enable more flexible and / or efficient compositing of various virtual and / or real objects onto a sphere.
[0086] The entire frame may be encoded as a full-frame video layer at a lower quality, lower frame rate, and / or lower resolution. One or more viewports may be encoded as viewport layers into one or more quality representations. Viewport layers may be encoded independently of the full-frame video layer. Viewport layers can be coded more efficiently using scalable coding, for example, using the scalable extension of HEVC (SHVC), where the full-frame video layer is used as a reference layer to encode one or more viewport layer representations using inter-layer prediction. The user can always request the full-frame video layer (for example, as a fallback layer). When the user WTRU has sufficient additional resources (e.g., bandwidth and / or computing resources), one or more high-quality enhancement viewports may be requested to overlap the full-frame video layer.
[0087] Pixels outside the viewport may be subsampled. The frame rate of areas outside the viewport may not be directly reduced. Layer-based 360 video streaming can separate the viewport from the entire 360° video frame. 360 video streaming may allow the viewport to be encoded at a higher bitrate, higher resolution, and / or higher frame rate, while the full-frame video layer 360° video may be encoded at a lower bitrate, lower resolution, and / or lower frame rate.
[0088] The full-frame video layer can be encoded at a lower resolution and upsampled for the overlay on the client side. Upsampling the full-frame video layer on the client side can reduce storage costs and / or provide acceptable quality at lower bitrates. With respect to the viewport layer, multiple quality representations of the viewport can be generated for fine-grained quality adaptation.
[0089] Figure 14 shows an exemplary layer-based 360° video overlay. For example, a high-quality viewport may be overlaid on a 360° full-frame video layer. The full-frame video layer stream can contain the entire 360° video at a lower quality. The enhancement layer stream can contain an enhanced viewport at a higher quality.
[0090] One or more directional low-pass filters may be applied across overlay boundaries (e.g., horizontal and / or vertical boundaries) to smooth abrupt quality changes. For example, a 1D or 2D low-pass filter may be applied on a vertical boundary to smooth one or more horizontally adjacent pixels along the vertical boundary, and / or a similar low-pass filter may be applied on a horizontal boundary to smooth one or more vertically adjacent pixels along the horizontal boundary.
[0091] A full-frame video layer can contain the entire 360° video in different projection formats such as equirectangular projection, cubemap, offset cubemap, and pyramidal projection. The projection format may have multiple full-frame video layers (Representations) supporting different quality levels such as resolution, bitrate, and / or frame rate.
[0092] A viewport layer can contain multiple viewports. A viewport can have several quality representations with different resolutions, bitrates, and / or frame rates for adaptive streaming. Multiple representations can be provided to support fine-grained quality adaptation without incurring high storage and / or transmission costs (for example, because the size of each viewport is relatively small compared to the size of the entire 360° video).
[0093] Figure 15 shows an exemplary layer-based 360° video representation. For example, two or more Representations may be available in the Full Frame Video Layer for the entire 360° video with different resolutions (2048×1024@30fps and 4096×2048@30fps). Two or more target viewports (e.g., Viewport #1 and Viewport #2) may be available in the Viewport Layer. Viewports #1 and #2 may have Representations with different resolutions and / or different bitrates. A Viewport Layer Representation can identify a specific Full Frame Video Layer Representation as its dependent Representation using @dependencyId. The user can request the Target Viewport Representation and / or its dependent Full Frame Video Layer Full Frame Representation to composite the final 360° video for rendering.
[0094] The MPEG-DASH SRD element can be used to support layer-based 360° video streaming. Depending on the application, the MPD author can use one or more SRD values to describe the spatial relationship between the full-frame video layer (360° full video) and the viewport layer video. The SRD element can specify the spatial relationship of spatial objects. The SRD element may not specify how one or more viewport videos are overlaid on the full-frame video layer. Viewport overlays may be specified to improve the streaming quality of layer-based 360° video streaming.
[0095] A viewpoint overlay can be used with a Role descriptor applied to an AdaptationSet element. A Role element with @schemeIdURI equal to "urn:mpeg:dash:viewport:overlay:2016" can signal which Representation is associated with the viewport layer image and / or which Representation is associated with the full-frame image layer. The @value of a Role element can contain one or more overlay indicators. For example, one or more overlay indicators could contain "f" and / or "v", where "f" indicates that the associated image Representation is the full-frame image layer and "v" indicates that the associated image Representation is the viewport image that will be overlaid on the full-frame image layer. One or more viewport images can be overlaid on the full-frame image layer. @dependencyId can specify a particular full-frame image layer Representation for the associated viewport image overlay composition. The original full-frame video layer resolution may be indicated by one or more SRD parameters (e.g., total_width and total_height of the associated AdaptationSet). The original resolution of the viewport may be indicated by one or more SRD parameters, object_width and object_height of the associated AdaptationSet. When the Representation resolution indicated by @width and @height is lower than the corresponding resolution specified by the SRD, the reconstructed video may be upsampled to align the full-frame video layer and enhancement layer video for overlay. Table 12 shows exemplary MPD SRD annotations for a layer-based 360° video representation example shown in Figure 15.
[0096] Table 12 shows that two or more full-frame representations identified by @id can be included in the same AdaptationSet. The first Representation (@id=2) may contain a full resolution of 4096×2048, and the second Representation (@id=1) may contain a half resolution. An AdaptationSet can contain an SRD that describes that the AdaptationSet elements span the entire reference space, since the object_width and object_height parameters are equal to the total_width and total_height parameters. An AdaptationSet can contain two or more Representations that have different resolutions but represent the same spatial portion of the source (for example, the first Representation may have a resolution of 4096×2048, and the second Representation may have a resolution of 2048×1024). An AdaptationSet Role element whose @schemeIdUri is equal to "urn:mpeg:dash:viewport:overlay:2016" and whose @value is equal to "b1" can indicate that one or more AdaptationSet elements are full-frame video layers. When the full-frame video layer Representation resolution specified by the values of the attributes @width and @height is not equal to the values of the SRD parameters @total_width and @total_height, the full-frame video layer video may be upscaled to match the full-frame video layer Representation resolution.
[0097] Regarding enhancement layers, viewport #1Representation may be included in the first AdaptationSet, and viewport #2Representation may be included in the second AdaptationSet. Each viewport may have a different resolution and / or bandwidth. An AdaptationSet may contain an SRD that describes that the image in the AdaptationSet element represents only a portion of the entire 360° image (for example, because its object_width and object_height parameters are smaller than its total_width and total_height, respectively). An AdaptationSet's Role element with @schemeIdUri equal to "urn:mpeg:dash:viewport:overlay:2016" and @value equal to "f" can indicate that the AdaptationSet element is a viewport image. When the viewport Representation resolution specified by the values of the attributes @width and @height is not equal to the values of the SRD parameters @object_width and @object_height, the viewport image may be upscaled to match the original full-frame resolution.
[0098] One or more (e.g., all) AdaptationSets can use the same first parameter source_id to indicate that the images in the AdaptationSet are spatially related to each other within a 4096x2048 reference space. Non-SRD-aware clients can only view full-frame image layers by using SupplementalProperty instead of EssentialProperty. SRD-aware clients can form a 360° image by overlaying viewport #1 (@id=3 / 4 / 5) and / or viewport #2 (@id=6 / 7 / 8) images onto any of the full-frame image layers (@id=l / 2) (for example, by upsampling if necessary).
[0099] [Table 12-1]
[0100] [Table 12-2]
[0101] Figure 16 shows an exemplary layer-based 360 video overlay. For example, an exemplary layer-based 360 video overlay may be associated with the Representations listed in Table 12. A client may request a half-resolution full-frame video layer Representation (@id=1) and an enhancement layer viewport #1Representation (@id=5). The half-resolution full-frame video layer can be upsampled to full resolution 4096×2048 because its Representation resolutions @width=2048 and @height=1024 are smaller than the associated SRD parameters @total_width=4096 and @total_height=2048. A high-quality viewport can be overlaid on the 360 video at positions signaled by object_x (e.g., 100) and object_y (e.g., 100) so as to be signaled by the associated SRD parameters.
[0102] The video transformations shown in Figure 16 can perform upscaling, projection transformation, and / or layout transformation (for example, when the representation resolution, projection, and / or layout do not match between the full-frame video layer and the viewport video).
[0103] Figure 17 shows an exemplary layer-based 360 video overlay flowchart.
[0104] Full-frame video layers and enhancement layer representations may have different frame rates. Since the attribute @frameRate, along with @width, @height, and @bandwidth, is specified as a common attribute and element for AdaptationSet, Representation, and Sub-Representation elements, @width, @height, and @frameRate can be signaled to AdaptationSet, Representation, and / or Sub-Representation. When @frameRate is signaled at the AdaptationSet level, one or more Representations assigned to the AdaptationSet may share the same @frameRate value. Table 13 shows an MPD example for Representations with different frame rates at the AdaptationSet level.
[0105] [Table 13-1]
[0106] [Table 13-2]
[0107] [Table 13-3]
[0108] Multiple target viewport images may be requested and / or overlaid on a 360° full image sphere. Based on the user orientation, the user may request multiple high-quality viewports simultaneously (for example, to facilitate fast viewport changes and / or to reduce overall system latency). One or more (e.g., all) viewport layers may share the same full-frame image layer 360° image as the reference space. The client can overlay the multiple viewports it receives onto the same reference space for rendering. When projection and layout differ between the full-frame image layer and the enhancement layer viewport, the full-frame image layer image may be used as the anchor format, and / or the enhancement layer viewport may be converted to the anchor format (for example, so that the composite position can be aligned between the full-frame image layer and the viewport image).
[0109] Figure 18 shows an exemplary multiple viewport overlay. For example, a layer-based overlay with two enhanced viewports may be signaled as shown in Table 13. A client may request a full-frame video layer 360° video as the reference space (e.g., Representation@id=1) and / or two high-quality enhanced layer viewports (e.g., Representation@id=5 and @id=8). A composite video with two enhanced viewports may be generated based on the decision that both viewports may be viewed for a short period (e.g., during the same segment duration). Additional projection and / or layout transformations may be performed if the video type, projection, and / or layout differ between the full-frame video layer video and the enhanced layer viewports. Layer-based 360° video overlays may be able to adapt to viewport changes more efficiently. The segment length of the full-frame video layer may be increased to improve encoding efficiency (e.g., since the full-frame video layer video is always delivered). The segment length of the viewport video layer can be kept shorter to accommodate rapid switching between viewports.
[0110] The layer-based 360 video overlays described herein can use an SRD to describe the spatial relationships between the viewport video and / or the overlay, and to describe the overlay procedure for forming a composite video. One or more spatial parameters specified in the SRD, such as source_id, object_x, object_y, object_width, object_height, total_width, and / or total_height, can be merged into the overlay @value to simplify the MPD structure. The @value of the Role element can contain a comma-separated list of viewport indicators, source_id, object_x, object_y, object_width, object_height, total_width, and / or total_height. Table 14 shows an example of an MPD with a merged overlay.
[0111] [Table 14-1]
[0112] [Table 14-2]
[0113] A layer-based 360 video overlay can be provided by one or more descriptors such as SupplementalProperty and / or EssentialProperty where @schemeIdUri is equal to "urn:mpeg:dash:viewport:overlay:2016".
[0114] Different video representations can be in different projection formats (for example, one representation might be equirectangular while another is a cubemap, or one representation might be spherical while another is linear). Video properties can be specified from the OMAF file format. One or more common attributes, @videoType, specified for AdaptationSet, Representation, and / or Sub-Representation can indicate spherical, light-field, and / or linear images. For spherical images, different projection formats and / or projection + layout combinations, such as equirectangular, cubemap (combined with different layouts shown in Figure 10), and / or pyramidal map, can be signaled as common attributes @projection and / or @layout for AdaptationSet, Representation, and / or Sub-Representation elements. Elements and attributes can be provided using SupplementalProperty and / or EssentialProperty elements.
[0115] Table 15 shows examples of the semantics of image types and / or projection attributes. Projection attributes can help users request appropriate images based on their capabilities. For example, a client who does not support spherical and / or light field images may request only linear images. In the example, a client who supports spherical images may select an appropriate format from the set of available formats (e.g., equirectangular projection instead of cubemap).
[0116] [Table 15]
[0117] [Table 16]
[0118] [Table 17]
[0119] [Table 18]
[0120] The layout format of Table 18 can be illustrated in the following figure.
[0121] Figure 19 shows an exemplary equirectangular projection representation with a half-height pole on the side.
[0122] Figure 20 shows an exemplary equirectangular projection representation with full-height poles on the sides.
[0123] Figure 21 shows an exemplary single-column layout for a cube representation format (for example, one with region_id).
[0124] Figure 22 shows an exemplary 2x3 layout for a cube representation format (for example, one with a region_id).
[0125] Figure 23 shows an exemplary 180° layout for a cubemap.
[0126] A 180° layout for cubemap projection can reduce the resolution of the back half of the cube by 25% (for example, by halving the width and height of the back half area), and / or a layout can be used in which, as a result, areas a, b, c, d, and e shown in Figure 23 belong to the front half and areas f, g, h, i, and j belong to the back half.
[0127] Table 19 shows an MPD example that supports both equirectangular projection and cubemap projection using such common attributes and elements.
[0128] [Table 19-1]
[0129] [Table 19-2]
[0130] Figure 24A is a diagram of an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), and single-carrier FDMA (SC-FDMA).
[0131] As shown in Figure 24A, the communication system 100 may include radio transceiver units (WTRUs) 102a, 102b, 102c, and / or 102d (sometimes referred to generally or collectively as WTRUs 102), radio access networks (RANs) 103 / 104 / 105, core networks 106 / 107 / 109, public switched telephone networks (PSTNs) 108, the internet 110, and other networks 112, but it will be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d may be configured to transmit and / or receive radio signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, and consumer electronics.
[0132] The communication system 100 may also include base stations 114a and 114b. Each of the base stations 114a and 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as the core networks 106 / 107 / 109, the Internet 110, and / or network 112. For example, base stations 114a and 114b could be transceiver base stations (BTS), node B, enode B, home node B, home enode B, site controllers, access points (APs), and wireless routers, etc. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b can include any number of interconnected base stations and / or network elements.
[0133] Base station 114a may be part of RAN 103 / 104 / 105, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals within a specific geographic area, which may be called a cell (not shown). A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, for example, one per sector of the cell. In another embodiment, base station 114a may utilize multiple-input multiple-output (MIMO) technology, and therefore multiple transceivers may be available per sector of the cell.
[0134] Base stations 114a and 114b can communicate with one or more WTRUs 102a, 102b, 102c, and 102d via air interfaces 115 / 116 / 117, where air interfaces 115 / 116 / 117 can be any suitable radio communication link (e.g., radio frequency (RF), microwave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interfaces 115 / 116 / 117 can be established using any suitable radio access technology (RAT).
[0135] More specifically, as described above, the communication system 100 can be a multiple access system and can utilize one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a and WTRU 102a, 102b, 102c in RAN 103 / 104 / 105 can implement radio technologies such as Universal Mobile Communications System (UMTS) Terrestrial Radio Access (UTRA) that can establish air interfaces 115 / 116 / 117 using broadband CDMA (WCDMA®). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink Packet Access (HSDPA) and / or High Speed Uplink Packet Access (HSUPA).
[0136] In another embodiment, base stations 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish air interfaces 115 / 116 / 117 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A).
[0137] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as IEEE 802.16 (e.g., Global Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM®), Extended Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0138] In Figure 24A, base station 114b can be, for example, a wireless router, home node B, home e-node B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as workplaces, homes, vehicles, and campuses. In one embodiment, base station 114b and WTRU 102c, 102d can establish a wireless local area network (WLAN) by implementing wireless technology such as IEEE 802.11. In another embodiment, base station 114b and WTRU 102c, 102d can establish a wireless personal area network (WPAN) by implementing wireless technology such as IEEE 802.15. In yet another embodiment, base station 114b and WTRU 102c, 102d can establish a picocell or femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, etc.). As shown in Figure 24A, base station 114b can have a direct connection to the internet 110. Therefore, base station 114b does not need to be required to access the internet 110 via core network 106 / 107 / 109.
[0139] RAN103 / 104 / 105 can communicate with core networks 106 / 107 / 109, which can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. For example, core networks 106 / 107 / 109 can provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 24A, it will be understood that RAN103 / 104 / 105 and / or core networks 106 / 107 / 109 can communicate directly or indirectly with other RANs that utilize the same RAT or a different RAT as RAN103 / 104 / 105. For example, in addition to connecting to RANs 103 / 104 / 105 which can utilize E-UTRA radio technology, core networks 106 / 107 / 109 can also communicate with other RANs (not shown) that utilize GSM radio technology.
[0140] Core networks 106 / 107 / 109 can also serve as gateways for WTRUs 102a, 102b, 102c, and 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing basic telephone services (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and Internet Protocol (IP) in the TCP / IP Internet Protocol suite. Networks 112 may include wired or wireless networks owned and / or operated by other service providers. For example, network 112 may include another core network connected to one or more RANs that have the same or different RATs as RAN 103 / 104 / 105.
[0141] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode capability. For example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different radio networks via different radio links. For example, WTRU 102c, shown in Figure 24A, may be configured to communicate with base station 114a, which can utilize cellular-based radio technology, and base station 114b, which can utilize IEEE 802 radio technology.
[0142] Figure 24B is a system diagram of an exemplary WTRU102. As shown in Figure 24B, the WTRU102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and other peripherals 138. It will be understood that the WTRU102 may include any partial combination of the above elements while maintaining consistency with the embodiment. Furthermore, the embodiments intend that base stations 114a and 114b, and / or nodes that base stations 114a and 114b may represent, such as, but not limited to, transceiver stations (BTS), node B, site controller, access point (AP), home node B, evolved home node B (eNodeB), home evolved node B (HeNB), home evolved node B gateway, and proxy node, may include some or all of the elements shown in Figure 24B and described herein.
[0143] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to a transceiver 120, and the transceiver 120 can be coupled to a transmit / receive element 122. Although Figure 24B shows the processor 118 and transceiver 120 as separate components, it will be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0144] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interfaces 115 / 116 / 117. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of radio signals.
[0145] In addition, although the transmit / receive element 122 is shown as a single element in Figure 24B, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Therefore, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals via the air interfaces 115 / 116 / 117.
[0146] The transceiver 120 may be configured to modulate the signal transmitted by the transmit / receive element 122 and demodulate the signal received by the transmit / receive element 122. As described above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as UTRA and IEEE 802.11.
[0147] The processor 118 of the WTRU102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (for example, a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and can receive user input data from them. The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can access information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and can store data in them. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, and a secure digital (SD) memory card, etc. In other embodiments, the processor 118 can access information from memory, such as on a server or home computer (not shown), rather than from memory physically located on the WTRU 102, and can store data therein.
[0148] The processor 118 can receive power from the power supply 134 and may be configured to distribute and / or control power to other components within the WTRU 102. The power supply 134 can be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.
[0149] The processor 118 may be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interfaces 115 / 116 / 117 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information by any suitable location determination method while maintaining consistency with the embodiments.
[0150] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photography or video), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, and an internet browser.
[0151] Figure 24C is a system diagram of RAN103 and core network 106 according to an embodiment. As described above, RAN103 can communicate with WTRU102a, 102b, and 102c via air interface 115 using UTRA radio technology. RAN103 can also communicate with core network 106. As shown in Figure 24C, RAN103 may include nodes B140a, 140b, and 140c, each of which may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via air interface 115. Each of nodes B140a, 140b, and 140c may be associated with a specific cell (not shown) within RAN103. RAN103 may also include RNC142a and 142b. It will be understood that RAN103 can include any number of nodes B and RNC while maintaining consistency with the embodiment.
[0152] As shown in Figure 24C, nodes B140a and B140b can communicate with RNC142a. In addition, node B140c can communicate with RNC142b. Nodes B140a, B140b, and B140c can communicate with their respective RNC142a and B142b via the Iub interface. RNC142a and B142b can communicate with each other via the Iur interface. Each of RNC142a and B142b may be configured to control their respective connected nodes B140a, B140b, and B140c. In addition, each of RNC142a and B142b may be configured to implement or support other functionalities such as outer loop power control, load control, admission control, packet scheduling, handover control, macro diversity, security functions, and data encryption.
[0153] The core network 106 shown in Figure 24C may include a media gateway (MGW) 144, a mobile switching center (MSC) 146, a serving GPRS support node (SGSN) 148, and / or a gateway GPRS support node (GGSN) 150. Although each of the above elements is shown as part of the core network 106, it will be understood that any of these elements may be owned and / or operated by an entity different from the core network operator.
[0154] The RNC142a in RAN103 may be connected to the MSC146 in the core network 106 via the IuCS interface. The MSC146 may be connected to the MGW144. The MSC146 and MGW144 can provide WTRU102a, 102b, and 102c with access to a circuit-switched network such as PSTN108, thereby facilitating communication between WTRU102a, 102b, and 102c and conventional land-line communication devices.
[0155] RNC142a in RAN103 can also be connected to SGSN148 in core network 106 via the IuPS interface. SGSN148 can be connected to GGSN150. SGSN148 and GGSN150 can provide WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0156] As described above, the core network 106 may also be connected to network 112, which may include other wired or wireless networks owned and / or operated by other service providers.
[0157] Figure 24D is a system diagram of RAN104 and core network 107 according to an embodiment. As described above, RAN104 can communicate with WTRU102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN104 can also communicate with core network 107.
[0158] RAN104 may include e-nodes B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of e-nodes B while maintaining consistency with the embodiment. Each of e-nodes B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, e-nodes B160a, 160b, and 160c can implement MIMO technology. Thus, e-node B160a can, for example, use multiple antennas to transmit radio signals to and receive radio signals from WTRU102a.
[0159] Each of the e-nodes B160a, 160b, and 160c may be associated with a cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, and user scheduling on the uplink and / or downlink. As shown in Figure 24D, the e-nodes B160a, 160b, and 160c can communicate with each other via the X2 interface.
[0160] The core network 107 shown in Figure 24D may include a Mobility Management Gateway (MME) 162, a Serving Gateway 164, and a Packet Data Network (PDN) Gateway 166. Although each of the above elements is shown as part of the core network 107, it will be understood that any of these elements may be owned and / or operated by an entity different from the core network operator.
[0161] The MME162 may be connected to each of the e-nodes B160a, 160b, and 160c within RAN104 via the S1 interface and can act as a control node. For example, the MME162 can be responsible for user authentication of WTRU102a, 102b, and 102c, bearer activation / deactivation, and selection of the serving gateway during the initial attachment of WTRU102a, 102b, and 102c. The MME162 can also provide control plane functionality for switching between RAN104 and other RANs (not shown) utilizing other radio technologies such as GSM or WCDMA.
[0162] The serving gateway 164 may be connected to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface. The serving gateway 164 can generally route and forward user data packets to and from WTRU102a, 102b, and 102c. The serving gateway 164 can also perform other functions, such as anchoring the user plane during handover between e-nodes B, triggering paging when downlink data is available to WTRU102a, 102b, and 102c, and managing and remembering the context of WTRU102a, 102b, and 102c.
[0163] The serving gateway 164 may be connected to the PDN gateway 166, which provides WTRUs 102a, 102b, and 102c with access to a packet-switched network such as the Internet 110, thereby facilitating communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.
[0164] The core network 107 can facilitate communication with other networks. For example, the core network 107 can provide WTRU 102a, 102b, and 102c with access to a circuit-switched network such as PSTN 108, thereby facilitating communication between WTRU 102a, 102b, and 102c and conventional land-line communication devices. For example, the core network 107 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the core network 107 and PSTN 108, or can communicate with such an IP gateway. In addition, the core network 107 can provide WTRU 102a, 102b, and 102c with access to network 112, which may include other wired or wireless networks owned and / or operated by other service providers.
[0165] Figure 24E is a system diagram of RAN105 and core network 109 according to an embodiment. RAN105 can be an access service network (ASN) that communicates with WTRU102a, 102b, and 102c via air interface 117 using IEEE 802.16 wireless technology. As will be further described below, communication links between different functional entities of WTRU102a, 102b, 102c, RAN105, and core network 109 can be defined as reference points.
[0166] As shown in Figure 24E, RAN105 may include base stations 180a, 180b, 180c and an ASN gateway 182, but it will be understood that RAN105 may include any number of base stations and ASN gateways while maintaining consistency with the embodiment. Each of the base stations 180a, 180b, and 180c may be associated with a cell (not shown) within RAN105, and each may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 117. In one embodiment, the base stations 180a, 180b, and 180c may implement MIMO technology. Thus, base station 180a may, for example, use multiple antennas to transmit radio signals to and receive radio signals from WTRU102a. Base stations 180a, 180b, and 180c can also provide mobility management functions such as handoff triggering, tunnel establishment, radio resource management, traffic classification, and quality of service (QoS) policy enforcement. The ASN gateway 182 can act as a traffic aggregation point and is responsible for paging, subscriber profile caching, and routing to the core network 109.
[0167] The air interface 117 between WTRU102a, 102b, 102c and RAN105 may be defined as an R1 reference point implementing the IEEE 802.16 specification. In addition, each of WTRU102a, 102b, and 102c may establish a logical interface (not shown) with the core network 109. The logical interface between WTRU102a, 102b, 102c and the core network 109 may be defined as an R2 reference point that can be used for authentication, authorization, IP host configuration management, and / or mobility management.
[0168] The communication links between base stations 180a, 180b, and 180c may be defined as R8 reference points, including protocols to facilitate WTRU handover and data transfer between base stations. The communication links between base stations 180a, 180b, and 180c and the ASN gateway 182 may be defined as R6 reference points. The R6 reference points may include protocols to facilitate mobility management based on mobility events associated with each of WTRU 102a, 102b, and 102c.
[0169] As shown in Figure 24E, RAN 105 may be connected to core network 109. The communication link between RAN 105 and core network 109 may be defined as an R3 reference point, including protocols to facilitate data transfer and mobility management capabilities, for example. Core network 109 may include a Mobile IP Home Agent (MIP-HA) 184, an Authentication, Authorization, and Billing (AAA) server 186, and a gateway 188. Although each of the above elements is shown as part of core network 109, it will be understood that any of these elements may be owned and / or operated by an entity different from the core network operator.
[0170] The MIP-HA can be responsible for IP address management and can enable WTRU102a, 102b, and 102c to roam between different ASNs and / or different core networks. The MIP-HA184 can provide WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices. The AAA server 186 can be responsible for user authentication and support for user services. The gateway 188 can facilitate inter-network connectivity with other networks. For example, the gateway 188 can provide WTRU102a, 102b, and 102c with access to circuit-switched networks such as the PSTN 108, facilitating communication between WTRU102a, 102b, and 102c and conventional landline communication devices. In addition, gateway 188 provides access to network 112 to WTRU 102a, 102b, and 102c, and network 112 may include other wired or wireless networks owned and / or operated by other service providers.
[0171] Each of the computing systems described herein may have one or more computer processors or hardware having memory comprising executable instructions for performing the functions described herein, including determining the parameters described herein and sending and receiving messages between entities (e.g., WTRUs and networks or clients and servers) in order to perform the functions described herein. The processes described above may be implemented in computer programs, software, and / or firmware embedded in computer-readable media for execution by the computer and / or processor.
[0172] Although not shown in Figure 24E, it will be understood that RAN105 may be connected to other ASNs, and core network 109 may be connected to other core networks. The communication link between RAN105 and other ASNs may be defined as an R4 reference point that can include protocols for coordinating the mobility of WTRU102a, 102b, and 102c between RAN105 and other ASNs. The communication link between core network 109 and other core networks may be defined as an R5 reference point that can include protocols for facilitating inter-network connectivity between the home core network and the visited core network.
[0173] While features and elements are described above in specific combinations, it will be understood by those skilled in the art that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). Software and associated processors can be used to implement radio frequency transceivers for use in WTRUs, WTRU terminals, base stations, RNCs, or any host computer. [Industrial applicability]
[0174] This invention can be used in virtual reality (VR) devices. [Explanation of Symbols]
[0175] 100 Communication Systems 102, 102a~102d WTRU 103, 104, 105 RAN 106, 107, 108 Core Network 108 PSTN 110 Internet 112 Other networks 114a, 114b base station 118 processors 120 Wireless Transceivers (Transceivers) 122 Antenna
Claims
1. Receiving a media presentation description (MPD) file associated with a 360-degree video, wherein the MPD file includes a first representation and a second representation, the first representation including a first set of spatial regions of the 360-degree video associated with a first quality level, the second representation including a second set of spatial regions of the 360-degree video associated with a second quality level, and the second quality level being higher than the first quality level. Monitor the user's orientation using tracking information from at least one sensor, Based on the monitored user orientation position, determine the frequency measurement associated with the change in the user's viewing direction, Selecting a first subset of spatial regions from the second set of spatial regions based on the frequency measurement associated with the change in the user's viewing direction, Requesting the aforementioned first subset of the spatial domain from the media server, A device equipped with a processor configured to perform [a certain action].
2. The aforementioned processor, The apparatus according to claim 1, configured to select the first subset of spatial regions representing a single extended viewport of the user based on the determination that the frequency measurement is below a threshold.
3. The aforementioned processor, The apparatus according to claim 1, configured to select a first subset of spatial regions representing a user's multiple extended viewports based on the determination that the frequency measurement value exceeds a threshold.
4. The apparatus of claim 1, wherein the first subset of the spatial region is selected to cover a first portion of the expanded spatial region of the 360-degree image based on the determination that the frequency measurement is above a threshold, and is selected to cover a second portion of the expanded spatial region of the 360-degree image based on the determination that the frequency measurement is below a threshold, the first portion covers an expanded spatial region larger than the second portion, the first portion of the expanded spatial region of the 360-degree image represents a plurality of expanded viewports of the user, and the second portion of the expanded spatial region of the 360-degree image represents a single expanded viewport of the user.
5. The apparatus according to claim 1, wherein the frequency measurement value characterizes the user's viewing habits.
6. The apparatus according to claim 1, wherein the processor is configured to request a second subset of the spatial domain, the second subset including a spatial domain independent of the first subset of the spatial domain.
7. The apparatus according to claim 1, wherein the at least one sensor includes one or more of a gyroscope, an accelerometer, and a magnetometer.
8. Receiving a media presentation description (MPD) file associated with a 360-degree video, wherein the MPD file includes a first representation and a second representation, the first representation including a first set of spatial regions of the 360-degree video associated with a first quality level, the second representation including a second set of spatial regions of the 360-degree video associated with a second quality level, and the second quality level being higher than the first quality level. Monitor the user's orientation using tracking information from at least one sensor, Based on the monitored user orientation position, determine the frequency measurement associated with the change in the user's viewing direction, Selecting a first subset of spatial regions from a second set of spatial regions based on the frequency measurement associated with the change in the user's viewing direction, Requesting the aforementioned first subset of the spatial domain from the media server, A method for providing this.
9. Selecting the first subset of the spatial domain is The method of claim 8, comprising selecting the first subset of spatial regions representing a single expanded viewport of the user based on the determination that the frequency measurement is below a threshold.
10. Selecting the first subset of the spatial domain is The method of claim 8, comprising selecting the first subset of spatial regions representing the user's multiple extended viewports based on the determination that the frequency measurement value exceeds a threshold.
11. The method of claim 8, wherein the first subset of the spatial region is selected to cover a first portion of the expanded spatial region of the 360-degree image based on the determination that the frequency measurement is above a threshold, and is selected to cover a second portion of the expanded spatial region of the 360-degree image based on the determination that the frequency measurement is below a threshold, the first portion covers a larger expanded spatial region than the second portion, the first portion of the expanded spatial region of the 360-degree image represents a plurality of expanded viewports of the user, and the second portion of the expanded spatial region of the 360-degree image represents a single expanded viewport of the user.
12. The method of claim 8, wherein the frequency measurement values characterize the user's viewing habits.
13. The method of claim 8, comprising requesting a second subset of the spatial domain, wherein the second subset includes a spatial domain independent of the first subset of the spatial domain.
14. The method of claim 8, wherein the at least one sensor includes one or more of a gyroscope, an accelerometer, and a magnetometer.
Citation Information
Patent Citations
Video transmitting apparatus and video receiving apparatus
JP2005142654A
Moving image encoding device and moving image decoding device
JP2010212811A
Method for navigation in panoramic scene
JP2012070378A
Extensions of motion-constrained tile sets SEI message for interactivity
US20150016504A1
Method of displaying a region of interest in a video stream
US20150365687A1