Viewpoint Metadata for Omnidirectional Video
The system addresses the challenge of delivering high-quality omnidirectional videos by signaling viewpoint metadata in ISO BMFF and DASH, facilitating smooth transitions and efficient delivery of immersive content.
Patent Information
- Application Number
- JP2023206939
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-23
- Filing Date
- 2023-12-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-04-04
AI Technical Summary
Omnidirectional videos present challenges in providing a comfortable and immersive user experience due to high video quality and low latency requirements, while large video sizes hinder efficient delivery.
A system and method for signaling position information and viewpoint metadata in omnidirectional video presentations using ISO Base Media File Format and Dynamic Adaptive Streaming over HTTP (DASH) to manage multiple viewpoints, enabling smooth transitions and efficient delivery.
Enables a seamless and immersive user experience by allowing navigation between multiple viewpoints with efficient delivery of high-quality omnidirectional video content.
Smart Images

Figure 0007708839000009 
Figure 0007708839000010 
Figure 0007708839000011
Abstract
Description
Technical Field
[0001] Relates to viewpoint metadata for omnidirectional video.
Background Art
[0002] Cross - reference to related applications This application is a non - provisional application of U.S. Provisional Patent Application No. 62 / 653,363 (filed on April 5, 2018) and U.S. Provisional Patent Application No. 62 / 675,524 (filed on May 23, 2018), both entitled "Viewpoint Metadata for Omnidirectional Video" and incorporated herein by reference in their entirety, and claims the benefit under 35 U.S.C. § 119(e).
[0003] Omnidirectional video or 360° video is a rapidly growing new format emerging in the media industry. It becomes usable due to the growing availability of VR devices and can provide viewers with a greater sense of presence.
Prior Art Documents
Non - Patent Documents
[0004]
Non - Patent Document 1
Non - Patent Document 2
Non - Patent Document 3
[0005] Compared with conventional rectilinear videos (2D or 3D), 360° videos present a new set of difficult technical problems for video processing and delivery. Enabling a comfortable and immersive user experience requires high video quality and very low latency, but large video sizes can be an obstacle to delivering high-quality 360° videos. [Means for Solving the Problems]
[0006] ISO Base Media File Format Within the ISO / IEC 14496 MPEG-4 standard, there are several parts that define file formats for the storage of time-based media. These are all based on and derived from the ISO Base Media File Format (ISO BMFF) described in ISO / IEC 14496-12, "Coding of Audio-Visual Objects, Part 12: ISO Base Media File Format", 2015. ISO BMFF is a structural media-independent definition. ISO BMFF mainly includes the structure and media data information for timed presentations of media data such as audio and video. Support also exists for un-timed data, such as metadata at different levels within the file structure. Next, the logical structure of the file is that of a video containing a set of tracks parallel in time. The time structure of the file is such that the tracks contain sequences of samples in time, and these sequences are mapped to the timeline of the entire video. ISO BMFF is based on the concept of a box-structured file. A box-structured file consists of a series of boxes (also called atoms) having a certain size and type. The type is a 32-bit value and is usually chosen to be four printable characters known as a four-character code (4CC). Un-timed data can be included in a metadata box at the file level, or added to a video box or one of the streams of timed data called tracks within the video.
[0007] Dynamic Streaming over HTTP (DASH) MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH) is a delivery format that dynamically adapts to changing network conditions. MPEG-DASH is described in ISO / IEC 23009-1, "Dynamic adaptive streaming over HTTP (DASH), Part 1: Media Presentation Description and Segment Formats," in May 2014. Dynamic HTTP streaming requires various bitrate alternatives of multimedia content to be available on the server. In addition, multimedia content can be composed of several media components (e.g., audio, video, text, etc.) each having different characteristics. In MPEG-DASH, these characteristics are described by a Media Presentation Description (MPD).
[0008] Figure 2 shows the MPD hierarchical data model. The MPD describes a sequence of Periods, and a consistent set of encoded versions of the components of the media content does not change within a Period. Each Period has a start time and a duration and consists of one or more AdaptationSets.
[0009] An AdaptationSet represents a set of encoded versions of one or several media content components having common characteristics such as language, media type, picture aspect ratio, role, accessibility, and rating characteristics. For example, an AdaptationSet can include different bitrates of the video component of the same multimedia content. Another AdaptationSet can include different bitrates of the audio component of the same multimedia content (e.g., low-quality stereo and high-quality surround sound, etc.). Each AdaptationSet usually includes multiple Representations.
[0010] A representation describes an encoded version of one or several media components that can be delivered, differing from other representations by bitrate, resolution, number of channels, or other characteristics. Each representation consists of one or more segments. Attributes of the Representation element, such as @id, @bandwidth, @qualityRanking, and @dependencyId, are used to specify properties of the associated representation. A Representation can also include sub-representations that are part of the representation and that can be used to describe and extract partial information from the representation. Sub-representations can provide the ability to access lower-quality versions of the representation in which they are included.
[0011] A segment is the largest unit of data that can be retrieved in a single HTTP request. Each segment has a URL, which is an addressable location on the server and which can be downloaded using an HTTP GET, or an HTTP GET with byte ranges.
[0012] To use this data model, the DASH client parses the MPD XML document and selects a set of adaptation sets suitable for its environment based on the information provided in each AdaptationSet element. Within each adaptation set, the client typically selects one representation based on the value of the @bandwidth attribute, but also taking into account the client's decoding and rendering capabilities. The client downloads the initialization segment of the selected representation and then accesses the content by requesting the entire segment or a byte range of the segment. After the presentation has started, the client continues to consume the media content by continuously requesting media segments or parts of media segments and playing the content according to the media presentation timeline. The client can switch representations considering updated information from its environment. The client should play the content continuously over time. When the client is to consume the media included in a segment towards the end of the media notified in the representation, the media presentation ends, a new period starts, or the MPD is fetched again.
[0013] Descriptors in DASH MPEG-DASH uses descriptors to provide application-specific information about media content. Descriptor elements are all structured in the same way, i.e., they contain an @schemeIdUri attribute to provide a URI to identify the scheme, an optional @value attribute, and an optional @id attribute. The semantics of the elements are specific to the scheme used. The URI identifying the scheme can be a URN or a URL. The MPD does not provide any specific information on how these elements are to be used. It is up to the application using the DASH format to instantiate the descriptor elements with the appropriate scheme information. A DASH application using one of these elements first defines a scheme identifier in the form of a URI and then defines the value space for the element when that scheme identifier is used. If structured data is used, any extension elements or attributes can be defined in a different namespace. Descriptors can appear at several levels within the MPD. The presence of an element at the MPD level means that the element is a child of the MPD element. The presence of an element at the Adaptation Set level indicates that the element is a child element of the AdaptationSet element. The presence of an element at the Representation level indicates that the element is a child element of the Representation element.
[0014] Omnidirectional media format The Omnidirectional Media Format (OMAF) is a system standard developed by MPEG as Part 2 of MPEG-I, and is a set of standards for the encoding, representation, storage, and delivery of immersive media. OMAF enables omnidirectional media applications and defines a media format that mainly focuses on 360° video, images, audio, and related timed-metadata tracks. The final draft international standard (FDIS) of OMAF was released in early 2018 and is described in ISO / IEC JTC1 / SC29 / WG11 N17399 "FDIS 23090-2 Omnidirectional Media Format" in February 2018.
[0015] As part of Phase 1b of MPEG-I, an extension of OMAF that supports several new features including 3DoF plus motion parallax and support for multiple viewpoints was planned in 2019. The requirements for Phase 1b were released in February 2018 and are described in ISO / IEC JTC1 / SC29 / WG11 N17331 "MPEG-I Phase 1b Requirements" in February 2018. OMAF and the MPEG-I Phase 1b requirements describe the following concepts. · The Field-of-view (FoV) is the extent of the captured / recorded content or the observable world in a physical display device.
[0016] · The Viewpoint is the point from which the user views the scene, which usually corresponds to the camera position. A slight head movement does not necessarily imply a different viewpoint. · A Sample is all the data associated with a single point in time. · A track is a timed sequence of samples with time specifications in an ISO-based media file. In the case of media data, a track corresponds to a sequence of images or sampled audio. · A box is an object-oriented building block in an ISO-based media file defined by a unique type identifier and length.
[0017] In some embodiments, in an omnidirectional video presentation, a system and method for signaling position information for one or more viewpoints are provided. In some embodiments, the method includes receiving a manifest (e.g., an MPEG-DASH MPD) for the omnidirectional video presentation, where the video presentation has at least one omnidirectional video associated with a viewpoint; determining, based on the manifest, whether a timed-metadata track for the viewpoint position is provided for the viewpoint; and determining the viewpoint position based on the information in the timed-metadata track in response to a determination that the timed-metadata track is provided.
[0018] In some embodiments, the step of determining whether a timed-metadata track for the viewpoint position is provided includes determining whether a flag in the manifest indicates that the viewpoint position is dynamic.
[0019] In some embodiments, the manifest includes coordinates indicating a first viewpoint position.
[0020] In some embodiments, the timed-metadata track is identified in the manifest, and the method further includes the step of incorporating the timed-metadata track.
[0021] In some embodiments, the time-stamped metadata track includes a viewpoint position in Cartesian coordinates. In other embodiments, the time-stamped metadata track includes a viewpoint position in longitude and latitude coordinates.
[0022] In some embodiments, the method further includes displaying a user interface to the user, the user interface enabling the user to select an omnidirectional video based on the viewpoint position of the omnidirectional video. The omnidirectional video is displayed to the user in response to the user selection of the omnidirectional video.
[0023] In some embodiments, the omnidirectional video presentation includes at least a first omnidirectional video and a second omnidirectional video. In such embodiments, the step of displaying the user interface can include displaying the first omnidirectional video to the user and displaying a user interface element or other indication of the second omnidirectional video at a position of the first omnidirectional video that corresponds to the position of the viewpoint of the second omnidirectional video.
[0024] A method for signaling information regarding various viewpoints in a multi-viewpoint omnidirectional media presentation is described herein. In some embodiments, a container file (which can use the ISO base media file format) including a number of tracks is generated. The tracks are grouped using track group identifiers, with each track group identifier associated with a different viewpoint. In some embodiments, a manifest (such as an MPEG-DASH MPD) is generated, where the manifest includes a viewpoint identifier identifying the viewpoint associated with each stream. In some embodiments, the metadata included in the container file and / or the manifest provides information regarding one or more of the following: the position of each viewpoint, the valid range of each viewpoint, the intervals during which each viewpoint is available, the transition effect for transitions between viewpoints, and the projection formats recommended for different field-of-view ranges.
[0025] In some embodiments, the method comprises the step of generating a container file (e.g., an ISO-based media file format file). At least first and second 360-degree video data are received, where the first video data represents a view from a first perspective and the second 360-degree video data represents a view from a second perspective. The container file is generated for at least the first video data and the second video data. In the container file, the first video data is organized into a first set of tracks and the second video data is organized into a second set of tracks. Each of the tracks in the first set of tracks includes a first track group identifier associated with the first perspective, and each of the tracks in the second set of tracks includes a second track group identifier associated with the second perspective.
[0026] In some such embodiments, each of the tracks in the first set of tracks includes each instance of a view group type box that includes the first track group identifier, and each of the tracks in the second set of tracks includes each instance of a view group type box that includes the second track group identifier.
[0027] In some embodiments, the container file is organized into a hierarchical box structure and the container file includes at least a first perspective information box and a perspective list box that identifies the second perspective information box. The first perspective information box includes at least (i) the first track group identifier and (ii) an indication of the time period during which video from the first perspective is available. The second perspective information box includes at least (i) the second track group identifier and (ii) an indication of the time period during which video from the second perspective is available. The indication of the time interval can be a list of instances of each perspective available interval box.
[0028] In some embodiments, the container file is organized into a hierarchical box structure, and the container file includes at least a first viewpoint information box and a viewpoint list box that identifies the second viewpoint information box. The first viewpoint information box includes at least (i) a first track group identifier and (ii) an indication of the position of the first viewpoint. The second viewpoint information box includes at least (i) a second track group identifier and (ii) an indication of the position of the second viewpoint. The indication of the position can include, among other options, Cartesian coordinates or latitude and longitude coordinates.
[0029] In some embodiments, the container file is organized into a hierarchical box structure, and the container file also includes at least a first viewpoint information box and a viewpoint list box that identifies the second viewpoint information box. The first viewpoint information box includes at least (i) a first track group identifier and (ii) an indication of the effective range of the first viewpoint. The second viewpoint information box includes at least (i) a second track group identifier and (ii) an indication of the effective range of the second viewpoint.
[0030] In some embodiments, the container file is organized into a hierarchical box structure, and the container file includes a transition effect list box that identifies at least one transition effect box. Each transition effect box includes (i) an identifier of the source viewpoint, (ii) an identifier of the destination viewpoint, and (iii) an identifier of the transition type. The identifier of the transition type can, among other options, identify a basic transition, a viewpoint path transition, or an auxiliary information viewpoint transition. In the case of a viewpoint path transition, a path viewpoint transition box including a list of viewpoint identifiers can be provided. In the case of an auxiliary information viewpoint transition, an auxiliary information viewpoint transition box including a track identifier can be provided.
[0031] In some embodiments, the container file is organized into a hierarchical box structure that includes metaboxes, and each metabox identifies at least one recommended projection list box. Each recommended projection list box can include information that identifies (i) a projection type and (ii) a field of view range corresponding to the projection type. The information that identifies the field of view range can include (i) a minimum horizontal field of view angle, (ii) a maximum horizontal field of view angle, (iii) a minimum vertical field of view angle, and (iv) a maximum vertical field of view angle.
[0032] In some embodiments, a method for generating a manifest, such as an MPEG-DASH MPD, is provided. At least first 360-degree video data representing a view from a first viewpoint and second 360-degree video data representing a view from a second viewpoint are received. A manifest is generated. In the manifest, at least one stream in a first set of streams is identified, and each stream in the first set represents at least a portion of the first video data. In the manifest, at least one stream in a second set of streams is also identified, and each stream in the second set represents at least a portion of the second video data. Each stream in the first set is associated in the manifest with a first viewpoint identifier, and each stream in the second set is associated in the manifest with a second viewpoint identifier.
[0033] In some embodiments, each stream in the first set is associated in the manifest with each adaptation set having the first viewpoint identifier as an attribute, and each stream in the second set is associated in the manifest with each adaptation set having the second viewpoint identifier as an attribute.
[0034] In some embodiments, each stream in the first set is associated in the manifest with each compliant set having a first viewpoint identifier in a first descriptor, and each stream in the second set is associated in the manifest with each compliant set having a second viewpoint identifier in a second descriptor.
[0035] In some embodiments, the manifest further includes an attribute indicating a valid range for each viewpoint. In some embodiments, the manifest further includes an attribute indicating a position for each viewpoint. The attribute indicating a position can include Cartesian coordinates or latitude and longitude coordinates.
[0036] In some embodiments, the manifest further includes information indicating, for each viewpoint, at least one time period during which video for each viewpoint is available.
[0037] In some embodiments of a method for generating a manifest, the first video data and the second video data are received in a container file, and in the container file, the first video data is compiled into a first set of tracks and the second video data is compiled into a second set of tracks, each track in the first set of tracks includes a first track group identifier associated with a first viewpoint, and each track in the second set of tracks includes a second track group identifier associated with a second viewpoint. The viewpoint identifier used in the manifest can be equal to each track group identifier in the container file.
[0038] Some embodiments may be implemented by a client device, such as a head-mounted display or a device equipped with other display devices for 360-degree video. In some such methods, a manifest that identifies a plurality of 360-degree video streams is received, where the manifest includes information identifying the viewpoint positions of each respective stream. A first video stream identified in the manifest is obtained and displayed. A user interface element indicating the viewpoint position of a second video stream identified in the manifest is overlaid on the display of the first video stream. In response to the selection of the user interface element, the second video stream is obtained and displayed.
[0039] In some such embodiments, the manifest further includes information identifying at least one valid range of the identified stream, and the client further displays an indication of the valid range.
[0040] In some embodiments, the manifest further includes information identifying the available period of the second video stream, and the user interface element is displayed only during the available period.
[0041] In some embodiments, the manifest further includes information identifying the transition type for the transition from the first video stream to the second video stream. In response to the selection of the user interface element, the client presents a transition having the identified transition type, and after the presentation of the transition, the second video stream is displayed.
[0042] In some embodiments, the manifest further includes information identifying the position of at least one virtual viewpoint. In response to the selection of the virtual viewpoint, the client synthesizes a view from the virtual viewpoint and displays the synthesized view. One or more synthesized views can be used in the transition.
[0043] A method for selecting a projection format is further described. In some embodiments, a client receives a manifest that identifies a plurality of 360-degree video streams. The manifest includes information identifying each projection format of each of the video streams. The manifest further includes information identifying each range of field-of-view sizes for each of the projection formats. The client determines a field-of-view size for display. The client then selects at least one of the video streams such that the determined field-of-view size is included in the identified range of field-of-view sizes for the projection format of the selected video stream. The client obtains at least one of the selected video streams and displays the obtained video stream using the determined field-of-view size.
[0044] Further included in the present disclosure is a system comprising a processor and a non-transitory computer-readable medium storing instructions operable to perform any of the methods described herein when executed by the processor. Further included in the present disclosure is a non-transitory computer-readable storage medium storing one or more container files or manifests generated using the methods disclosed herein.
Advantages of the Invention
[0045] A method and system for signaling information regarding various viewpoints in a novel, multi-viewpoint omnidirectional media presentation are provided.
Brief Description of the Drawings
[0046]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Mode for Carrying Out the Invention
[0047] Exemplary network for implementing an embodiment FIG. 1A is a diagram showing an exemplary communication system 100 in which one or more of the disclosed embodiments can be implemented. The communication system 100 can be a plurality of access systems that provide content such as voice, data, video, messaging, paging, etc. to a plurality of wireless users. The communication system 100 enables such content to be accessed by sharing system resources including wireless bandwidth among a plurality of wireless users. For example, the communication system 100 can use one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), and the like.
[0048] As shown in FIG. 1A, communication system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104, a core network (CN) 106, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate in and / or communicate within a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d can each be referred to as a “station” and / or “STA,” but are configured to transmit and / or receive wireless signals and can also be a user equipment (UE), mobile station, fixed or mobile subscriber unit, subscription-based unit, pager, cellular phone, personal digital assistant (PDA), smartphone, laptop, netbook, personal computer, wireless sensor, hotspot or Mi-Fi device, Internet of Things (IoT) device, watch or other wearable, head-mounted display (HMD), vehicle, drone, medical device and applications (e.g., remote surgery), industrial device and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), home electronics device, device operating in a commercial and / or industrial wireless network, etc. Any of the WTRUs 102a, 102b, 102c, and 102d can be interchangeably referred to as a UE.
[0049] The communication system 100 can also include base station 114a and / or base station 114b. Each of base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as CN106 / 115, the Internet 110, and / or other network 112. By way of example, base stations 114a, 114b can be a base transceiver station (BTS), Node B, eNodeB, home Node B, home eNodeB, gNB, NR Node B, site controller, access point (AP), wireless router, and the like. Although base stations 114a, 114b are each shown as a single element, it will be understood that base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0050] The base station 114a can be part of the RAN 104 / 113, which can also include other base stations and / or network elements (not shown) such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 114a and / or the base station 114b can be configured to transmit and / or receive radio signals at one or more carrier frequencies that can be referred to as a cell (not shown). These frequencies can be an authorized spectrum, an unlicensed spectrum, or a combination of authorized and unlicensed spectra. A cell can provide coverage for providing wireless services to a specific geographical area that can be relatively fixed or variable over time. A cell can be further divided into cell sectors. For example, the cell associated with the base station 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, i.e., one for each sector of the cell. In an embodiment, the base station 114a can use multiple-input multiple-output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0051] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d via a wireless interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The wireless interface 116 can be established using any suitable radio access technology (RAT).
[0052] More specifically, as described above, the communication system 100 can be a plurality of access systems, and can use one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in RAN104 / 113, and the WTRUs 102a, 102b, 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that can establish radio interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0053] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the radio interface 116 using Long Term Evolution (LTE), and / or LTE-Advanced (LTE-A), and / or LTE-Advanced Pro (LTE-A Pro).
[0054] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access that can establish the radio interface 116 using New Radio (NR).
[0055] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c can implement both LTE radio access and NR radio access using, for example, the dual connectivity (DC) principle. Accordingly, the radio interfaces utilized by the WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and / or by transmissions sent between multiple types of base stations (e.g., eNBs and gNBs).
[0056] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.
[0057] The base station 114b in FIG. 1A can be, for example, a wireless router, a home Node B, a home eNodeB, or an access point, and can utilize any suitable RAT to facilitate wireless connection in a localized area such as a workplace, home, vehicle, campus, industrial facility, skywalk (e.g., used for drones), driveway, and similar locations. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0058] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, 102d. The data can have various Quality of Service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. CN 106 / 115 can provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or implement high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as RAN 104 / 113, or a different RAT. For example, in addition to being connected to a RAN 104 / 113 that can utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0059] CN106 / 115 can also act as a gateway for WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN 108 can include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 110 can include a global system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 can include another CN connected to one or more RANs that can use the same RAT as the RAN 104 / 113 or a different RAT.
[0060] Some, or all, of the WTRU102a, 102b, 102c, 102d in the communication system 100 can include a multimode function (e.g., the WTRU102a, 102b, 102c, 102d can include multiple transceivers for communicating with various wireless networks via various wireless links). For example, the WTRU102c shown in Figure 1A can be configured to communicate with a base station 114a that can use cellular-based wireless technology and a base station 114b that can use IEEE802 wireless technology.
[0061] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 can include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. The WTRU 102 can include any sub-combination of the foregoing elements, although it will be understood to be consistent with embodiments.
[0062] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DCP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 1B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0063] The transmit / receive element 122 can be configured to transmit signals to, or receive signals from, a base station (e.g., base station 114a) via the wireless interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be a light emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF and optical signals. It will be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0064] Although the transmit / receive element 122 is shown as a single element in FIG. 1B, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the wireless interface 116.
[0065] The transceiver 120 can be configured to modulate signals transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As previously mentioned, the WTRU 102 can have a multimode functionality. Thus, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs such as, for example, NR and IEEE 802.11.
[0066] The processor 118 of the WTRU 102 can be coupled to, and can receive user input data from, a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit, or an organic light emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can access information from, and store data in, any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 can include a random access memory (RAM), a read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from, and store data in, a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0067] The processor 118 can receive power from a power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 can include one or more dry cells (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, a fuel cell, and the like.
[0068] Processor 118 can also be coupled to a GPS chipset 136 that can be configured to provide position information (e.g., longitude and latitude) regarding the current location of WTRU 102. In addition to, or instead of, information from GPS chipset 136, WTRU 102 can receive position information from a base station (e.g., base stations 114a, 114b) via radio interface 116 and / or can determine its position based on the timing of signals received from two or more neighboring base stations. It will be understood that WTRU 102 can obtain position information by any suitable positioning method while remaining consistent with the embodiments.
[0069] Processor 118 can further be coupled to other peripheral devices 138 that can include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripheral devices 138 can include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. Peripheral devices 138 can include one or more sensors, where the sensors can be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0070] The WTRU 102 can include full-duplex radio, in which case some or all of the transmissions and receptions of signals (e.g., associated with a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception)) coincide and / or can occur simultaneously. The full-duplex radio can include an interference management unit that can reduce and / or substantially eliminate self-interference by signal processing by hardware (e.g., choke) or by a processor (e.g., by a separate processor (not shown) or by processor 118). In embodiments, the WRTU 102 can include half-duplex radio, in which case some or all of the transmissions and receptions of signals (e.g., associated with a particular subframe for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0071] In FIGS. 1A-1B, the WTRU is shown as a wireless terminal, but in some representative embodiments, it is contemplated that such a terminal can use a wired communication interface to the communication network (e.g., temporarily or permanently).
[0072] In a representative embodiment, the other network 112 can be a WLAN.
[0073] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for a BSS and one or more stations (STAs) associated with the AP. The AP can access or have an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or from the BSS. Traffic from outside the BSS to an STA can reach and be delivered to the STA through the AP. Traffic originating from an STA to a destination outside the BSS can be sent to the AP and delivered to each destination. Traffic between STAs within the BSS can, for example, be sent through the AP, in which case the source STA can send the traffic to the AP and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS is considered to be and / or can be called peer-to-peer traffic. Peer-to-peer traffic can be sent between the source and destination STAs (e.g., directly between them) using a direct link setup (DLS). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) can communicate directly with each other. The IBSS mode of communication may also be referred to herein as the "ad hoc" mode of communication.
[0074] When using the 802.11ac infrastructure operation mode or a similar operation mode, the AP can send beacons on a fixed channel such as the primary channel. The primary channel can have a fixed width (e.g., a bandwidth of 20 MHz), or a width that is dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, for example, in an 802.11 system, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) can sense the primary channel. If the primary channel is sensed / detected as busy and / or determined to be busy by a particular STA, the particular STA may back off. Only one STA (e.g., only one station) can transmit at any given time in a given BSS.
[0075] A High Throughput (HT) STA can use a 40 MHz wide channel for communication, which is formed, for example, by combining the primary 20 MHz channel with an adjacent or non - adjacent 20 MHz channel.
[0076] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining adjacent 20 MHz channels. A 160 MHz channel can be formed by combining eight adjacent 20 MHz channels or by combining two non-adjacent 80 MHz channels, which can be referred to as an 80+80 configuration. In the case of the 80+80 configuration, after channel encoding, the data can pass through a segment parser that can split the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time domain processing can be performed separately for each stream. The streams can be mapped to two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0077] The sub-1 GHz operation mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using the non-TVWS spectrum. According to an exemplary embodiment, 802.11ah can support meter type control / machine type communication, such as MTC devices in a macro coverage area. The MTC device can have some functions, such as limited functionality including support for some and / or limited bandwidths (e.g., support only for that). The MTC device can include a battery having a battery life exceeding a threshold (e.g., to maintain a very long battery life).
[0078] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STA that supports the least bandwidth operating mode among all STAs operating in the BSS. In the example of 802.11ah, the primary channel can be 1 MHz wide for an STA (e.g., an MTC type device) that supports the 1 MHz mode (e.g., only supports) even when the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the state of the primary channel. For example, if the primary channel is busy due to an STA (supporting only the 1 MHz operating mode) transmitting to the AP, the entire available frequency band may be considered busy even if most of the frequency band remains idle and available.
[0079] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is from 6 MHz to 26 MHz according to national regulations.
[0080] In the figures of FIGS. 1A-1B, and the corresponding descriptions of FIGS. 1A-1B, one or more of the functions described herein with respect to one or more of WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other device described herein, or all, can be implemented by one or more emulation devices (not shown). The emulation device can be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, the emulation device can be used to test other devices and / or to simulate the network and / or WTRU functions.
[0081] The emulation device can be designed to perform one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices can implement one or more, or all, of the functions, but are implemented and / or deployed, either fully or partially, as part of a wired and / or wireless communication network, and / or are deployed, to test other devices within the communication network. One or more emulation devices can implement one or more, or all, of the functions, but are temporarily implemented / or deployed as part of a wired and / or wireless communication network. The emulation device can be directly coupled to another device to perform the test and / or can perform the test using wireless communication over the air.
[0082] One or more emulation devices can perform one or more functions, including all functions, but are not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be used in a test scenario in a test laboratory and / or a non-deployed (e.g., for testing) wired and / or wireless communication network to perform tests on one or more components. One or more emulation devices can be used as test equipment. For transmitting and / or receiving data, wireless communication can be used by the emulation device directly through RF coupling and / or via an RF circuit (which can include one or more antennas).
[0083] Detailed description Problems addressed in some embodiments Omnidirectional video provides a 360-degree experience that allows viewers to watch video in all directions around a central viewing position. However, viewers have generally been limited to a single viewing perspective and have not been able to navigate the scene by changing their viewing perspective. For large events such as the opening ceremony of the Olympic Games, an NFL or NBA tournament, or a carnival parade, a single 360° camera is not sufficient to capture the entire scene. By capturing the scene from multiple perspectives and allowing the user to switch between different perspectives while watching the video, a more enhanced experience can be provided. Figure 3 shows a user interface that can be presented to the user in some embodiments to show the available perspectives. In this example, the user interface displays a top-down view of the venue and provides an indication of the positions of the available perspectives. In this case, perspective 302 is the active perspective (the perspective from which the user is currently experiencing the presentation) and is highlighted. Other perspectives such as perspectives 304, 306, 308, 310, 312, 314, 316, etc. are displayed and their availability can be indicated, but are not currently selected by the user.
[0084] During regeneration, the user interface, such as that shown in FIG. 3, is superimposed on a frame rendered at one of the four corners, for example, and the user can select various viewpoints using a user input device such as a touch screen or an HMD controller. Next, a viewpoint switch is triggered, and the user's view is transitioned so that a frame from the target viewpoint is rendered on the display. In some embodiments, a transition effect (e.g., fade-out / fade-in) accompanies the transition between viewpoints.
[0085] FIG. 4 shows another user design example where the positions of available viewpoints are shown using icons presented as an overlay on content 400 displayed on a head-mounted display. The position of each viewpoint icon in the user's view corresponds to the spatial position of the available viewpoint. In the example of FIG. 4, icons 406, 414 can be displayed to correspond to viewpoints 306, 314 of FIG. 3, respectively. The viewpoint icons can be rendered using the correct depth effect to allow the user to perceive each viewpoint position in the 3D space within the scene. For example, icon 416 (corresponding to viewpoint position 316) can be displayed in a larger size than icons 406, 414 to indicate that the viewpoint corresponding to icon 416 is closer to the current viewpoint. The user can select a viewpoint icon to switch the user's view of the rendered scene to the associated viewpoint.
[0086] In an exemplary embodiment, to enable support for multiple viewpoints, information about the available viewpoints is signaled to a player (which can be an omnidirectional media player with a DASH client operating on a user device such as an HMD). This information can include aspects such as the number of available viewpoints, the position and extent of each viewpoint, and when the video data is available for each viewpoint. Further, since most omnidirectional media presentations are experienced via a head-mounted display, sudden changes in viewpoints may feel unnatural to viewers immersed in the virtual environment. Therefore, it is preferable to support a viewpoint transition effect that provides a smooth transition when the user changes their viewpoint. These transitions can also be used by content producers to guide the user experience.
[0087] Grouping of Viewpoint Media Components In some embodiments, media samples for omnidirectional media content having multiple viewpoints are stored in several tracks within a container file. A video player that plays or streams the content operates to identify which track belongs to which viewpoint. To enable this, a mapping is made between the media tracks in the file and the viewpoints to which they belong. In some embodiments, this mapping is signaled at the media container (file format) level. In some embodiments, this mapping is signaled at the transport protocol level (DASH).
[0088] Signaling at the Media Container Level (File Format) In ISO / IEC 14496-12 (ISO BMFF), the TrackGroupBox is defined to enable the grouping of several tracks in a container file that share certain characteristics or have specific relationships. The track group box contains zero or more boxes, and the specific characteristics or relationships are indicated by the box types of the boxes it contains. The boxes it contains include identifiers that can be used to conclude that tracks belong to the same track group. Tracks that contain the same type of boxes within the TrackGroupBox and have the same identifier value within these contained boxes belong to the same track group. aligned(8) class TrackGroupBox extends Box('trgr') { } The track group type is defined by extending the TrackGroupTypeBox that stores the track_group_id identifier and the four-character code identifying the group type, track_group_type. The pair of track_group_id and track_group_type identifies the track group within the file.
[0089] To group together several media tracks belonging to a single viewpoint, in some embodiments, a new group type (ViewpointGroupTypeBox) is defined as follows: aligned(8) class ViewpointGroupTypeBox extends TrackGroupTypeBox('vpgr') { / / additional viewpoint data can be defined here }
[0090] In some embodiments, the media has a ViewpointGroupTypeBox within a TrackGroupBox, and tracks belonging to the same viewpoint have the same value of track_group_id in each ViewpointGroupTypeBox. The 3DoF+ omnidirectional media player can thus identify the available viewpoints by parsing each track in the container and examining the number of unique track_group_id values within the ViewpointGroupTypeBox for each track.
[0091] Transport-level signaling (DASH) The OMAF standard defines the interfaces related to delivery for DASH. In some embodiments, information related to different viewpoints is signaled in the media presentation descriptor. In DASH, each media component is represented by an AdaptationSet element in the MPD. In some embodiments, AdaptationSet elements belonging to the same viewpoint are grouped by defining additional attributes for the AdaptationSet element or by adding descriptors to the AdaptationSet where a viewpoint identifier is provided.
[0092] Some descriptors are defined in the MPEG-DASH standard. These include SupplementalProperty descriptors that can be used by media presentation authors, indicating that the descriptor contains auxiliary information that can be used by DASH clients for optimized processing. The semantics of the information being signaled are specific to the manner in which it is used, which is identified by the @schemeIdUri attribute. In the present disclosure, several new XML elements and attributes are described for signaling information related to viewpoints. The new elements can be defined in the same namespace as those defined in the latest version of the OMAF standard (urn:mpeg:mpegI:omaf:2017), or in a separate new namespace (urn:mpeg:mpegI:omaf:2019) to distinguish between OMAF v1 and OMAF v2 features. For the sake of explanation, the namespace (urn:mpeg:mpegI:omaf:2017) is used for the remainder of this document.
[0093] Embodiments are described in which an @schemeIdUri attribute is added to a SupplementalProperty element equal to "urn:mpeg:mpegI:omaf:2017:ovp" to identify and describe the viewpoint to which a media component belongs. Such a descriptor is referred to herein as an OMAF Viewpoint (OVP) descriptor. In some embodiments, at the adaptation set level, there can be at most one OVP descriptor. The OVP descriptor can have an @viewpoint_id attribute that holds a value representing a unique viewpoint identifier. Examples of the semantics for @viewpoint_id are given in Table 1. AdaptationSet elements having the same @viewpoint_id value can be recognized by the player as belonging to the same viewpoint.
[0094] [Table 1]
[0095] Signaling of Viewpoint Information In some of the methods described herein, in order for a player to identify attributes belonging to different viewpoints (e.g., spatial relationships between viewpoints, availability of viewpoints, etc.), additional metadata describing the viewpoints is signaled in a container file (or in the case of streaming, in an MPD file). Examples of viewpoint attributes signaled in some embodiments include viewpoint position, viewpoint effective range, viewpoint type, and viewpoint availability. The viewpoint position specifies the viewpoint position within the 3D space of the captured scene. The viewpoint effective range is the distance from the viewpoint within which an object can be rendered at a certain quality level. The certain quality level can be, for example, a minimum quality level, a quality level exceeding a known quality threshold, a guaranteed quality level, or a quality level recognized or considered acceptable by the provider of the omnidirectional media content. For example, an object within the effective range is of sufficient size in the rendered image, provides a resolution that offers good quality, and guarantees an acceptable viewing experience for the user. The viewpoint effective range depends on the characteristics of the capture device (e.g., camera sensor resolution, field of view, etc.). The effective range can be at least partially determined by the camera lens density representing the number of lenses integrated into a 360-degree video camera.
[0096] Figure 5 shows examples of the effective ranges of various cameras. In this example, viewpoints 502, 504, 506, 508, 510, 512, 514, and 516 are shown along with dotted circles 503, 505, 507, 509, 511, 513, 515, and 517, indicating the effective range for each viewpoint. The omnidirectional cameras located at viewpoints 502 and 510 can include more lenses to cover a wider area, and thus, the effective ranges of viewpoints 502 and 510 can cover penalty areas 520 and 522, as shown in Figure 5. In this example, cameras along the sides of the field can have fewer lenses, and thus, the effective ranges of these viewpoints (504, 506, 508, 512, 514, 516) can be smaller than those of viewpoints 502 and 510. Generally, an omnidirectional camera having more lenses, more component cameras, or high-quality component cameras (e.g., component cameras having a high-quality optical system, high resolution, and / or high frame rate, etc.) can be associated with a higher effective range.
[0097] In another embodiment, the viewpoint effective range can be at least partially determined by camera lens parameters such as focal length, aperture, depth of field, and focus distance. The effective range can define a minimum range and a maximum range, and the effective range is between the minimum range and the maximum range where no stitching error occurs.
[0098] Viewpoints can be classified as actual viewpoints or virtual viewpoints. An actual viewpoint is a viewpoint where an actual capture device is placed to capture a scene from that viewpoint position. A virtual viewpoint refers to a viewpoint where the rendering of the viewport at that position requires further processing such as view synthesis and can utilize auxiliary information and / or video data from one or more other (e.g., actual) viewpoints. The availability of a viewpoint specifies at which time during a presentation media data is available for the viewpoint.
[0099] User interactions with the viewport scene, such as zooming in or out, can be supported within the valid range. The virtual viewpoint only needs to be identifiable within the valid range of one or more cameras. The valid range can also be used as a criterion for generating a transition path. For example, the transition from viewpoint A to viewpoint B can include multiple transition viewpoints if the valid ranges of the viewpoints cover the transition path.
[0100] Media Container Level Signaling of Viewpoint Information In ISO BMFF, information related to viewpoints for static viewpoints can be signaled in a "meta" box (MetaBox) at the file level. The "meta" box holds static metadata and contains only one mandatory "hdlr" box (HandlerBox), which declares the structure or format of the MetaBox. In some embodiments, for OMAF v2 metadata, a four-character code "omv2" is used for the handler_type value in the "hdlr" box. To identify the viewpoints available in a file, some embodiments use a box called OMAFViewpointListBox, which contains a list of OMAFViewpointInfoBox instances. Each OMAFViewpointInfoBox holds information about a specific viewpoint. An example of the syntax of OMAFViewpointListBox is as follows. Box Type: ’ovpl’ Container: MetaBox Mandatory: No Quantity: Zero or one aligned(8) class OMAFViewpointListBox extends Box(’ovpl’) { unsigned int(16) num_viewpoints; OMAFViewpointInfoBox viewpoints[]; }
[0101] An example of the semantics of the OMAFViewpointListBox is as follows: num_viewpoints indicates the number of viewpoints in the media file. viewpoints is a list of OMAFViewpointInfoBox instances.
[0102] An example of the syntax of the OMAFViewpointInfoBox is given below. Box Type: ’ovpi’ Container: OMAFViewpointListBox Mandatory: No Quantity: Zero or more aligned(8) class OMAFViewpointInfoBox extends Box(’ovpi’) { unsigned int(16) viewpoint_id; bit(1) effective_range_flag; bit(1) virtual_viewpoint_flag; bit(1) dynamic_position_flag; bit(5) reserved; if (effective_range_flag == 1) { unsigned int(32) effective_range; } unsigned int(32) num_availability_intervals; OMAFViewpointPositionGlobalBox(); / / optional OMAFViewpointPositionCartesianBox(); / / optional OMAFViewpointAvailabilityIntervalBox availability_intervals[]; Box other_boxes[];}
[0103] An example of the semantics of the OMAFViewpointInfoBox is as follows: The viewpoint_id is a unique identifier for the viewpoint.
[0104] The virtual_viewpoint_flag indicates whether the viewpoint is a virtual viewpoint (there is no capture device placed at the viewpoint position) or a captured viewpoint. The information necessary to generate a virtual viewpoint is signaled in the OMAFVirtualViewpointConfigBox.
[0105] The dynamic_position_flag indicates whether the position is static or dynamic. If this flag is set, the position of the viewpoint is provided using a timed-metadata track. Otherwise, the position is indicated by the OMAFViewpointPositionGlobalBox and / or the OMAFViewpointPositionCartesianBox in this OMAFViewpointInfoBox.
[0106] The effective_range is the radius that defines a volumetric sphere centered on the viewpoint within which the viewpoint provides rendering at a certain quality (e.g., minimum level of quality, quality exceeding a known quality threshold, guaranteed quality level, or quality level recognized or considered acceptable by the provider of the omnidirectional media content, etc.).
[0107] The num_availability_intervals indicates the number of time intervals during which this viewpoint is available.
[0108] availability_intervals is a list of OMAFViewpointAvailabilityIntervalBox instances.
[0109] In some embodiments, when the viewpoint position in space changes over time, the position information is signaled using a time - specified metadata track. The time - specified metadata track is a track within a media container (ISO BMFF) file, and the samples represent dynamic metadata information. For dynamic viewpoint position information, some embodiments use a time - specified metadata track with a sample entry type of “vpps”. The sample entry for this track can be as follows. aligned(8) class OMAFDynamicViewpointSampleEntry extends MetadataSampleEntry(‘vpps’) { unsigned int(16) viewpoint_id; unsigned int(3) coordinate_system_type; bit(5) reserved; } An example of the semantics of OMAFDynamicViewpointSampleEntry is as follows.
[0110] viewpoint_id is the identifier of the viewpoint whose position is defined by the sample of this time - specified metadata track.
[0111] coordinate_system_type indicates the coordinate system used to define the position of the viewpoint.
[0112] In some embodiments, the samples for the viewpoint position metadata track have the following structure. aligned(8) class OMAFViewpointPositionSample { if (coordinate_system_type == 1) { ViewpointPositionGlobalStruct(); } else if (coordinate_system_type == 2) { ViewpointPositionCartesianStruct(); } }
[0113] The sample format can depend on the coordinate system type defined in the sample entry of the timed - metadata track. ViewpointPositionGlobalStruct and ViewpointPositionCartesianStruct are described in more detail below.
[0114] Transport Protocol Level Signaling of Viewpoint Information To identify and describe the set of viewpoints available in a media presentation, some embodiments include a SupplementaryProperty descriptor at the Period level. This descriptor can have a @schemeIdUri equal to "urn:mpeg:mpegI:omaf:2017:ovl", and is referred to herein as the OMAF Viewpoint List (OVL) descriptor. In some embodiments, at the Period level, there can be at most one OVL descriptor. The OVL descriptor can include at least one ovp element. The ovp element has an @id attribute with a value representing a unique viewpoint identifier and can include sub - elements with information about the viewpoint.
[0115] Table 2 is a list of examples of elements and attributes used to signal viewpoint information in an MPD file for a DASH client. Further details are given below.
[0116] [Table 2 - 1]
[0117]
Table 2-2
[0118] In Table 2 and other tables in this disclosure, elements are in bold, attributes are not in bold and are preceded by @. "M" indicates that the attribute is mandatory in the specific embodiment shown in the table, "O" indicates that the attribute is optional in the specific embodiment shown in the table, "OD" indicates that the attribute is an optional one with a default value in the specific embodiment shown in the table, and "CM" indicates that the attribute is conditionally mandatory in the specific embodiment shown in the table. <minoccurs> .. <maxoccurs>(N=not limited).
[0119] The data types for various elements and attributes are those defined in the XML schema. The XML schema for ovp is provided in the following section, "XML Schema for DASH Signaling".
[0120] Viewpoint position The "actual" viewpoints correspond to 360° video cameras placed at various positions to capture a scene from different advantageous points. In some embodiments, a viewpoint can represent a view from a virtual position. A virtual position can represent a point not associated with the location of a physical camera. A virtual position can represent a point where synthetic content can be rendered, or where content captured by one or more cameras at other (actual) viewpoints can be transformed, processed, or combined to synthesize a virtual view. To provide useful information to the player regarding the camera settings and their layout used to capture the scene, in some embodiments, the spatial relationship between viewpoints is signaled by providing the position of each viewpoint. The position information can be represented in various ways in various embodiments. In some embodiments, global geolocation coordinates similar to those used by the GPS system can be used to identify the location of the camera / viewpoint. Alternatively, a Cartesian coordinate system can be used for positioning.
[0121] Media Container Level Signaling of Viewpoint Position Described herein are two examples of boxes that can be used to identify the position of a viewpoint when present in an OMAFViewpointInfoBox, namely, the OMAFViewpointPositionGlobalBox and the OMAFViewpointPositionCartesianBox. In some embodiments, these boxes are optional. An exemplary syntax for the proposed position boxes is given below. Additional boxes can also be introduced to provide position information based on other coordinate systems. Box Type: ’vpgl’ Container: OMAFViewpointInfoBox Mandatory: No Quantity: Zero or one aligned(8) class ViewpointPositionGlobalStruct() { signed int(32) longitude; signed int(32) latitude; signed int(32) altitude; } aligned(8) class OMAFViewpointPositionGlobalBox extends Box(’vpgl’) { ViewpointPositionGlobalStruct(); }
[0122] In some embodiments, double precision or floating point types are used for the longitude, latitude, and / or altitude values. Box Type: ’vpcr’ Container: OMAFViewpointInfoBox Mandatory: No Quantity: Zero or one aligned(8) class ViewpointPositionCartesianStruct() { signed int(32) x; signed int(32) y; signed int(32) z; } aligned(8) class OMAFViewpointPositionCartesianBox extends Box('vpcr') { ViewpointPositionCartesianStruct(); }
[0123] Transport Protocol Level Signaling of Viewpoint Position To signal the position of a viewpoint, in some embodiments, an ovp:position element can be added to the ovp element. This element can contain an ovp:position:global element and / or an ovp:position:cartesian element. In some embodiments, at most one of each of these elements is present within the ovp:position element. The attributes of the ovp:position:global element provide the position of the viewpoint in terms of degrees, with respect to global geolocation coordinates. In some embodiments, the ovp:position:global element has three attributes, namely, @longitude, @latitude, and @altitude. In some embodiments, the @altitude attribute is optional and may not be present. The attributes of the ovp:position:cartesian attribute provide the position of the viewpoint with respect to Cartesian coordinates. In some embodiments, three attributes are defined for the ovp:position:cartesian element: @x, @y, and @z, where only @z is optional.
[0124] Viewpoint Availability In some cases, the viewpoint may not be available for the entire duration of the media presentation. Thus, in some embodiments, the availability of the viewpoint is signaled before media samples for that viewpoint are processed. This allows the player to process samples for tracks belonging to a particular viewpoint only when the viewpoint is available.
[0125] The change in the availability of viewpoints over time is shown in FIG. 6. At time t1, only viewpoints 601, 602, 603, and 604 are available. Subsequently, during the presentation at time t2, a penalty shot is awarded to one of the teams, and most of the players are near the right goal. At that point in time, two additional viewpoints 605 and 606 become available to the user until time t3. The time interval between t2 and t3 is the interval available for viewpoints 605 and 606. Using the viewpoint availability information (e.g., received from the server), the player or streaming client operates to show the user the availability of additional viewpoints at time t2 during playback, for example, by using the UI shown in FIG. 3 or FIG. 4. When the available interval starts, the player can present the user with options to switch to any or all of the available (e.g., newly available) viewpoints during the available interval. As shown in FIG. 6, the user may be given the option to switch to viewpoint 605 or 606 starting from time t2. At the end of the available interval, the player can remove the option to switch to viewpoints that are no longer available after the available interval ends. In some embodiments, when the available interval ends (e.g., at time t3 shown in FIG. 6), if the user is still in one of these viewpoints, the user can return to the viewpoint where the user was before switching to a viewpoint that is no longer available (e.g., viewpoint 605 or 606 shown in FIG. 6). In some embodiments, the available interval for viewpoints can also be signaled for virtual viewpoints. However, the availability of these viewpoints depends on the availability of other reference viewpoints used to support the rendering of the virtual viewpoint as well as any auxiliary information.
[0126] Media Container-Level Signaling of Viewpoint Availability In some embodiments, a box (OMAFViewpointAvailaibilityIntervalBox) is introduced to signal the available intervals. Zero or more instances of this box may be present in the OMAFViewpointInfoBox. When no OMAFViewpointAvailaibilityIntervalBox instance exists for a viewpoint, this indicates that the viewpoint is available for the entire duration of the presentation. Box Type: ’vpai’ Container: OMAFViewpointInfoBox Mandatory: No Quantity: Zero or more aligned(8) class OMAFViewpointAvailabilityIntervalBox extends Box(‘vpai’) { bit(1) open_interval_flag; bit(7) reserved; unsigned int(64) start_time; / / mandatory unsigned int(64) end_time; }
[0127] An example of the semantics for the OMAFViewpointAvailabilityIntervalBox is as follows: The open_inverval_flag is a flag that indicates whether the available interval is an open interval (value 1) where the viewpoint is available from start_time until the end of the presentation, or a closed interval (value 0). If the flag is set (value 1), the end_time field does not exist in this box.
[0128] start_time is the presentation time when the viewpoint is available (corresponding to the composition time for the first sample in the interval).
[0129] end_time is the subsequent presentation time when the viewpoint is no longer available (corresponding to the composition time of the last sample in the interval).
[0130] Transport protocol level signaling of viewpoint availability In some embodiments, one or more ovp:availability elements may be added to an instance of the ovp element to signal the availability of the viewpoint in the MPD file. This element means the available period and has two attributes, @start and @end, indicating the presentation time when the viewpoint is available and the presentation time of the last sample of the available interval, respectively.
[0131] Virtual viewpoint In some embodiments, virtual viewpoints are generated using an omnidirectional virtual view synthesis process. In some embodiments, this process utilizes one or more input (reference) viewpoints, their associated depth maps, and additional metadata that describes a transformation vector between the input viewpoint position and the virtual viewpoint position. In some such embodiments, each pixel of the input omnidirectional viewpoint maps pixels of the reference viewpoint's equirectangular frame to points in 3D space and then projects them to a target virtual viewpoint to map them to positions in the virtual viewpoint sphere. One such view synthesis process is described in great detail in "Extended VSRS for 360-degree video", MPEG121, Gwangju, Korea, January 2018, m41990 and is shown in FIG. 8. In the example of FIG. 8, point 802 has a position described by angular coordinates (φ, θ), and depth z, relative to input viewpoint 804. In generating virtual viewpoint 806 displaced by vector (Tx, Ty, Tz) from input viewpoint 804, the angular coordinates (φ’, θ’) for point 802 are found with respect to virtual viewpoint 806. The displacement vector (Tx, Ty, Tz) can be determined based on the viewpoint position signaled in a container file, manifest, time-stamped metadata track, or other.
[0132] In various embodiments, various techniques can be used to generate virtual viewpoints. Virtual viewpoint frames synthesized from various reference viewpoints can then be merged together using a blending process to generate the final equirectangular frame at the virtual viewpoint. Holes that appear in the final frame due to occlusion of the reference viewpoints can be processed using repair and hole-filling steps.
[0133] The virtual viewpoint is a non-captured viewpoint. The viewport can be rendered at the virtual viewpoint using video data from other viewpoints and / or other auxiliary information. In some embodiments, the information used to render the scene from the virtual viewpoint is signaled in the OMAFVirtualViewpointConfigBox that exists in the OMAFViewpointInfoBox when the virtual_viewpoint flag is set. In some embodiments, the OMAFVirtualViewpointConfigBox can be defined as follows. Box Type: vvpc’ Container: OMAFViewpointInfoBox Mandatory: No Quantity: Zero or more aligned(8) class OMAFVirtualViewpointConfigBox extends Box(‘vvpc’) { unsigned int(5) synthesis_method; unsigned int(3) num_reference_viewpoints; unsigned int(16) reference_viewpoints[]; / / optional boxes but no fields }
[0134] Examples of semantics for the OMAFVirtualViewpointConfigBox fields are given below.
[0135] synthesis_method indicates which synthesis method is used to generate the virtual viewpoint. The value of synthesis_method can be an index into a list of view synthesis methods. For example, depth-image-based rendering, image-warping-based synthesis, etc.
[0136] The num_reference_viewpoints indicates the number of viewpoints used as a reference in the synthesis of virtual viewpoints.
[0137] The reference_viewpoints is a list of viewpoint ids used as a reference when synthesizing the viewport for this viewpoint.
[0138] In another embodiment, the identifier of the track containing the information required for the synthesis process is directly signaled in the virtual viewpoint configuration box, which can be implemented as follows. aligned(8) class OMAFVirtualViewpointConfigBox extends Box('vvpc') { unsigned int(5) synthesis_method; unsigned int(3) num_reference_tracks; unsigned int(16) reference_track_ids[]; / / optional boxes but no fields }
[0139] An example of the semantics of the OMAFVirtualViewpointConfigBox fields for this embodiment is as follows.
[0140] The synthesis_method indicates which synthesis method is used to generate the virtual viewpoint. The value of synthesis_method is an index into the list of view synthesis methods. For example, depth-image-based rendering, image-warping-based synthesis, etc.
[0141] The num_reference_tracks indicates the number of tracks in the container file used as a reference in the synthesis of virtual viewpoints.
[0142] The reference_track_ids is a list of track identifiers for the tracks used in the composition of the viewport for this perspective.
[0143] Signaling of Viewpoint Groups In large-scale events such as the FIFA World Cup, some events may be held in parallel at different venues or locations. For example, some competitions may be held at different stadiums, perhaps in different cities. In some embodiments, viewpoints can be grouped based on the geo-location of the event / venue. In some embodiments, the ViewpointGroupStruct structure is used to store information about groups of viewpoints within a media container file. An example of the syntax of this structure is as follows. aligned(8) class ViewpointGroupStruct() { unsigned int(8) viewpoint_group_id; signed int(32) longitude; signed int(32) latitude; unsigned int(8) num_viewpoints; unsigned int(16) viewpoint_ids[]; string viewpoint_group_name; }
[0144] Examples of the semantics of the fields of ViewpointGroupStruct are as follows.
[0145] The viewpoint_group_id is a unique id that identifies the viewpoint group.
[0146] The longitude is the longitude coordinate of the geo-location of the event / venue where the viewpoint is located.
[0147] latitude is the latitude coordinate of the geolocation of the event / venue where the viewpoint is located.
[0148] num_viewpoints is the number of viewpoints within the viewpoint group.
[0149] viewpoint_ids is an array having the ids of the viewpoints that are part of the viewpoint group.
[0150] viewpoint_group_name is a string having a name that describes the group.
[0151] To signal the viewpoint groups available within a media container file, an OMAFViewpointGroupsBox can be added to the MetaBox within an ISO BMFF container file. An example of the syntax of the OMAFViewpointGroupsBox is given below. Box Type: ’ovpg’ Container: MetaBox Mandatory: No Quantity: Zero or one aligned(8) class OMAFViewpointGroupsBox extends Box(’ovpg’) { unsigned int(8) num_viewpoint_groups; ViewpointGroupStruct viewpoint_groups[]; }
[0152] Examples of the semantics for the fields of this box are as follows. num_viewpoint_groups is the number of viewpoint groups.
[0153] The viewpoint_groups is an array of ViewpointGroupStruct instances and provides information about each viewpoint group.
[0154] In the case of transport protocol level signaling (e.g., DASH), the ovg element is defined to signal the viewpoint groups available in the media presentation and can be signaled in the OVL descriptor described above. The OVL descriptor can contain one or more ovg elements. The ovg element has an @id attribute with a unique viewpoint group identifier and values representing other attributes that describe the group. Table 3 lists the attributes of an example of the ovg element.
[0155]
Table 3
[0156] Signaling of Viewpoint Transition Effects In this specification, the following examples of transition types are disclosed, namely, basic transitions, viewpoint path transitions, and auxiliary information transitions. A basic transition is a predefined transition that can be used when switching from one viewpoint to another. An example of such a transition is a fade-to-black transition, in which case the rendered view gradually fades out to black and then fades in to the frame from the new viewpoint. A viewpoint path transition allows the content producer to specify a path that the player can follow as it crosses other viewpoints when switching to the target viewpoint. An auxiliary information transition is a transition that relies on auxiliary information provided by the content producer on a separate track. For example, the auxiliary track can contain depth information that can be used to render intermediate virtual views as the viewport moves from the first viewpoint to the target viewpoint.
[0157] In some embodiments, the transition can be based on the rendering of intermediate virtual views. This can be done, for example, using a view synthesis process such as depth-image-based rendering (DIBR) as described in C. Fehn, "Depth-image-based rendering (DIBR), compression, and transmission for a new approach on 3D-TV", SPIE Stereoscopic Displays and Virtual Reality Systems XI, vol. 5291, May 2004, pages 93 - 104. DIBR uses depth information to project pixels in a 2D plane to their positions in 3D space and re-projects them onto another plane. Since there is no capture device at these intermediate viewpoints (e.g., no 360-degree camera), they are referred to herein as virtual viewpoints. The number of intermediate virtual viewpoints rendered between the source and destination viewpoints determines the smoothness of the transition and also depends on the capabilities of the player / device and the availability of auxiliary information for these intermediate viewpoints.
[0158] FIG. 7 shows an embodiment using virtual viewpoints. In the example of FIG. 7, only viewpoints 702, 704, 706, 708, 710, 712, 714, 716 are viewpoints using the capture device, and the remaining intermediate viewpoints (703, 705, 707, 709, 711, 713, 715, 717) are virtual viewpoints. Other types of auxiliary information include the following, namely, a point cloud stream, additional reference frames from neighboring viewpoints (to improve the quality of the virtual view), and occlusion information (to support the hole filling step in the view synthesis process and improve the quality of the virtual view obtained at the intermediate viewpoints). In some embodiments, the point cloud stream is used to enable the rendering of virtual views at any viewpoint position between the source and destination viewpoints. In some embodiments, the point cloud is rendered using the techniques described in Paul Rosenthal, Lars Linsen, "Image-space point cloud rendering", Proceedings of Computer Graphics International, pages 136-143, 2008.
[0159] Media Container Level Signaling of Viewpoint Transition Effects Some embodiments operate to signal the transition effects between pairs of viewpoints in a container file as a list of boxes in a new OMAFViewpointTransitionEffectListBox that can be placed in the MetaBox at the file level. In some embodiments, at most one instance of this box exists within the MetaBox. The boxes in the OMAFViewpointTransitionEffectListBox are instances of the OMAFViewpointTransitionBox. Examples of the syntax of two boxes are given below. Box Type: ’vptl’ Container: MetaBox Mandatory: No Quantity: Zero or one aligned(8) class OMAFViewpointTransitionEffectListBox extends Box(‘vptl’) { OMAFViewpointTransitionBox transitions[]; } Box Type: ’vpte’ Container: OMAFViewpointTransitionEffectListBox Mandatory: No Quantity: One or more aligned(8) class OMAFViewpointTransitionEffectBox extends Box(‘vpte’) { unsigned int(16) src_viewpoint_id; / / mandatory unsigned int(16) dst_viewpoint_id; / / mandatory unsigned int(8) transition_type; / / mandatory / / additional box to specify the parameters of the transition }
[0160] Examples of semantics for the fields of the OMAFViewpointTransitionBox are as follows: src_viewpoint_id is the id of the source viewpoint.
[0161] dst_viewpoint_id is the id of the destination viewpoint.
[0162] The transition_type is an integer that identifies the type of transition. A value of 0 indicates a basic transition. A value of 1 indicates a viewpoint path transition. A value of 2 indicates an auxiliary information transition. The remaining values are reserved for future transitions.
[0163] In some embodiments, additional boxes that are related to a particular type of transition and provide further information can be present in the OMAFViewpointTransitionBox. The additional boxes can be defined for each of the previously defined transition types. When the transition_type field of the OMAFViewpointTransitionBox is equal to 0, the OMAFBasicViewpointTransitionBox is present. This box contains only one field, namely, basic_transition_type, whose value indicates a specific transition from a predefined set of basic transitions. The OMAFPathViewpointTransitionBox is present when the transition_type field of the OMAFViewpointTransitionBox is equal to 1. This box contains a list of viewpoint identifiers that the player can follow when the user requests a transition to the target viewpoint. In some embodiments, the field can also be provided to indicate the speed of the transition along the path. The OMAFAuxiliaryInfoViewpointTransitionBox is present when the transition_type field of the OMAFViewpointTransitionBox is equal to 2. This box contains two fields, namely, a type field that specifies the nature of the transition (e.g., generating a virtual viewpoint, etc.) and an aux_track_id that provides a reference to one of the tracks in a file that contains the time-specification auxiliary information used to implement the transition effect. Examples of the syntax of the three aforementioned boxes are given below. aligned(8) class OMAFBasicViewpointTransitionBox extends Box('vptb') { unsigned int(8) basic_transition_type; } aligned(8) class OMAFPathViewpointTransitionBox extends Box('vptp') { unsigned int(16) intermediate_viewpoints[]; } aligned(8) class OMAFAuxiliaryInfoViewpointTransitionBox extends Box('vpta') { unsigned int(8) type; unsigned int(32) aux_track_id; }
[0164] Transport protocol level signaling of viewpoint transition effects (e.g., DASH) Viewpoint transition effect information signaled at the container level can also be signaled at the transport protocol level in the manifest file. If the container file contains viewpoint transition effect information, this information preferably matches the information signaled in the manifest file. In some embodiments, the viewpoint transition effects are signaled within an OVL descriptor, such as those described above. The transition effect between a pair of viewpoints can be signaled by the ovp:transition element. In one example, this element has three attributes: @src, @dst, and @type. These attributes specify the id of the source viewpoint, the id of the destination viewpoint, and the type of the transition effect, respectively. For some types of transition effects, the ovp:transition element can include child elements that provide additional information to be used by the client to render these transitions.
[0165] Table 4 lists examples of elements and attributes that can be used to signal a viewpoint transition effect in an MPD file.
[0166] [Table 4]
[0167] Signaling of the recommended projection format for the FoV Within different FoV ranges, different projection formats may be advantageous. For example, in a 90° field of view, a linear projection format can operate well, but when using a linear projection in a larger field of view such as 130°, an undesirable stretching effect may be visible. Conversely, projection formats such as a "planetoid" stereoscopic projection or a fisheye projection format may not operate well at a 90° FoV, but can present an appropriate rendering experience at higher FoV degrees.
[0168] In some embodiments, an OMAFRecommendedProjectionListBox is provided as additional metadata information in the "meta" box to signal the recommended projection format for a range of the device's field of view (FoV) values. This box contains one or more OMAFRecommendedProjectionBox instances. The OMAFRecommendedProjectionBox defines the horizontal and vertical FoV ranges and provides the projection types recommended for the specified FoV range. A player or streaming client that receives this signaling can determine the size of the field of view of the device on which the player or streaming client is operating (e.g., it can look up the device's FoV capabilities from a local database or obtain this characteristic via an API call to the HMD's operating system). The player or streaming client can then compare this determined field of view size with the FoV ranges defined in the OMAFRecommendedProjectionBoxes to determine which of the recommended projection types corresponds to the device's field of view. The player or streaming client can then request the content in the determined recommended projection format. Examples of the syntax for these boxes are provided below. Box Type: ’orpl’ Container: MetaBox Mandatory: No Quantity: Zero or one aligned(8) class OMAFRecommendedProjectionListBox extends Box(‘orpl’) { OMAFRecommendedProjectionBox recommendations[]; } Box Type: ’orpr’ Container: OMAFRecommendedProjectionListBox Mandatory: No Quantity: One or more aligned(8) class OMAFRecommendedProjectionBox extends Box('orpr') { bit(3) reserved = 0; unsigned int(5) projection_type; unsigned int(32) min_hor_fov; unsigned int(32) min_ver_fov; unsigned int(32) max_hor_fov; unsigned int(32) max_ver_fov; }
[0169] An example of the semantics of the OMAFRecommendedProjectionBox field is as follows: projection_type indicates the mapping type to the spherical coordinate system of the picture to be projected, as specified by the OMAF standard. The value of projection_type can be the index of a list of rendering projection methods, including orthographic projection, planetary projection, azimuthal equidistant projection, fisheye projection, etc.
[0170] min_hor_fov and min_ver_fov provide the minimum horizontal and vertical display fields of view in units of 2 - 16 degrees. min_hor_fov can range from 0 to 360×216, including both ends. min_ver_fov can range from 0 to 180×216, including both ends.
[0171] max_hor_fov and max_ver_fov provide the maximum horizontal and vertical display fields of view in units of 2 - 16 degrees. max_hor_fov can range from 0 to 360×216, including both ends. max_ver_fov can range from 0 to 180×216, including both ends.
[0172] For a specific FoV, when a projection format is recommended, min_hor_fov is equal to max_hor_fov and min_ver_fov is equal to max_ver_fov.
[0173] In another embodiment, the content author or content provider can provide information identifying the recommended viewport for devices with various FoV configurations having appropriate projection recommendations. Various devices with various FoVs can use the recommended projection format to render 360 - video content according to the recommended viewport.
[0174] OMAF describes the recommended viewport information box (RcvpInfoBox) as follows. class RcvpInfoBox extends FullBox('rvif',0,0) { unsigned int(8) viewport_type; string viewport_description; } viewport_type specifies the type of the recommended viewport as listed in Table 5.
[0175]
Table 5
[0176] In some embodiments, additional types of recommended viewports (which can be assigned to, for example, type 2, etc.) are used based on the FOV of the rendering device. In some embodiments, the viewport_description of the RcvpInfoBox can be used to indicate the recommended rendering projection method and the corresponding rendering FOV range. In some embodiments, an optional box is added to the RcvpInfoBox based on the viewport_type to indicate additional parameters used for the corresponding recommended type. For example, the OMAFRecommendedProjectionBox can be signaled when the viewport type is associated with the FOV. class RcvpInfoBox extends FullBox('rvif',0,0) { unsigned int(8) viewport_type; string viewport_description; Box[] other_boxes; / / optional }
[0177] In another embodiment, the recommended viewport can be adapted to multiple recommended types or subtypes to provide the user with a flexible choice. For example, viewing statistics can be further divided into statistics by measurement period (e.g., weekly, monthly), geography (country, city), or age (youth, adult). Table 6 shows a hierarchical recommendation structure that can be used in some embodiments.
[0178]
Table 6
[0179] In some embodiments, an inductive RcvpInfoBox structure is used to support the hierarchical recommendation structure. The other_boxes field proposed in the RcvpInfoBox structure can include RcvpInfoBoxes for specifying subtypes such as the following. class RcvpInfoBox extends FullBox('rvif',0,0) { unsigned int(8) viewport_type; string viewport_description; RcvpInfoBox(); / / optional ; }
[0180] A single directors cut recommended viewport can propose multiple tracks, each of which can support one or more recommended rendering projection methods for the FOV range. An exemplary structure of the RcvpInfoBox is shown below. The value of viewport_type for the first RcvpInfoBox is 0, indicating that such a recommended viewport is by directors cut, and the value of viewport_type in the second RcvpInfoBox (e.g., 1) can indicate that the track associated with this directors cut recommended viewport is recommended for a device with a specific rendering FOV. One or more instances of the OMAFRecommendedProjectionBox can be signaled to provide the recommended projection methods for the corresponding FOV range. RcvpInfoBox{ viewport_type = 0; / / recommended director’s cut RcvpInfoBox { Viewport_type=1; / / recommeded for device FOV OMAFRecommendedProjectionBox(); / / projection method 1 OMAFRecommendedProjectionBox(); / / projection method 2 viewport_description; } viewport_description; }
[0181] In DASH MPD, SupplementalProperty and / or EssentialProperty descriptors with @schemeIdUri equal to "urn:mpeg:dash:crd" can be used to provide a Content Recommendation Description (CRD). The @value of a SupplementalProperty or EssentialProperty element using the CRD scheme can be implemented as a comma-separated list of values for the CRD parameters, as shown in Table 7.
[0182]
Table 7
[0183] XML Schema for DASH Signaling Examples of XML schemas for DASH signaling that can be used in some embodiments are as follows. <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http: / / www.w3.org / 2001 / XMLSchema" targetnamespace="urn:mpeg:mpegI:omaf:2017" xmlns:omaf="urn:mpeg:mpegI:omaf:2017" elementformdefault="qualified"> <xs:element name="ovp" type="omaf:viewpointType" / > <xs:element name="ovg" type="omaf:viewpointGroupType" / > <xs:complextype name="viewpointType"> <xs:attribute name="id" type="xs:string" use="required" / > <xs:attribute name="effective_range" type="xs:unsignedInt" use="optional" / > <xs:attribute name="virtual" type="xs:boolean" use="optional" default="false" / > <xs:attribute name="synthesisMethod" type="xs:unsignedByte" use="optional" / > <xs:attribute name="refViewpointIds" type="xs:boolean" use="optional" / > <xs:attribute name="dynamicPosition" type="xs:boolean" use="optional" default="false" / > <xs:element name="position" type="omaf:viewpointPositionType" minOccurs="0" maxOccurs="1" / > <xs:element name="availability" type="omaf:viewpointAvailabilityType" maxOccurs="unbounded" / > <xs:element name="transition" type="omaf:vpTransitionType" minOccurs="0" maxOccurs="unbounded" / > < / xs:complextype> <xs:complextype name="viewpointPositionType"> <xs:element name="global" type="omaf:viewpointGlobalPositionType" maxOccurs="1" / > <xs:element name="cartesian" type="omaf:viewpointCartesianPositionType" maxOccurs="1" / > < / xs:complextype> <xs:complextype name="viewpointGlobalPositionType" use="optional" maxoccurs="1"> <xs:attribute name="longitude" type="xs:double" use="required" / > <xs:attribute name="latitude" type="xs:double" use="required" / > <xs:attribute name="altitude" type="xs:double" use="optional" default="0" / > < / xs:complextype> <xs:complextype name="viewpointCartesianPositionType" use="optional" maxoccurs="1"> <xs:attribute name="x" type="xs:int" use="required" / > <xs:attribute name="y" type="xs:int" use="required" / > <xs:attribute name="z" type="xs:int" use="optional" default="0" / > < / xs:complextype> <xs:complextype name="viewpointAvailabilityType" use="optional" maxoccurs="unbounded"> <xs:attribute name="start" type="xs:unsignedLong" use="required" / > <xs:attribute name="end" type="xs:unsignedLong" use="optional" / > < / xs:complextype> <xs:complextype name="vpTransitionType" use="optional" maxoccurs="unbounded"> <xs:attribute name="src" type="xs:string" use="required" / > <xs:attribute name="dst" type="xs:string" use="required" / > <xs:attribute name="type" type="xs:unsignedByte" use="required" / > <xs:element name="omaf:vpBasicTransitionType" use="optional" maxOccurs="1" / > <xs:element name="omaf:vpPathTransitionType" use="optional" maxOccurs="1" / > <xs:element name="omaf:vpAuxTransitionType" use="optional" maxOccurs="1" / > < / xs:complextype> <xs:complextype name="vpBasicTransitionType"> <xs:attribute name="type" type="unsignedByte" use="required" / > < / xs:complextype> <xs:complexttype name="vpPathTransitionType"> <xs:attribute name="viewpoints" type="xs:string" use="required" / > < / xs:complexttype> <xs:complextype name="vpAuxTransitionType"> <xs:attribute name="auxIdList" type="xs:string" use="required" / > < / xs:complextype> <xs:complextype name="viewpointGroupType"> <xs:attribute name="id" type="xs:string" use="required" / > <xs:attribute name="name" type="xs:string" use="optional" / > <xs:attribute name="longitude" type="xs:double" use="required" / > <xs:attribute name="latitude" type="xs:double" use="required" / > <xs:attribute name="viewpointIds" type="xs:string" use="required" / > < / xs:complextype> < / xs:schema>
[0184] Further Embodiments In some embodiments, the method includes receiving at least first 360-degree video data representing a view from a first perspective and second 360-degree video data representing a view from a second perspective, and generating a container file (e.g., an ISO-based media file format file) for at least the first video data and the second video data. In the container file, the first video data is organized into a first set of tracks, and the second video data is organized into a second set of tracks, where each track in the first set of tracks includes a first track group identifier associated with the first perspective, and each track in the second set of tracks includes a second track group identifier associated with the second perspective.
[0185] In some such embodiments, each track in the first set of tracks includes each instance of a perspective group type box that includes the first track group identifier, and each track in the second set of tracks includes each instance of a perspective group type box that includes the second track group identifier.
[0186] In some embodiments, where the container file is organized into a hierarchical box structure and the container file includes a view list box that identifies at least a first perspective information box and a second perspective information box, the first perspective information box includes at least (i) the first track group identifier and (ii) an indication of a time interval during which video is available from the first perspective, and the second perspective information box includes at least (i) the second track group identifier and (ii) an indication of a time interval during which video is available from the second perspective. The indication of the time interval can be a list of instances of an available interval box for each perspective.
[0187] In some embodiments where the container file is organized into a hierarchical box structure and the container file includes a view list box that identifies at least a first view information box and a second view information box, the first view information box includes at least (i) a first track group identifier, and (ii) an indication of the position of the first view, and the second view information box includes at least (i) a second track group identifier, and (ii) an indication of the position of the second view. The indication of the position can include Cartesian coordinates or latitude and longitude coordinates.
[0188] In some embodiments where the container file is organized into a hierarchical box structure and the container file includes a view list box that identifies at least a first view information box and a second view information box, the first view information box includes at least (i) a first track group identifier, and (ii) an indication of the valid range of the first view, and the second view information box includes at least (i) a second track group identifier, and (ii) an indication of the valid range of the second view.
[0189] In some embodiments where the container file is organized into a hierarchical box structure and the container file includes a transition effect list box that identifies at least one transition effect box, each transition effect box includes an identifier of a source view, an identifier of a destination view, and an identifier of a transition type. The identifier of the transition type can identify a basic transition or a view path transition. When the identifier of the transition type identifies a path view transition box, the path view transition box can include a list of view identifiers. When the identifier of the transition type identifies an auxiliary information view transition box, the auxiliary information view transition box can include a track identifier.
[0190] In some embodiments where the container file is organized into a hierarchical box structure that includes metaboxes, the metabox identifies at least one recommended projection list box, and each recommended projection list box includes information that identifies (i) a projection type and (ii) a corresponding field of view. The information that identifies the corresponding field of view can include a minimum horizontal field of view angle, a maximum horizontal field of view angle, a minimum vertical field of view angle, and a maximum vertical field of view angle.
[0191] Some embodiments include a non-transitory computer storage medium that stores a container file generated according to any of the methods described herein.
[0192] In some embodiments, the method includes receiving at least first 360-degree video data representing a view from a first viewpoint and second 360-degree video data representing a view from a second viewpoint, and generating a manifest such as an MPEG-DASH MPD. In the manifest, at least one stream in a first set of streams is identified, each stream in the first set represents at least a portion of the first video data, at least one stream in a second set of streams is identified, each stream in the second set represents at least a portion of the second video data, each stream in the first set is associated in the manifest with a first viewpoint identifier, and each stream in the second set is associated in the manifest with a second viewpoint identifier.
[0193] In some such embodiments, each stream in the first set is associated in the manifest with each conformant set that has a first viewpoint identifier as an attribute, and each stream in the second set is associated in the manifest with each conformant set that has a second viewpoint identifier as an attribute. The attribute can be an @viewpoint_id attribute.
[0194] In some embodiments, each stream in the first set is associated in the manifest with each conforming set having a first viewpoint identifier in a first descriptor, and each stream in the second set is associated in the manifest with each conforming set having a second viewpoint identifier in a second descriptor. The first and second descriptors can be SupplementalProperty descriptors.
[0195] In some embodiments, the manifest includes an attribute indicating the valid range of each viewpoint. In some embodiments, the manifest includes an attribute indicating the position for each viewpoint. The attribute indicating the position can include Cartesian coordinates or latitude and longitude coordinates. In some embodiments, the manifest includes information indicating at least one time period during which video is available for each viewpoint.
[0196] In some embodiments, the first video data and the second video data are received in a container file (such as an ISO base media file format file), in which case the first video data is compiled into a first set of tracks and the second video data is compiled into a second set of tracks, and each track in the first set of tracks includes a first track group identifier associated with the first viewpoint, and each track in the second set of tracks includes a second track group identifier associated with the second viewpoint. The viewpoint identifiers used in the manifest are equal to each track group identifier in the container file.
[0197] In some embodiments, the method comprises receiving a manifest that identifies a plurality of 360-degree video streams, the manifest including, for each identified stream, information that identifies the viewpoint position of each stream; obtaining and displaying a first video stream identified by the manifest; and overlaying, on the display of the first video stream, a user interface element that indicates the viewpoint position of a second video stream identified by the manifest. In some embodiments, the method includes obtaining and displaying the second video stream in response to a selection of the user interface element.
[0198] In some embodiments where the manifest further includes information that identifies at least one valid range of an identified stream, the method further includes displaying an indication of the valid range. In some embodiments where the manifest further includes information that identifies the available period of the second video stream, the user interface element is displayed only during the available period.
[0199] In some embodiments, the manifest includes information that identifies a transition type for a transition from the first video stream to the second video stream. In response to a selection of the user interface element, the method includes presenting a transition having the identified transition type and obtaining and displaying the second video stream, where the second video stream is displayed after the presentation of the transition.
[0200] In some embodiments where the manifest further includes information that identifies the position of at least one virtual viewpoint, the method further includes synthesizing a view from the virtual viewpoint in response to a selection of the virtual viewpoint and displaying the synthesized view.
[0201] In some embodiments, the method comprises receiving a manifest (MPEG-DASH MPD) that identifies a plurality of 360-degree video streams, the manifest including information that identifies each projection format of each video stream, the manifest further including information that identifies each range of field-of-view sizes for each of the projection formats, determining a field-of-view size for display, selecting at least one of the video streams such that the determined field-of-view size is included in the identified range of field-of-view sizes for the projection format of the selected video stream, and obtaining at least one of the selected video streams and displaying the obtained video stream using the determined field-of-view size.
[0202] Further embodiments include a system comprising a processor and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to perform any of the methods described herein.
[0203] Various hardware elements of one or more of the described embodiments are referred to as "modules," and it should be noted that each module performs (i.e., implements, executes, and the like) the various functions described herein in connection with that module. As used herein, a module can include hardware that one of ordinary skill in the art would consider appropriate for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices, etc.). Each described module can also include executable instructions for performing one or more of the functions described as being performed by that module, and these instructions can take the form of, or include, hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, or can be stored on one or more any suitable non-transitory computer-readable media commonly referred to as RAM, ROM, etc.
[0204] Although functions and elements have been described above in certain combinations, one of ordinary skill in the art will understand that each function or element may be used alone or in any combination with other functions and elements. Additionally, the methods described herein may be implemented by a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, magnetic media such as read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor may be used in connection with software to implement a radio frequency transceiver that may be used in a WTRU, UE, terminal, base station, RNC, or any host computer.< / maxoccurs> < / minoccurs>
Claims
1. Receiving, from a server, information identifying a group of viewpoints including information identifying a first group of viewpoints and a second group of other viewpoints; Receiving, from the server, information identifying one or more omnidirectional videos captured from respective viewpoints belonging to the first group and covering a first event occurring at a certain venue, and information identifying one or more omnidirectional videos captured from respective viewpoints belonging to the second group and covering a second event occurring at a different venue; Based on the received information identifying the group of viewpoints, transitioning from a first omnidirectional video of the one or more omnidirectional videos captured from respective viewpoints belonging to the first group to a second omnidirectional video of the one or more omnidirectional videos captured from respective viewpoints belonging to the second group; A method comprising the steps of.
2. The method according to claim 1, wherein the first event and the second event occur simultaneously.
3. For a group of viewpoints, the received information identifying the group of viewpoints The method according to claim 1, comprising an identification value identifying the group of viewpoints.
4. For a group of viewpoints, the received information identifying the group of viewpoints The method according to claim 1, comprising a location value identifying a location of an event covered by one or more omnidirectional videos captured by respective viewpoints of the group.
5. The method according to claim 4, wherein the location value includes a longitude value and a latitude value of the geolocation of the event.
6. For a group of viewpoints, the received information identifying the group of viewpoints The method according to claim 1, comprising a number value identifying the number of viewpoints in the group.
7. For a group of viewpoints, the received information identifying the group of viewpoints The method according to claim 1, comprising an array of identification values identifying respective viewpoints in the group.
8. For a group of viewpoints, the received information identifying the group of viewpoints The method according to claim 1, comprising a string identifying a name describing the group.
9. A system, comprising: At least one processor; When implemented by the at least one processor, the system receives from a server information identifying a group of viewpoints including information identifying a first group of viewpoints and a second group of other viewpoints, receives from the server information identifying one or more omnidirectional videos captured from each viewpoint belonging to the first group and covering a first event occurring at a certain venue, and information identifying one or more omnidirectional videos captured from each viewpoint belonging to the second group and covering a second event occurring at a different venue, transitions from a first omnidirectional video of the one or more omnidirectional videos captured from each viewpoint belonging to the first group to a second omnidirectional video of the one or more omnidirectional videos captured from each viewpoint belonging to the second group based on the received information identifying the group of viewpoints a memory storing instructions A system comprising. **Claim 10** The system according to claim 9, wherein the first event and the second event occur simultaneously. **Claim 11** For a group of viewpoints, the received information identifying the group of viewpoints The system according to claim 9, comprising an identification value identifying the group of viewpoints. **Claim 12** For a group of viewpoints, the received information identifying the group of viewpoints The system according to claim 9, comprising a location value identifying the location of an event covered by one or more omnidirectional videos captured by each viewpoint of the group. **Claim 13** The system according to claim 12, wherein the location value includes longitude and latitude values of the geolocation of the event. **Claim 14** For a group of viewpoints, the received information identifying the group of viewpoints The system according to claim 9, comprising a number value identifying the number of viewpoints in the group. **Claim 15** For a group of viewpoints, the received information identifying the group of viewpoints The system according to claim 9, comprising an array of identification values identifying each viewpoint in the group. **Claim 16** For a group of viewpoints, the received information identifying the group of viewpoints The system according to claim 9, comprising a string identifying a name describing the group.
17. receiving, by at least one processor, information identifying a group of viewpoints including information identifying a first group of viewpoints and a second group of other viewpoints from a server; receiving, from the server, information identifying one or more omnidirectional videos captured from respective viewpoints belonging to the first group and covering a first event occurring at a certain venue, and information identifying one or more omnidirectional videos captured from respective viewpoints belonging to the second group and covering a second event occurring at a different venue; transitioning, based on the received information identifying the group of viewpoints, from a first omnidirectional video of the one or more omnidirectional videos captured from respective viewpoints belonging to the first group to a second omnidirectional video of the one or more omnidirectional videos captured from respective viewpoints belonging to the second group; and a non-transitory computer-readable medium comprising instructions for implementing a method including the above steps.
Citation Information
Patent Citations
IEC14496-12
IEC23009-1
Media container file
JP2012505570A
Method of transmitting omnidirectional video, method of receiving omnidirectional video, device for transmitting omnidirectional video, and device for receiving omnidirectional video
US20180063505A1
Display device and information processing terminal device
WO2017159063A1