Metrics and messages to improve experience for 360-degree adaptive streaming

By exchanging information in a 360° video system and using dynamic adaptive streaming technology, viewport display and switching are optimized, solving the problems of high video size and high bandwidth requirements, and achieving low-latency rendering and large-scale delivery.

CN115766679BActive Publication Date: 2025-10-24INTERDIGITAL VC HOLDINGS INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202211176352.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-02-28
Filing Date
2018-03-23
Publication Date
2025-10-24
Estimated Expiration
2038-03-23

AI Technical Summary

Technical Problem

The high video size and high bandwidth requirements of 360° video result in excessive bandwidth consumption during delivery, making it difficult to achieve low-latency rendering and large-scale delivery.

Method used

By exchanging information between streaming clients, content origin servers, measurement/analysis servers, and other video streaming auxiliary network components, using Dynamic Adaptive Streaming (DASH) technology, the performance of different clients, players, devices, and networks can be monitored and compared, troubleshooting can be performed in real time or offline, and different viewports can be displayed in segments via HTTP to optimize video quality and latency.

Benefits of technology

It enables efficient display and switching of viewports under different network conditions, reduces viewport switching latency, improves user experience, reduces transmission bandwidth consumption, and supports low-latency rendering and large-scale delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766679B_ABST
    Figure CN115766679B_ABST
Patent Text Reader

Abstract

Provided is a method for receiving and displaying media content. The method can include requesting a set of DASH video segments associated with different viewports and qualities. The method can include displaying the DASH video segments. The method can include determining a latency metric based on a time difference between the display of a DASH video segment and one of: a device beginning to move, the device terminating movement, the device determining that the device has begun to move, the device determining that the device has stopped moving, or a display of a different DASH video segment. The different DASH video segment can be associated with one or more of a different quality or a different viewport.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese Invention Patent Application No. 201880031651.3, filed on March 23, 2018, entitled "Metrics and Messages to Improve Experience for 360-Degree Adaptive Streaming".

[0002] Cross Reference to Related Applications

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 475,563, filed on March 23, 2017, U.S. Provisional Patent Application No. 62 / 525,065, filed on June 26, 2017, and U.S. Provisional Patent Application No. 62 / 636,795, filed on February 28, 2018, the contents of which are incorporated herein by reference. BACKGROUND

[0004] 360° video is a fast growing format emerging in the media industry. 360° video can be enabled by the growing availability of virtual reality (VR) devices. 360° video can provide a whole new level of presence to the viewer. Compared to straight video (e.g., 2D or 3D), 360° video presents very difficult engineering challenges in video processing and / or delivery. Enabling a comfortable and / or immersive user experience can require very high video quality and / or very low latency. Large video size of 360° video can hinder delivering 360° video at scale in a quality manner.

[0005] 360° video applications and / or services can encode the entire 360° video into a standards-compliant stream in order to perform progressive download and / or adaptive streaming. By delivering the entire 360° video to the client, low latency rendering can be achieved (e.g., the client can have access to the entire 360° video content and / or can choose to render the portion it desires to view without further constraints). From the server's perspective, the same stream can support multiple users using possibly different viewports. The size of the video can be high, thereby incurring high transmission bandwidth when delivering the video (as an example, because the entire 360° video is encoded at high quality, e.g., 4K@60fps or 6K@90fps per eye). As an example, this high bandwidth consumption during delivery can not be effective since a user can only view a small portion of the entire picture (e.g., a viewport). SUMMARY

[0006] Systems, methods, and instrumentalities can be provided for exchanging information between one or more streaming clients, one or more content origin servers, one or more metrics / analysis servers, and / or other video streaming auxiliary network components.

[0007] For example, streaming clients can generate and report metrics that they support in a consistent manner, and various network components can generate and send streaming assistance messages to streaming clients or to each other. The performance of different clients, players, devices, and networks can be monitored and compared. Problems can be debugged in real-time or offline, and one or more faults can be isolated in real-time or offline.

[0008] For example, viewport view statistics can be reported from a client to a metrics server with dynamic adaptive streaming over hypertext transfer protocol (DASH) messages with server and network assistance. Information associated with a VR device can be reported. A streaming client can measure and report one or more latency parameters, such as a viewport switch latency, an initial latency, and a settling latency. One or more head mounted device (HMD) precision and sensitivity can be measured and exchanged between network components, such as origin servers, content delivery network servers, and clients or metrics servers. Initial rendering orientation information can be sent to a VR device with one or more DASH messages.

[0009] While dynamic adaptive streaming over hypertext transfer protocol (DASH) terminology is used herein to describe systems, methods, and instrumentalities, those skilled in the art will appreciate that the systems, methods, and instrumentalities apply equally to other streaming standards and implementations.

[0010] A device for receiving and displaying media content can display a first viewport. The first viewport can be composed of at least a portion of one or more dynamic adaptive streaming over HTTP (DASH) video segments. The first viewport can have a first quality. The device can determine a motion of the device at a first time. The motion of the device can include one or more of a movement of the device outside of the first viewport, a change in an orientation of the device, or a zoom operation. The device can detect the motion of the device when a change in one or more viewport switch parameters is equal to or greater than a threshold. As an example, the viewport switch parameters can include one or more of a center azimuth, a center elevation, a center tilt, an azimuth range, and an elevation range.

[0011] The device can display a second viewport. The second viewport can be composed of at least a portion of one or more DASH video segments at a second time. The second viewport can have a second quality that is less than the first quality. Alternatively, the second viewport can have a second quality that is equal to or greater than the first quality. The device can determine a viewport switch latency metric based on the first time and the second time, and can send the viewport switch latency metric.

[0012] The device can display a third viewport. The third viewport can be composed of at least a portion of one or more DASH video segments at a third time. The third viewport can have a third quality that is greater than the second quality. The third quality can be equal to or greater than the first quality, or the third quality can be equal to or greater than a quality threshold but less than the first quality. In one example, an absolute difference between the first quality and the third quality can be equal to or less than a quality threshold. Each DASH video segment can have a respective quality. The device can determine the first, second, or third quality based on the relevant portion of the one or more DASH video segments used to compose the respective viewport and the respective quality. For example, the device can determine the quality of a viewport using an average, a maximum, a minimum, or a weighted average of the respective quality of each of the one or more DASH video segments used to compose the viewport. The device can determine a quality viewport switch latency metric based on a difference between the first time and the third time, and can send the quality viewport switch latency metric.

[0013] The device can determine a viewport loss and whether the viewport loss affects at least one of the displayed first viewport, the second viewport, or the third viewport. The device can determine information associated with the viewport loss (e.g., based on a condition that the viewport loss affects at least one of the displayed first viewport, the second viewport, or the third viewport). The information associated with the viewport loss can include a viewport loss reason and a DASH video segment associated with the viewport loss. The device can send a viewport loss metric indicating the information associated with the viewport loss. The information associated with the viewport loss can include one or more of a time at which the viewport loss was determined, a source URL of a packet, and an error type of the viewport loss. The cause of the viewport loss can include one of a server error, a client error, a packet loss, a packet error, or a packet drop.

[0014] A device can request a set of DASH video segments associated with different viewports and qualities. The device can display the DASH video segments. The device can determine a latency metric based on a time difference between displaying a DASH video segment and one of: the device starting to move, the device ceasing to move, the device determining that the device has started to move, the device determining that the device has stopped moving, or displaying a different DASH video segment. The different DASH video segments can be associated with different qualities and / or different viewports.

[0015] A device for receiving and displaying media content can display a first viewport at a first time. The first viewport can include at least a portion of one or more DASH video segments. The first viewport can have a first quality. The device can detect that the device is in motion at a second time. The device can determine that the device has stopped moving at a third time while the device is associated with a second viewport. The device can display the second viewport at a fourth time. The second viewport can include at least a portion of one or more DASH video segments. The second viewport can be associated with a second quality. The device can display a third viewport at a fifth time. The third viewport can include at least a portion of one or more DASH video segments. The third viewport can be associated with a third quality that is greater than the second quality. The third quality can be equal to or greater than the first quality, or the third quality can be greater than a quality threshold but less than the first quality. The device can determine and transmit a latency metric based on a time difference between two or more of the first time, the second time, the third time, the fourth time, and the fifth time. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1A is a system diagram illustrating an example communications system in which one or more disclosed embodiments can be implemented;

[0017] Figure 1B is a system diagram illustrating an example RAN that can be used within the communications system 100 shown in Figure 1A FIG. 1C is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that can be used within the communications system 100 shown in FIG. 1 A according to an embodiment;

[0018] Figure 1C is a system diagram illustrating an example RAN that can be used within the communications system 100 shown in Figure 1A FIG. 1C is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that can be used within the communications system 100 shown in FIG. 1 A according to an embodiment;

[0019] Figure 1D is a system diagram illustrating an example RAN that can be used within the communications system 100 shown in Figure 1A FIG. 1C is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that can be used within the communications system 100 shown in FIG. 1 A according to an embodiment;

[0020] Figure 2An example portion of a 360° video displayed on a head mounted device (HMD) is described.

[0021] Figure 3 An example equirectangular projection for a 360° video is described.

[0022] Figure 4 An example 360° video mapping is described.

[0023] Figure 5 An example media presentation description (MPD) hierarchical data model is described.

[0024] Figure 6 An example of SAND messages exchanged between a network component and a client is shown.

[0025] Figure 7 An example of a VR metrics client reference model is shown.

[0026] Figure 8 An example of a 3D motion tracking range is shown.

[0027] Figure 9 An example of a 2D motion tracking intersection range with multiple sensors is shown.

[0028] Figure 10 An example of a viewport switching event is shown.

[0029] Figure 11 An example of a 360 video zoom operation is shown.

[0030] Figure 12 An example of weighted viewport quality for sub-picture scenarios is shown.

[0031] Figure 13 An example of weighted viewport quality for region-wise quality ranking (RWQR) encoding scenarios is shown.

[0032] Figure 14 An example of an equal quality viewport switching event is shown.

[0033] Figure 15 An example of a 360 video zoom example is shown.

[0034] Figure 16 An example of a latency interval is shown.

[0035] Figure 17 An example of a message flow of SAND messages between a DANE and a DASH client or between a DASH client and a metrics server is shown. DETAILED DESCRIPTION

[0036] Figure 1A is a diagram illustrating an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, broadcast, etc. to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tailing unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multi-carrier (FBMC).

[0037] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112. However, it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, any of the WTRUs 102a, 102b, 102c, 102d may be referred to as a “station” and / or “STA,” which may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, medical equipment and applications (e.g., remote surgery), industrial equipment and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated process chain environments), consumer electronic devices, and devices operating on commercial and / or industrial wireless networks, etc. The WTRUs 102a, 102b, 102c, 102d may be interchangeably referred to as UEs.

[0038] The communications system 100 can also include a base station 114a and / or a base station 114b. Each of the base stations 114a and / or 114b can be any type of device configured to wirelessly interface with one or more of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CN 106 / 115, the Internet 110, and / or the other networks 112. By way of example, the base stations 114a, 114b can be a base transceiver station (BTS), a Node-B, an eNode B, a Home Node B, a Home eNode B, a gNB, a NR Node B, a site controller, an access point (AP), a wireless router, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b can include any number of interconnected base stations and / or network elements.

[0039] The base station 114a can be part of the RAN 104 / 113, which can also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base stations 114a and / or 114b can be configured to transmit and / or receive wireless signals on one or more carrier frequencies of the RAN 104 / 113. These frequencies can be licensed or unlicensed frequencies, or a combination thereof. The cell can provide service to a particular geographical area that is relatively fixed or can change over time, depending on the location of the base stations 114a, 114b. The cell can further be divided into cell sectors. For example, the cell associated with the base stations 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, one for each sector of the cell. In an embodiment, the base station 114a can employ multiple-input multiple-output (MIMO) techniques. Thus, the base station 114a can utilize multiple transceivers for each sector of the cell. For example, the base station 114a can utilize beamforming to transmit and / or receive signals.

[0040] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).

[0041] More specifically, as noted above, the communications system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c can implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish the air interface 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0042] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro.

[0043] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as NR Radio Access, which can establish the air interface 116 using New Radio (NR).

[0044] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for instance using dual connectivity (DC) principles. Thus, the air interface utilized by WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., a eNB and a gNB).

[0045] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 IX, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.

[0046] Figure 1A The base station 114b in FIG. 10 can be a wireless router, Home Node B, Home eNode B, or access point, for example, and can utilize any suitable RAT for facilitating wireless connectivity access points employing the IEEE 802.11 specification family, e.g., 802.11a, 802. 1 lb, 802.1 lg, 802.11h, 802.1 lad, 802.1 laac, 802.1 lbac, and / or the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As shown, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b can not be required to access the Internet 110 via the CN 106 / 115. Figure 1A

[0047] The RAN 104 / 113 can be in communication with the CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have varying quality of service (QoS) requirements, such as differing throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 106 / 115 can provide call control, billing services, mobile location-based services, pre-paid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 10, a Figure 1A ​It is to be understood that the RAN 104 / 113 and / or the CN 106 / 115 can each include other elements that are not explicitly shown, as well. For example, the RAN 104 / 113 and / or the CN 106 / 115 can each also include relay nodes, backhaul links, load balancers, serving gateways, routing entities, and the like. It is to be understood that the RAN 104 / 113 and / or the CN 106 / 115 can be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the CN 106 / 115, the RAN 104 / 113 can also be

[0048] The CN 106 / 115 can also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or the other networks 112. The PSTN 108 can include circuit-switched telephone networks that provide infrastructure for the provision of voice telephony, facsimile, and / or other

[0049] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 can include multi-mode capabilities, e.g., the WTRUs 102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links. For example, the WTRU 102c shown in Figure 1 A can be configured to communicate with the base station 114a, which can employ a cellular-based radio technology, and with the base station 114b, which can employ an IEEE 802 radio technology. Figure 1A The WTRU 102c shown in Figure 1 A can be configured to communicate as a cellular telephone. The WTRU 102c can include a cellular transceiver 152, a modem 154, and a

[0050] Figure 1B Figure 1 B shows a system diagram of an example WTRU 102. As shown in Figure 1B The WTRU 102 can include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and other peripherals 138, among others. It will be appreciated that the WTRU 102 can include any sub-combination of the foregoing elements while remaining consistent with an embodiment.

[0051] The processor 118 can be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled Figure 1B The processor 118 and the transceiver 120 are depicted as separate components, however, it will be appreciated that the processor 118 and the transceiver 120 can be integrated in an electronic package or chip.

[0052] The transmit / receive element 122 can be configured to transmit signals to, or receive signals from, a base station (e.g., the base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 122 can be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and receive both RF and light signals. It will be appreciated that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0053] Although the transmit / receive element 122 is depicted in the Figure 1B WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0054] The transceiver 120 can be configured to modulate the signals that are to be transmitted by the transmit / receive element 122 and to demodulate the signals that are received by the transmit / receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. Thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11, for example.

[0055] The processor 118 of the WTRU 102 can be coupled to, and can receive user input data from, the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can access information from, and store data in, any appropriate memory, such as the non-removable memory 130 and / or the removable memory 132. The non-removable memory 130 can include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).

[0056] The processor 118 can receive power from the power source 134, and can be configured to distribute and / or control the power to the other components in the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 can include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.

[0057] The processor 118 can also be coupled to the GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 can receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on

[0058] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, frequency modulation (FM) radio units, digital music players, media players, video game console modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. The peripherals 138 may include one or more sensors, which may be one or more of the following: a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a contact sensor, a magnetometer, a barometer, a posture sensor, a biometric sensor, and / or a humidity sensor.

[0059] The WTRU 102 may include a full-duplex radio in which the reception or transmission of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. A full-duplex radio may include an interface management unit that reduces and / or substantially eliminates self-interference through signal processing by hardware (e.g., chokes) or by a processor (e.g., a separate processor (not shown) or by the processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio that transmits or receives some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)).

[0060] Figure 1C 1 is a system diagram illustrating the RAN 104 and the CN 106 according to one embodiment. As described above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c using an E-UTRA radio technology over the air interface 116. The RAN 104 may also be in communication with the CN 106.

[0061] The RAN 104 can include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 can include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c can implement MIMO technology. Thus, the eNode-B 160a, for example, can use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a.

[0062] Each of the eNode-Bs 160a, 160b, 160c can be associated with a particular cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, and the like. As shown, the eNode-Bs 160a, 160b, 160c can communicate with one another over an X2 interface. Figure 1C

[0063] Figure 1C The CN 106 can include the mobility management gateway (MME) 162, serving gateway (SGW) 164, and packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements are depicted as part of the CN 106, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.

[0064] The MME 162 can be connected to each of the eNode-Bs 160a, 160b, 160c in the RAN 104 via an S1 interface and can serve as a control node. For example, the MME 142 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 can also provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM or WCDMA.

[0065] ​The SGW 164 can be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 can generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 can perform other functions, such as anchoring user planes during inter-eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing contexts of the WTRUs 102a, 102b, 102c, and the like.

[0066] The SGW 164 can be connected to the PGW 166, which can provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0067] The CN 106 can also serve as a gateway for the WTRUs 102a, 102b, 102c to access the PSTN 108, the Internet 110, and / or the other networks 112. The PSTN 108 can include circuit-switched telephone networks that provide

[0068] While not shown in Figures 1A-1D WTRUs 102a, 102b, 102c are described as wireless terminals, it is contemplated that in certain representative embodiments such terminals can use (e.g., temporarily or permanently) wired communication interfaces with the communication network.

[0069] In representative embodiments, the other network 112 can be a WLAN.

[0070] A WLAN using an infrastructure Basic Service Set (BSS) mode can have an Access Point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an access or an interface to a Distribution System (DS) or a backhaul that carries traffic to and from the STAs. Traffic to STAs that is carried by the backhaul, but is not intended for the STAs (e.g., data packets that are a destination for an access terminal other than the STAs, or data packets that are to be passed through the STAs to another destination) can be adapted not to be transmitted from the access terminals to the AP. For example, the access terminals can be instructed not to transmit certain types of data packets (e.g., data packets that are not intended for the access terminals or data packets that are to be passed through the access terminals to another destination). In some embodiments, the AP can be a wireless station that also functions as a wireless access point. The STAs associated with the AP can include client stations, which can exchange data with the AP directly (e.g., using wireless links) and / or with other STAs via the AP (e.g., using wireless links and / or wired links). The AP can coordinate communications for the STAs associated with the AP. For example, the AP can coordinate the uplink and / or downlink communications for the STAs associated with the AP.

[0071] When using an 802.11 ac infrastructure mode of operation or similar mode of operation, the AP can transmit beacons on a fixed channel (e.g., a primary channel). The primary channel can have a fixed width (e.g., 20 MHz bandwidth) or a dynamically set width by signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish a connection with the AP. In some embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) with collision avoidance can be implemented (e.g., in 802.11 systems). For CSMA / CA, STAs (e.g., each of the STAs), including the AP, can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, the particular STA can back off. In a given BSS, only one STA (e.g., only one station) can transmit at any given time.

[0072] High Throughput (HT) STAs can use a 40 MHz wide channel for communication (e.g., by combining a 20 MHz wide primary channel with an adjacent or nonadjacent 20 MHz wide secondary channel).

[0073] Very High Throughput (VHT) STAs can support 20MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining 8 contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data can be passed through a segment parser that can divide the data into two streams. Inverse Fast Fourier Transform (IFFT) processing, and time domain processing, can be done on each stream separately. The streams can be mapped on to the two 80 MHz channels, and the data can be transmitted by a transmitting STA. At the receiver of the receiving STA, the above described operations can be done in reverse for the 80+80 configuration, and the combined data can be sent to the Medium Access Control (MAC).

[0074] 802.11af and 802.11ah support sub-1 GHz modes of operation. The channel operating bandwidths and carriers used in 802.11af and 802.11ah are reduced compared to 802.11η and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. In accordance with typical embodiments, 802.11ah can support meter type control / machine type communication (e.g., MTC devices in a macro coverage area). MTC devices can have certain capabilities, such as including restricted capabilities that support (e.g., only support) certain and / or limited bandwidths. MTC devices can include a battery, and the battery life of the battery is above a threshold (e.g., to preserve a long battery life).

[0075] For WLAN systems that can support multiple channels and channel bandwidths (e.g., 802.11η, 802.1 lac, 802.1 laf, and 802.1 lah), these systems include a channel that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a STA that is derived from all STAs operating in the BSS supporting the minimum bandwidth operating mode. In the example of 802.1 lah, even though the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes, the width of the primary channel can be 1 MHz for STAs (e.g., MTC type devices) that support (e.g., only support) 1 MHz mode. Carrier sensing and / or network allocation vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy (e.g., because a STA (which only supports 1 MHz operating mode) is transmitting to the AP), then the entire available frequency band can be considered busy even though most of the frequency band remains space and available for use.

[0076] In the United States, the available frequency bands for 802.1 lah are 902 MHz to 928 MHz. In Korea, the available frequency bands are 917.5 MHz to 923.5 MHz. In Japan, the available frequency bands are 916.5 MHz to 927.5 MHz. Depending on the country code, the total bandwidth available for 802.1 lah is 6 MHz to 26 MHz.

[0077] Figure 1D is a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As described above, the RAN 113 can be in communication with the WTRUs 102a, 102b, 102c over the air interface 116 using NR radio technologies. The RAN 113 can also be in communication with the CN 115.

[0078] The RAN 113 can include gNBs 180a, 180b, 180c, though it will be appreciated that the RAN 113 can include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, 180c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, 180c can implement MIMO technology. For example, gNBs 180a, 180b, 180c can utilize beamforming to transmit and / or receive wireless signals to and / or from the gNBs 180a, 180b, 180c. Thus, the gNB 180a, for example, can use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a. In an embodiment, the gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, the gNB 180a can transmit multiple component carriers (not shown) to the WTRU 102a. A subset of these component carriers can be on the licensed frequency spectrum while the remaining component carriers can be on the unlicensed frequency spectrum. In an embodiment, the gNBs 180a, 180b, 180c can implement Coordinated Multi-Point (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).

[0079] The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing different amounts of OFDM symbols and / or lasting different absolute lengths of time).

[0080] The gNBs 180a, 180b, 180c can be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In the standalone configuration, the WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c without also accessing other RANs (e.g., eNode-Bs 160a, 160b, 160c). In the standalone configuration, the WTRUs 102a, 102b, 102c can utilize one or more of gNBs 180a, 180b, 180c as a mobile anchor point. In the standalone configuration, the WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using signal(s) in an unlicensed band. In the non-standalone configuration, the WTRUs 102a, 102b, 102c can communicate / be connected with the gNBs 180a, 180b, 180c while also communicating with / being connected to another RAN (e.g., eNode-Bs 160a, 160b, 160c). For example, WTRUs 102a, 102b, 102c can implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c in a substantially simultaneous manner. In the non-standalone configuration, eNode-Bs 160a, 160b, 160c can serve as mobile anchor points for the WTRUs 102a, 102b, 102c and the gNBs 180a, 180b, 180c can provide additional coverage and / or throughput to the WTRUs 102a, 102b, 102c served by the gNBs 180a, 180b, 180c.

[0081] Each of the gNBs 180a, 180b, 180c can be associated with a certain cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support of network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards User Plane Function (UPF) 184a, 184b, routing of control plane information towards Access and Mobility Management Function (AMF) 182a, 182b and the like. As shown, the gNBs 180a, 180b, 180c can communicate with one another over an Xn interface. Figure 1D As shown, the gNBs 180a, 180b, 180c can be in communication with one another over an Xn interface.

[0082] Figure 1DThe illustrated CN 115 can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements are depicted as part of the CN 115, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.

[0083] The AMF 182a, 182b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and can serve as a control node. For example, the AMF 182a, 182b can be responsible for authenticating WTRUs 102a, 102b, 102c, supporting for network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing the WTRU 102a, 102b, 102c registration area, terminating NAS signaling, and mobility management, etc. The AMF 162 can utilize network slicing to tailor the CN support provided to the WTRU 102a, 102b, 102c based on the type of services being utilized by the WTRU 102a, 102b, 102c. For example, different network slices can be established for different use cases such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, and / or services for machine type communication (MTC) access, etc. The AMF 162 can provide control plane functionality for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.

[0084] The SMF 183a, 183b can be connected to AMF 182a, 182b in the CN 115 via an N11 interface. The SMF 183a, 183b can also be connected to UPF 184a, 184b in the CN 115 via an N4 interface. The SMF 183a, 183b can select and control the UPF 184a, 184b and configure the routing of traffic through the UPF 184a, 184b. The SMF 183a, 183b can perform other functions, such as managing and allocating IP address, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and the like. The PDU session type can be IP-based, non-IP based, Ethernet-based, and the like.

[0085] The UPF 184a, 184b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which can provide the WTRUs 102a, 102b, 102c with access to packet- switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPF 184, 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering of downlink packets, and the like.

[0086] The CN 115 can facilitate communications with other networks. For example, the CN 115 can include, or can communicate with, an IP gateway for facilitating communications between the CN 115 and the PSTN 108, or other networks. The CN 115 can also provide the WTRUs 102a, 102b, 102c with access to the other networks, which can include other wired and / or wireless networks that are owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c can be connected to a local DN 185a, 185b through the UPF 184a, 184b via the N3 interface and an N6 interface between the UPF 184a, 184b and the DN 185a, 185b.

[0087] In view of the Figures 1A-1D and the corresponding description of Figures 1A-1D one or more or all of the functions described with reference to one or more of the accompanying drawings can be executed by one or more simulation devices (not shown): WTRU 102a-d, base station 114a-b, eNodeB 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN 185a-b and / or any other device described herein. These simulation devices can be one or more devices configured to simulate one or more functions described herein. For example, these simulation devices can be used to test other devices and / or to simulate a network and / or WTRU functionality.

[0088] The one or more emulation devices can perform one or more or all of the functions while being implemented / deployed in a laboratory about, and / or a carrier network. For example, the one or more emulation devices can perform one or more or all of the functions while being implemented / deployed as part of a wired and / or wireless communication network with or without partial or full deployment. The one or more emulation devices can perform one or more or all of the functions while being implemented / deployed as part of a wired and / or wireless communication network with full or partial deployment. The emulation devices can be coupled directly to the other devices to perform the testing, and / or can perform the testing using over-the-air wireless communications.

[0089] The one or more emulation devices can perform one or more or all of the functions while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices can be used in a testing laboratory and / or a testing scenario that is not a deployed (e.g., testing) wired and / or wireless communication network to implement tests on one or more components. The one or more emulation devices can be test equipment. The emulation devices can transmit and / or receive data using direct RF coupling and / or wireless communications via RF circuitry (which can include one or more antennas, for example).

[0090] Figure 2 An example portion of a 360° video displayed on a head mounted device (HMD) is depicted. As an example, as shown in Figure 2 As an example, as shown in FIG. 1, a portion of a 360° video can be presented to a user when viewing the 360° video. The portion of the video can change as the user looks around and / or zooms in or out of the video image. The portion of the video can change based on feedback provided by the HMD and / or other types of user interfaces (e.g., wireless transmit / receive unit (WTRU)). The spatial region of the entire 360° video can be referred to as a viewport. The viewport can be presented to the user in whole or in part. The viewport can have one or more different qualities than other portions of the 360° video.

[0091] A 360° video can be captured and / or rendered on a sphere (e.g., to enable a user to select an arbitrary viewport). Sphere video formats can not be directly deliverable with regular video codecs. A 360° video (e.g., a sphere video) can be compressed by projecting the sphere video onto a 2D plane using a projection method. The projected 2D video can be coded (e.g., using regular video codecs). Examples of projection methods can include equirectangular projection. Figure 3An equirectangular projection for 360° video is described. As an example, an equirectangular projection method can use one or more of the following equations to map a first point P having coordinates (θ, φ) on a sphere to a second point P having coordinates (u, v) on a two-dimensional plane,

[0092] u = φ / (2*pi) + 0.5 Equation 1

[0093] v = 0.5 - θ / (pi) Equation 2

[0094] Figure 3 One or more examples related to 360° video mapping are described. As an example, a 360 video can be converted to a 2D plane video (as an alternative, for example, to reduce bandwidth requirements) by using one or more other projection methods, such as mapping. For example, the one or more other projection methods can include a pyramid map, a cube map, and / or an offset cube map. By using the one or more other projection methods, a spherical video can be presented with less data. For example, a cube map projection method can use 20% less pixels than an equirectangular projection (as an example, because the cube map projection introduces relatively less warping). The one or more other projection methods, such as a pyramid projection, can perform sub-sampling on one or more pixels that are less likely to be viewed by a user (as an example, thereby reducing the size of the 2D video being projected).

[0095] Viewport-specific representations can be used as an alternative approach for storing and / or delivering an entire 360 video between a server and a client. Figure 3 The one or more projection methods shown, such as a cube map and / or a pyramid projection, can provide a non-uniform quality representation for different viewports (for example, certain viewports can be represented with higher quality than other viewports). It can be necessary to generate and / or store multiple versions of the same video with different target viewports at the server side (for example, to support all viewports of a spherical video). For example, in an implementation of VR video delivery at Facebook, multiple versions of the same video were used with different target viewports, Figure 3, an offset cubemap format is shown. The offset cubemap can provide the highest resolution (e.g., highest quality) for the front viewport, the lowest resolution (e.g., lowest quality) for the rear viewport, and a medium resolution (e.g., medium quality) for one or more side views. The server can store multiple versions of the same content (e.g., to accommodate client requests for different viewports of the same content). As an example, the same content can have a total of 150 different versions (e.g., 30 viewports multiplied by 5 resolutions per viewport). During delivery (e.g., streaming), a client can request a specific version corresponding to the client's current viewport. The specific version can be delivered by the server. A client requesting a specific version corresponding to the client's current viewport can save transmission bandwidth, possibly including increased storage requirements on the server, and / or potentially increase latency when / when the client changes from a first viewport to a second viewport. This latency issue can be severe when such viewport changes frequently.

[0096] The client device can receive a video stream encoded in accordance with the Omnidirectional Media Format (OMAF). A sub-picture can be a picture that represents a spatial subset of the original content. A sub-picture bitstream can be a bitstream that represents a spatial subset of the original content. A viewing orientation can be a triplet of azimuth, elevation, and tilt angles that characterizes the orientation of a user when consuming audio-visual content. A viewport can be an area of ​​one or more omnidirectional images and / or one or more videos suitable for display and for viewing by a user. A viewpoint can be the center point of a viewport. Content coverage can include one or more spherical areas covered by the content represented by a track or image item.

[0097] OMAF can specify viewport-independent video profiles and viewport-dependent video profiles. For viewport-independent video streaming, the 360 ​​video picture can be encoded into a single-layer bitstream. The entire encoded bitstream can be stored on the server and, if necessary, can be fully transmitted to the OMAF player and fully decoded by the decoder. The area of ​​the decoded picture corresponding to the current viewport can be rendered to the user. For viewport-dependent video streaming, a video streaming method or a combination of video streaming methods can be used. The OMAF player can receive video encoded according to a single-layer stream-based method. In a single-layer stream-based method, multiple single-layer streams can be generated. As an example, each stream can contain the entire omnidirectional video, but each stream can have different high-quality encoded areas indicated by regional quality ranking (RWQR) metadata. Depending on the current viewport, a stream containing a high-quality encoded area that matches the current viewport position can be selected and / or can be transmitted to the OMAF player.

[0098] An OMAF player can receive video encoded according to a sub-picture stream based approach. In the sub-picture stream based approach, a 360 video can be partitioned into sub-picture sequences. Each sub-picture sequence can cover a subset of the spatial region of the omnidirectional video content. Each sub-picture sequence can be encoded into a single layer bitstream in a mutually independent manner. The OMAF player can select one or more sub-pictures to be performed streaming (e.g., based on the orientation / viewport metadata of the OMAF player). The stream for the current viewport can be received, decoded, and / or rendered with better quality or higher resolution compared to the quality or resolution covering the remaining and currently unrendered region indicated by the region-wise quality ranking (RWQR) metadata.

[0099] HTTP streaming has also become a dominant approach in commercial deployments. For example, streaming platforms such as Apple's HTTP Live Streaming (HLS), Microsoft's Smooth Streaming (SS), and / or Adobe's HTTP Dynamic Streaming (HDS) can use HTTP streaming as the underlying delivery method. HTTP streaming standards for multimedia content can enable a standard-based client to stream content from any standard-based server (as an example, thereby enabling interoperability between servers and clients of different vendors). Dynamic adaptive streaming over HTTP (MPEG-DASH) can be a versatile delivery format that can provide the best possible video experience to end users by dynamically adapting to changing network conditions. DASH can be built on top of the HTTP / TCP / IP stack. DASH can define a manifest format, a media presentation description (MPD), and a segment format for the ISO base media file format as well as MPEG-2 transport streams.

[0100] Dynamic HTTP streaming can require that alternative versions of a multimedia content at different bit rates can be provided on a server. The multimedia content can include several media components (e.g., audio, video, text), some or all of which can have different characteristics. In MPEG-DASH, these characteristics can be described by the MPD.

[0101] The MPD can be an XML document that contains the metadata needed by a DASH client to construct appropriate HTTP-URLs to access video segments in an adaptive manner during a streaming session (related examples are described herein). Figure 5An example MPD hierarchical data model is described. The MPD can describe a series of Periods, where in a Period, the consistent set of encoded versions of media content components does not change. A (e.g., each) Period can have a start time and duration. A (e.g., each) Period can include one or more Adaptation Sets (e.g., Adaptation Set). As described herein with respect to Figure 8 A DASH streaming client can be a WTRU, as described.

[0102] An "Adaptation Set" can represent a set of encoded versions of one or several media content components that share one or more same attributes (e.g., language, media type, picture aspect ratio, role, accessibility, viewpoint, and / or rating attributes). A first Adaptation Set can include different bitrates of a video component of the same multimedia content. A second Adaptation Set can include different bitrates (e.g., lower quality stereo and / or higher quality surround sound) of an audio component of the same multimedia content. A (e.g., each) Adaptation Set can contain multiple Representations.

[0103] A "Representation" can describe a deliverable encoded version of one or several media components, whereby it can be distinguished from other Representations by bit rate, resolution, number of channels, and / or other characteristics. A (e.g., each) Representation can include one or more Segments. One or more attributes (e.g., @id, @bandwidth, @quality ranking, and @dependency Id) of a Representation element can be used to specify one or more attributes of the associated Representation.

[0104] A Segment can be the largest unit of data that can be retrieved with a single HTTP request. A (e.g., each) Segment can have a URL (e.g., an addressable location on a server). Segments (e.g., each Segment) can be downloaded using HTTP GET or HTTP GET with byte ranges.

[0105] A DASH client can parse the MPD XML document. The DASH client can select a selection of adaptation set (e.g., based on information provided in one or more adaptation set elements) that is suitable for the DASH client environment. Within one (e.g., each) adaptation set, the client can select a representation. The client can select the representation based on the value of the @bandwidth attribute, the client decoding capability, and / or the client rendering capability. The client can download the initialization segment of the selected representation. The client can access the content (e.g., by requesting the entire segment or a byte range of the segment). After the presentation starts, the client can continue to consume the media content. For example, the client can request (e.g., continuously request) media segments and / or media segment portions during the presentation. The client can play out the content according to the media presentation timeline. The client can switch from a first representation to a second representation based on updated information from the client environment. The client can play out the content continuously for two or more periods. When the client consumes the media contained in the segment up to the end of the media announced in the representation, the media presentation can terminate, a period can start, and / or the MPD can be (e.g., need to be) re-fetched.

[0106] Messages between a streaming client (e.g., a DASH client) and network components or between different network components can provide information about the real-time working characteristics of the network, servers, proxies, caches, CDNs, and the performance and status of the DASH client. Figure 6 An example of server and network assisted DASH (SAND) messages exchanged between network components and one or more clients is shown. As shown, parameter enhanced delivery (PED) messages can be exchanged between DASH aware network components (e.g., 602 or 604) or DASH assisted network components (DANE) (e.g., 602 or 604). Parameter enhanced reception (PER) messages can be sent from one or more DANE 602 and 604 to DASH clients 606. Status messages can be sent from one or more DASH clients 606 to one or more DANE 602 and 604. The status messages can provide real-time feedback from the DASH clients 606 to the one or more DANE 602 and 604 to support real-time operations. Metrics messages can be sent from one or more DASH clients 606 to a metrics server 608. The metrics messages can provide session summaries or longer time intervals.

[0107] DASH clients can generate and report supported metrics in a consistent manner. The performance of different clients, players, devices, and networks can be monitored and compared. Problems can be debugged in real time or offline, and one or more faults can be isolated in real time or offline. Entities that can generate, report, process, and visualize the metrics share a common understanding of the information contained in the messages, metrics, and reports.

[0108] When a DASH client streams omnidirectional media, some DASH metrics can be used in the same way as when the client streams traditional media. Omnidirectional media can be associated with higher bitrate requirements, higher rendering latency, and higher latency sensitivity. Certain metrics are available for use by DASH clients streaming this media type. For DASH entities that support SAND, special PED / PER messages can be used.

[0109] The VR device can measure its accuracy and sensitivity. The VR device can include one or more HMDs. Initial rendering orientation information can be sent to the VR device. Viewport viewing statistics can be reported (e.g., by the VR device). Regions of interest (ROIs) can be signaled at the file format level. ROIs can be signaled as events carried with (e.g., each) segment. Signaling for prefetching data in DANE components can be at the MPD level along with supplementary PED messages. ROI information can be carried in a timed metadata track or in an event stream element.

[0110] As an example, Figure 6 As shown, the VR Metrics Client Reference Model can be used with different types of information exchanged between one or more of a streaming client, an origin server, or a measurement / analysis server and other DANEs.

[0111] Semantics can be defined using an abstract syntax. Items in the abstract syntax can have one or more primitive types, including integers, real numbers, Boolean values, enumerations, strings and / or vectors, etc. Items in the abstract syntax can have one or more composite types. The composite types can include objects and / or lists, etc. Objects can include unordered sequences of key values ​​and / or value pairs. Key values ​​can have string type. The key values ​​can be unique in the sequence. Lists can include ordered lists of items. Multiple (e.g., two) timestamps can be used or defined. The types of timestamps can include real time (e.g., wall clock time) with Real-Time type and / or media time with Media-Time type.

[0112] The VR Metrics Client Reference Model may be used for metrics collection and / or processing. The VR Metrics Client Reference Model may use or define one or more Observation Points (OPs) to perform metrics collection and processing. Figure 7 This is an example of a VR metric client reference model. A VR device may include a VR client 718. A VR device or a VR player may be used based on the VR metric client reference model 700.

[0113] Figure 7 The functional components of the VR metric client reference model 700 are shown. These functional components may include one or more of the following: network access 706, media processing 708, sensors 710, media presentation 712, VR client control and management 714, or metric collection and processing (MCP) 716. VR applications can communicate with different functional components to provide an immersive user experience. Network access 706 can request VR content from network 704. Network access 706 can request one or more viewport segments with different qualities to perform viewport-dependent streaming (e.g., based on dynamic sensor data). Media processing component 708 can decode (e.g., each) media sample(s) and / or can pass one or more decoded media samples to media presentation 712. Media presentation 712 can present the corresponding VR content using one or more devices (e.g., one or more HMDs, speakers, or headphones).

[0114] The sensor 710 can detect the position and / or movement of the user. The sensor 710 can generate sensor data (e.g., based on the position and / or movement of the user). The VR client 718 and / or server can use the sensor data in various ways. For example, the sensor data can be used to determine a viewport segment to request (e.g., a specific high-quality viewport segment). The sensor data can be used to determine a portion or portions of a current media sample (e.g., audio and video) to be presented on a device (e.g., one or more HMDs and / or headphones). The sensor data can be used to apply post-processing to match the media sample presentation to the user movement (e.g., recent user movement).

[0115] The VR client control and / or management module 700 can configure the VR client 718 for the VR application. The VR client control and / or management module 700 can interact with some or all functional components. The VR client control and / or management module 700 can manage the workflow. The workflow can include an end-to-end media processing process. As an example, for the VR client 718, the workflow can include processing from receiving media data to presenting media samples.

[0116] The MCP 716 can aggregate data (e.g., data from various observation points (OP)). The data can include one or more of: user location, event time, and / or media sampling parameters, among others. The MCP 716 can derive one or more metrics used by or sent to the metrics server 702. For example, a viewport switch latency metric derivation can be based on input data from OP3 (e.g., motion trigger time) and / or OP4 (e.g., viewport display update time).

[0117] The metrics server 702 can collect data (e.g., metrics data). The metrics server 702 can use the data to perform one or more of: user behavior analysis, VR device performance comparison, or other debugging or tracking purposes. The entire immersive VR experience across networks, platforms, and / or devices will be enhanced.

[0118] A VR service application (e.g., a CDN server, a content provider, and / or a cloud gaming session) can use the metrics collected by the metrics server 702 to deliver different quality levels and / or immersive experiences to VR clients (e.g., VR devices that include the VR client 718) (e.g., based on client capabilities and performance). For example, a VR cloud gaming server can check performance metrics of a VR client, including one or more of: end-to-end latency, tracking sensitivity and accuracy, or field of view, among others. The VR cloud gaming server can determine whether the VR client 718 can provide an appropriate level of VR experience to a user. The VR cloud gaming server can authorize one or more eligible VR clients to access a VR game. The one or more eligible VR clients can provide an appropriate level of VR experience to a user.

[0119] The VR cloud gaming server can not authorize one or more unqualified VR clients to access a VR game. The one or more unqualified VR clients can not be able to provide a proper level of VR experience for the user. For example, if the VR cloud gaming server determines that a client (e.g., VR client 718) is experiencing a high end-to-end latency and / or does not have sufficient tracking sensitivity parameters, the VR cloud gaming server can not allow the VR client to join a live VR game session. The VR cloud gaming server can permit the client (e.g., VR client 718) to enter a regular game session. For example, the client can join a game that only employs a 2D version that does not have an immersive experience. The VR cloud gaming server can monitor dynamic performance changes of the client. When the VR metrics show a degradation in performance in real-time, the VR cloud gaming server can switch to a lower quality VR game content or a less immersive VR game. For example, if the VR cloud gaming server determines that the client has a relatively high quality viewport switch latency, the VR cloud gaming server can deliver tiles (e.g., tiles of a similar quality) to the client in order to reduce the quality disparity. The VR experience of the VR user can be affected by network conditions. A content delivery network (CDN) or edge server can push VR content and / or services to a cache that is closer to the client (e.g., VR user).

[0120] The MCP 716 can aggregate data from different observation points (OPs) in various ways. Exemplary OPs (OP1-OP5) are provided herein, however it should be understood that the MCP 716 can aggregate data from any order or combination of the OPs described herein as examples. The MCP 716 can be part of a client or reside in a metrics server (e.g., metrics server 702).

[0121] OP1 can correspond to or include an interface from the network access 706 to the MCP 716. The network access 706 can issue one or more network requests, receive VR media streams, and / or decapsulate VR media streams from the network 704. OP1 can contain a set of network connections. One (e.g., each) network connection can be defined by one or more of: a destination address, a time of initiation, a time of connection, or a time of closure. OP1 can include a series of transport network requests. One or more (e.g., each) transport network request can be defined by one or more of: a time of transport, content, or a TCP connection used to send the one or more network requests. For each network response, OP1 can contain one or more of: a time of receipt, response header content, and / or a time of receipt of each response body byte.

[0122] OP2 can correspond to or include an interface from media processing 708 to MCP 716. Media processing 708 can perform de-multiplexing and / or decoding (e.g., audio, image, video, and / or light field decoding). OP2 can include encoded media samples. An encoded media sample (e.g., each encoded media sample) can include one or more of: media type, media codec, media decode time, omnidirectional media metadata, omnidirectional media projection, omnidirectional media region packing, omnidirectional video viewport, frame packing, color space or dynamic range, etc.

[0123] OP3 can correspond to an interface from sensors 710 to MCP 716. Sensors 710 can acquire a user's head or body orientation and / or movement. Sensors 710 can acquire environmental data, such as light, temperature, magnetic field, gravity, and / or biometrics. OP3 can include a list of sensor data, where the sensor data includes one or more of: 6DoF (X, Y, Z, yaw, pitch, and roll), depth, or velocity, etc.

[0124] OP4 can correspond to an interface from media presentation 712 to MCP 716. The media presentation 712 can synchronize and / or present mixed natural and synthetic VR media elements (e.g., to provide a fully immersive VR experience to a user). Media presentation 712 can perform color conversion, projection, media composition, and / or view composition for (e.g., each) VR media element. OP4 can include decoded media samples. A sample (e.g., each media sample) can include one or more of: media type, media sample presentation timestamp, wall clock counter, actual rendering viewport, actual rendering time, actual playout frame rate, or audio-video synchronization, etc.

[0125] OP5 can correspond to an interface from VR client control and management 714 to MCP 716. VR client control and management component 714 can manage client parameters, such as display resolution, frame rate, field of view (FOV), eye-to-screen distance, lens separation distance, etc. OP5 can contain VR client configuration parameters. A parameter (e.g., each parameter) can include one or more of: display resolution, display density (e.g., in PPI), horizontal and vertical FOV (e.g., in degrees), tracking range (e.g., in millimeters), tracking accuracy (e.g., in millimeters), motion prediction (e.g., in milliseconds), media codec support, or OS support, etc.

[0126] If the user (e.g., a VR user) has been granted permission, the VR client control and management 714 can collect personal profile information of the user (e.g., in addition to the device information). The personal profile information can include one or more of the following: gender, age, location, race, or religion, etc. of the VR user. The VR service application can collect the personal profile information from the metrics server 702. The VR service application can collect (e.g., directly collect) the personal profile information from the VR client 718 (including a VR device as an example). The VR service application can identify VR content and / or services that are suitable for the VR user (e.g., a personal VR user).

[0127] A VR streaming service provider can deliver different sets of 360-degree video tiles to different users. The VR streaming service provider can deliver different sets of 360-degree video tiles based on characteristics of different potential viewers. For example, an adult can receive a full 360-degree video that includes access to restricted content. A minor user can view the same video but cannot access the restricted content (e.g., regions and / or tiles). The VR service provider can stop delivering certain VR content for a location and / or at a time (e.g., for a particular location and / or at a particular time) in order to comply with laws and regulations (e.g., laws or regulations of a local government). A VR cloud gaming server can provide different levels of immersion for games to users (e.g., men, women, elderly, and / or young) with different simulated motion sickness sensitivities, thereby avoiding potential simulated motion sickness. A VR device can control local VR rendering (e.g., by using personal profile information). For example, a VR service provider can multicast / broadcast a 360-degree video that includes one or more restricted scenes (e.g., in a non-discriminatory manner to all). A VR receiver (e.g., each VR receiver) can configure rendering based on personal profile information. A VR receiver can grant different levels of access to restricted content (e.g., one or more scenes) based on personal profile information. For example, a VR receiver can filter out or blur some or all restricted scenes (one or more) based on age, race, or religion information of a potential viewer. A VR receiver can determine whether to allow access to some or all restricted scenes (one or more) based on age, race, or religion information of a potential viewer. The methods and / or techniques herein can allow a VR service provider to provide VR content (e.g., a single VR content version) to more users and / or in an efficient manner.

[0128] The MCP 716 can obtain metrics from one or more OPs and / or derive VR metrics (e.g., a particular VR metric, such as latency).

[0129] A device (e.g., a VR device, a HMD, a phone, a tablet, or a personal computer) can generate and / or report viewing-related metrics and messages. The viewing-related metrics and messages can include one or more of the following: viewport view, viewport match, rendering device, initial rendering orientation, viewport switch latency, quality viewport switch delay (e.g., quality viewport switch delay), initial latency, settling latency, 6DoF coordinates, gaze data, frame rate, viewport loss, precision, sensitivity, and / or ROI, etc.

[0130] A device can use a metric to indicate which viewport(s) was / were requested and / or viewed at what time(s). The metric can include a viewport view metric. Which viewport(s) was / were requested and viewed at what time(s) can be indicated by a report on the metric (e.g., viewport view). A device or a metrics server can derive a viewport view metric (as shown in Table 1) from OP4. Table 1 shows an example of a report on viewport view.

[0131]

[0132] Table 1

[0133] An entry (e.g., entry "source") can include (e.g., specify) the origin of the original VR media sample. A device can infer the entry "source" from one or more of the following: the URL of the MPD, the media element representation for DASH streaming, or the initial URL of the media sample file for progressive download.

[0134] An entry (e.g., entry "timestamp") can include (e.g., specify) the media sample presentation time.

[0135] An entry (e.g., entry "duration") can specify a time interval for a client to report how long the client stayed in the corresponding viewport.

[0136] An entry (e.g., entry "viewport") can include (e.g., specify) the region of the omnidirectional media presented at media time t. The region can be defined by one or more of the following: center yaw, center pitch, static horizontal range, static vertical range (if present), or horizontal range and vertical range. For 6DoF, the entry "viewport" can include the user position or additional user position (x, y, z).

[0137] An entry (e.g., entry "viewport") can present a watched viewport using one or more annotations or expressions used in different specifications and implementations. For example, different users' viewport view metrics can be compared and / or correlated with each other by a metric server (e.g., metric server 702). As an example, a metric server can determine a more popular viewport. A network (e.g., CDN) can use this information (e.g., entry "viewport") to perform caching or for content generation. A device (e.g., VR device) can record one or more viewport views with certain precision or precisions in order to be played back. The recorded viewport views can be used later (e.g., by another viewer) to replay a particular version of a 360 video that was watched earlier. A network component (e.g., broadcast / multicast server) or one or more VR clients can synchronize viewports watched by multiple viewers based on viewport view metrics. Viewers can have similar or identical experiences (e.g., in real time).

[0138] A device (e.g., client) can log viewport view metrics (e.g., in a periodic manner) and / or can be triggered by one or more viewing orientation changes (e.g., detected by a sensor) to report viewport view metrics.

[0139] A device can use a metric to indicate whether a client follows a recommended viewport (e.g., director cut). The metric can include a viewport match metric. A device (e.g., client) can compare a client's viewing orientation to corresponding recommended viewport metadata and / or log a match event if the client viewing orientation matches the corresponding recommended viewport metadata. Table 2 shows an example of what can be reported in a report about a viewport match metric.

[0140]

[0141] Table 2

[0142] One entry (e.g., entry "IDENTIFIED") can be a Boolean parameter that specifies whether the viewport (e.g., the current viewport) matches the recommended viewport during the measurement interval. When the value of entry "IDENTIFIED" is true, any of the following entries or a combination thereof can be presented. One entry (e.g., entry "t- trackld") can be an identifier of the recommended viewport timing metadata sequence and / or can be used to identify which of the one or more recommended viewports is matched. One entry (e.g., entry "COUNT") can specify the cumulative number of matches of the recommended viewport specified by "track." One entry (e.g., entry "DURATION") can specify the time interval of consecutive viewport matches during the measurement interval. One entry (e.g., entry "TYPE") can specify the type of the matched viewport, where a value of 0 can indicate that the recommended viewport is a director's cut and / or a value of 1 can indicate that the recommended viewport is selected based on viewing statistics. The device (e.g., the client) can record the viewport matches (e.g., in a periodic manner) and / or can be triggered to report the viewport matches when the viewing orientation changes or the recommended viewport position and / or size changes.

[0143] The device can use the metrics to indicate the rendering device metrics and messages. The metrics can include the rendering device. As an example, a VR device type report about the metrics rendering device can indicate one or more of: brand, model, operating system (OS), resolution, density, refresh rate, codec, projection, packaging, tracking range, tracking accuracy, tracking sensitivity, and / or other device-specific information about the device (e.g., location of rendering omnidirectional media). The metrics server can determine based on the report whether the content is consumed on a TV or on an HMD or on a daydream-type HMD. The metrics server or client can associate the information with the content type for better analysis. If the VR device is an HMD, information such as HMD-specific supported information can be included in the metrics. The client (e.g., VR client) can determine the appropriate device information to report to the metrics server. The metrics server can request the appropriate information from the VR device. The HMD-specific supported information can include, but is not limited to, accuracy, sensitivity, maximum supported frame rate, maximum supported resolution, and / or a list of supported codecs / projection methods. Table 3 shows an example of a report about the VR device (e.g., metrics rendering device). The metrics server can associate information such as in Table 3 with the content type for analysis. For example, the client or metrics server can derive the metrics rendering device and / or one or more entries in Table 3 from OP5.

[0144]

[0145] Table 3

[0146] An entry (e.g., entry "Tracking Range") can include (e.g., specify) a maximum extent or maximum size of a 2D or 3D tracking range supported by the VR device for VR user motion detection. The maximum extent or maximum size of the 2D or 3D tracking range can mean that motion tracking and detection can be inaccurate outside the maximum extent or maximum size of the 2D or 3D tracking range. A minimum extent or minimum size can be a point. As an example, the tracking range can be defined as the width, length, and height (e.g., in millimeters) of the VR space. Figure 8 An example of a measured 3D tracking range is shown. The entry tracking range can be defined as the width 802, height 804, and length 806 of the 3D space 800, or as a circular range for a 2D range with (e.g., a single) tracking sensor. The entry "Tracking Range" can be defined as the width, length of a 2D intersection area with multiple tracking sensors. Figure 9 An example of a 2D motion tracking intersection range with multiple sensors (e.g., tracking sensors 904-910) is shown. As shown, the length and width of the 2D intersection area 902 can define the entry "Tracking Range". Figure 9 An example of a 2D motion tracking intersection range with multiple sensors (e.g., tracking sensors 904-910) is shown. As shown, the length and width of the 2D intersection area 902 can define the entry "Tracking Range".

[0147] An entry (e.g., entry "Tracking Precision") can include (e.g., specify) the tracking resolution unit (e.g., minimum tracking resolution unit) in millimeters for translational motion (x, y, z) and in millidegrees for rotational motion (yaw, pitch, roll).

[0148] An entry (e.g., entry "Tracking Sensitivity") can include and can be (e.g., specified) the minimum motion displacement (e.g., fine motion) (Δχ, Ay, Δζ) in millimeters for translation and in millidegrees (Δ yaw, Δ pitch, Δ roll) for rotation that the VR device detects.

[0149] An entry (e.g., entry "Tracking Latency") can indicate (e.g., specify) the time interval between a motion and the detection of the motion by the sensor and the reporting of the motion to the application.

[0150] An entry (e.g., entry "Rendering Latency") can indicate (e.g., specify) the time interval between a media sample being output from the rendering module and its presentation on the display and / or headphones.

[0151] A device (e.g., a VR device) can generate and / or report an initial rendering orientation message. A device can use a metric to indicate the orientation from which to start playback. The metric can include an initial rendering orientation. An initial rendering orientation report on the metric "initial rendering orientation" can indicate from which orientation playback should start. An initial rendering orientation report on the metric "initial rendering orientation" can indicate from which orientation playback should start. The initial orientation can be indicated at the beginning of the presentation and / or at certain points in the presentation (e.g., scene changes). The content generator can provide the indication. One or more initial rendering orientation messages can have a PER type. For example, the one or more initial rendering orientation messages can be sent from a DANE to a DASH client. The information (e.g., initial rendering orientation) can be carried in a timed metadata track, an event stream element, or a SAND-PER message. A metric server can use the initial rendering orientation metric to configure the HMD orientation to reset one or more coordinates of the rendering position for VR content (e.g., at scene changes). With the initial rendering orientation message, the HMD can move another video region to the front. Table 4 shows an example content for the initial rendering orientation message. The metric server or client can derive the metric and one or more entries contained in Table 4 from OP2.

[0152]

[0153] Table 4

[0154] A device can generate a metric to indicate the latency experienced by a client when switching from one viewport to another. The metric can include "viewport switch latency." For example, one or more viewport switch latency metric reports on the metric "viewport switch latency" can indicate the latency experienced by a client when switching from one viewport to another. The viewport switch latency metric can be used by a network or a client (e.g., a VR device) to measure the quality of experience (QoE) of a viewer. A (SAND) DANE component and / or origin server can use the viewport switch latency metric for one or more of: content encoding, packaging, caching, or delivery. Table 5 shows an example report for the viewport switch latency metric. The metric server or client can derive the viewport switch latency metric and / or one or more entries shown in Table 5 from OP3 and / or OP4.

[0155]

[0156] Table 5

[0157] An entry (e.g., entry "pan") can include (e.g., specify) a user's pan motion (e.g., forward / backward, up / down, left / right, and / or pan motion that can be represented by a displacement vector (Ax, Ay, Az)).

[0158] An entry (e.g., entry "rotation") can include (e.g., specify) a user's angular motion, e.g., yaw, pitch, roll, and / or angular motion that can be represented by an angular displacement (Ayaw, Apitch, Aroll) in radian or degree units.

[0159] An entry (e.g., entry "first viewport") can include (e.g., specify) a center coordinate of a first viewport and a size of the first viewport.

[0160] An entry (e.g., entry "second viewport") can indicate (e.g., specify) a center point coordinate(s) of a second viewport and a dimension of the second viewport.

[0161] An entry (e.g., entry "latency") can include (e.g., specify) a delay between an inputted translational and / or rotational motion and a corresponding viewport of one or more relevant VR media elements (video, audio, image, light field, etc.) that is updated or displayed on a VR presentation. For example, entry "latency" can specify a time interval between a time when a head motion from a first viewport to a second viewport occurs (e.g., a time when a client detects the head motion) and a time when a corresponding second viewport is presented on a display.

[0162] For example, a viewport switch latency can correspond to how quickly an HMD responds to a head movement, unpacks, and renders content for a shifted viewport. A viewport switch latency can correspond to a time between when a viewer's head movement is detected by an HMD and when the viewer's head movement affects what the viewer observes in a video. For example, a viewport switch latency can be measured between when a head moves from a current viewport to a different viewport and when the different viewport is displayed on an HMD. As an example, a viewport switch latency can be caused by a network or a client's insufficient processing power to retrieve and / or render a corresponding viewport. After receiving and analyzing one or more viewport switch latency metric reports along a timeline, a network and / or a client player can use different intelligent methods to minimize a viewport switch latency.

[0163] A device (e.g., an HMD, a phone, a tablet, or a personal computer) can determine a movement of the device (e.g., a head movement). The movement of the device can include one or more of the following: a movement of the device outside of a current viewport, a change in orientation, or a zoom operation, etc. The device can determine a viewport switch latency based on the movement of the device.

[0164] Figure 10 An example is shown regarding a viewport switch event. The device can determine a device movement based on the viewport switch event. A viewport can be represented by a set of parameters. As Figure 10As shown, the first viewport 1002 can be represented by a first azimuth range 1010, a first elevation range 1012, and a first center azimuth and / or a first center elevation 1006. The second viewport 1004 can be represented by a second azimuth range 1014, a second elevation range 1016, and a second center azimuth and / or a second center elevation 1008.

[0165] If the change in one or more parameters in the set of parameters is equal to or greater than a threshold, the device can determine a viewport switch event. If the change in one or more parameters in the set of parameters is equal to or greater than a threshold, the device can determine that the device moved. Upon determining a viewport switch event, the device can apply one or more constraints to maintain consistent measurements across different applications and / or devices.

[0166] As an example, the device can determine that the first viewpoint 1002 associated with the first rendering viewport is stable prior to the viewport switch based on one or more parameters associated with the first viewport 1002 (the first azimuth range 1010, the first elevation range 1012, and the first center azimuth and / or the first center elevation 1006). Likewise, as an example, the device can determine that the second viewpoint 1004 associated with the second rendering viewport is stable after the viewport switch event based on one or more parameters associated with the second viewport (the second azimuth range 1014, the second elevation range 1016, and the second center azimuth and / or the second center elevation 1008). As an example, as shown in FIG. 10, the device can determine that the first viewpoint 1002 is stable prior to the viewport switch event and the second viewpoint 1004 is stable after the viewport switch event based on the center azimuth, the center elevation, and / or the center tilt angle of the first viewpoint 1002 and the second viewpoint 1004. Figure 10 As shown, a viewpoint can be represented by one or more of the following: a center azimuth, a center elevation, and / or a center tilt angle. The center azimuth and the center elevation can indicate the center of the viewport. The center tilt angle can indicate the tilt angle of the viewport (e.g., in units of 2 -16 degrees relative to a global coordinate axis). As an example, the center azimuth, the center elevation, and / or the center tilt angle of the first viewpoint can not change by more than m units of 2 -16 degrees within n milliseconds prior to the viewport switch event, and the center azimuth, the center elevation, and / or the center tilt angle of the second viewpoint can not change by more than m units of 2 -16 degrees within n milliseconds after the viewport switch event. The m and the n can both be positive integers. The device can be configured with the m and the n. As an example, the device can receive the values of m and n (e.g., from a content server or from a measurement server). The device can use the m and the n for identifying a viewport switch event to log measurements and / or report the measurements.

[0167] A viewport switch event can be triggered when a viewpoint (e.g., a viewpoint for a viewer) moves outside of a current rendered or displayed viewport of a device. A device can determine that the device moves when a viewpoint (e.g., a viewpoint for a viewer) moves outside of a current rendered or displayed viewport. As an example, as shown in FIG. 10, a first viewport 1002 or a first rendered viewport can be represented by a center position 1006 (e.g., a first center azimuth, a first center elevation, a first center tilt) of the first viewport 1002 or the first rendered viewport and / or a first viewport or a first rendered viewport region (e.g., a first azimuth range 1010 and a first elevation range 1012). A second viewport 1004 or a second rendered viewport can be represented by a center position 1008 (e.g., a second center azimuth, a second center elevation, a second center tilt) of the second viewport 1004 or the second rendered viewport and / or a second viewport or a second rendered viewport region (e.g., a second azimuth range 1014 and a second elevation range 1016). Figure 10

[0168] A device can determine that a viewport switch event occurs when a parameter change (e.g., a difference between a parameter of the first viewport 1002 and a parameter of the second viewport 1004) is equal to or greater than a threshold. As an example, a viewport switch event can occur when a distance between the first center azimuth 1006 and the second center azimuth 1008 is equal to or greater than a threshold (e.g., the first azimuth range / 2). A viewport switch event can occur when a distance between the first center elevation 1006 and the second center elevation 1008 is equal to or greater than a threshold (e.g., the first elevation range / 2).

[0169] A viewport switch event can be triggered by a device detecting a viewing orientation change. A viewing orientation can be defined by a triple of azimuth, elevation, and tilt angles that characterize an orientation in which a user consumes audiovisual content. A viewing orientation can be defined by a triple of azimuth, elevation, and tilt angles that characterize a viewport orientation in which a user consumes an image or a video. As an example, a device can determine that the device moves when the device detects a viewing orientation change. As an example, a viewing orientation can be associated with a relative position or location of a viewport.

[0170] A viewport switch event can occur during one or more zoom operations. As an example, a device can determine that the device moves during one or more zoom operations. As an example, a zoom operation can occur even if a center position of a rendered or displayed viewport does not change, but a size of a corresponding viewport and / or a corresponding spherical region changes. Figure 11 ​An example of zooming a 360 video viewport is shown. For one or more zoom operations, the first viewport and second viewport parameters of the viewport switching delay can be replaced by the first viewport size 1102 and the second viewport size 1108, respectively. The parameters azimuth range and / or elevation range can specify one or more ranges (e.g., in 2-degrees) around the center point of the viewport to be rendered. -16 degrees). When the difference between the first viewport size 1102 and the second viewport size 1108 is equal to or greater than a threshold value, the device may determine that a viewport switching event has occurred. The first viewport size 1102 and the second viewport size 1108 may be determined by the azimuth range 1104 and / or elevation range 1106 for the first viewport and the azimuth range 1110 and / or elevation range 1112 for the second viewport. The device may separately determine the first viewport size 1102 before the viewport switching event and the second viewport size 1108 after the viewport switching after the first and second viewports are determined to be stable. For example, when the first viewport size 1102 does not change by more than m times of 2 within n milliseconds before the viewport switching event -16 The device can determine that the first viewport is stable when the second viewport size 1108 does not change by more than m times 2 within n milliseconds after the viewport switch event. -16 The device determines whether the secondary viewport is stable when the device determines that the secondary viewport is stable.

[0171] A device may generate a metric to indicate the delay or time it takes to restore the quality of a displayed viewport to a quality threshold after a change in the device's movement. For example, after the device moves from a first viewport to a second viewport, the device may determine a quality viewport switching delay metric based on the time interval between the first viewport and the second viewport (e.g., when the quality of the second viewport reaches a quality threshold). As an example, when the quality of the second viewport reaches the quality of the first viewport or reaches a quality that is better and / or higher than the quality of the first viewport, the quality viewport switching delay metric may be presented as an equal-quality viewport switching delay metric. The equal-quality viewport switching delay metric may be a subset of the quality viewport switching delay metric. For a quality viewport switching delay metric, the quality of the second viewport may or may not reach the quality of the first viewport. The quality of the second viewport may be less than the quality of the first viewport but greater than the quality threshold. The device may be configured to have a quality threshold or to receive a quality threshold (e.g., from a metric server). The device may determine or specify the quality threshold and / or report the quality threshold to the metric server.

[0172] The quality threshold can be defined in terms of a difference between a plurality of (e.g., two) qualities. As an example, the quality threshold can be reached when a difference (e.g., an absolute difference or a difference in absolute values between qualities or quality rankings) between a first viewport quality and a second viewport quality is equal to or less than a threshold value.

[0173] The device can generate a metric to indicate a latency experienced by the device to switch from a current viewport to a different viewport until the different viewport reaches a same or similar rendering quality as the current viewport. For example, the device can determine an equal quality viewport switch latency metric. One or more equal quality viewport switch latency metric reports regarding the metric "EQ Viewport Switch Latency" can indicate a latency experienced by the device to switch from a current viewport to a different viewport until the different viewport reaches a same or higher quality as the current viewport or until the different viewport reaches a quality that exceeds a quality threshold but is less than the quality of the current viewport. As an example, the metric server can use the equal quality viewport switch latency metric to measure the QoE of a viewer. Table 6 shows an example report regarding the equal quality viewport switch latency metric. The device or the metric server can derive the equal quality viewport switch latency metric and / or one or more entries shown in Table 6 from OP3 and / or OP4.

[0174] Table 6 shows an example of a report regarding the equal quality viewport switch latency metric.

[0175]

[0176] Table 6

[0177] For example, when a user moves their head to view viewport B, the user can be viewing viewport A at a certain quality (e.g., high quality (HQ)). The quality of viewport B can be at a different quality (e.g., at low quality (LQ) because of viewport adaptive streaming). Depending on one or more of the following, it can take some time to ramp up the quality of viewport B from LQ to HQ: motion detection, viewport request scheduling, and one or more network conditions. Before the quality of viewport B is ramped up from LQ to HQ, the user can move their head to viewport C. A quality viewport switch latency metric can correspond to the time between switching from presenting viewport A at a first quality to presenting viewport B and / or viewport C on a display (e.g., an HMD display or a regular display) at a quality that is higher than a quality threshold. An equal quality viewport switch latency metric can correspond to the time between switching from presenting viewport A at a first quality to presenting viewport B and / or C on a display (e.g., an HMD display or a regular display) at a quality that is the same as or higher than the first quality of viewport A. The quality threshold can be less than the first quality of viewport A. The quality threshold can be reached when the difference (e.g., absolute difference or difference between temporal quality or absolute value of quality ordering) between the first quality of viewport A and the quality of viewport B and / or viewport C is equal to or less than a threshold value.

[0178] An entry (e.g., entry "pan") of the metric "EQ viewport switch latency" can include (e.g., specify) HMD position pan in a number (e.g., 3) of degrees of freedom (forward / backward, up / down, left / right, and / or can be represented by one or more vectors (e.g., constant vectors (Δx, Δy, Δz)). An entry (e.g., entry "rotation") can include (e.g., specify) HMD circular motion in a number (e.g., 3) of degrees of freedom (e.g., yaw, pitch, roll, and / or can be represented by one or more vectors (e.g., constant vectors (Δyaw, Δpitch, Δroll)). An entry (e.g., entry "quality ordering") can include (e.g., specify) the quality of the viewport rendered to the user after (e.g., at the end point) the viewport switch. The quality of the viewport can indicate a rendering quality. The higher the quality ordering value, the lower the rendering quality. An entry (e.g., entry "latency") can include (e.g., specify) the delay between the inputted pan and / or rotation motion and the quality of the corresponding viewport updated on the VR presentation.

[0179] A viewport can be covered by multiple regions. The viewport can include one or more portions of one or more DASH video segments. For example, a viewport can be covered by multiple regions based on multiple DASH video segments. A region can include one or more DASH video segments or one or more portions of one or more DASH video segments. A viewport can be defined by quality. As an example, a region (e.g., each region) can be defined by quality as indicated by a quality ordering value based on one or more region-wise quality ordering (RWQR). The region can be an independently coded sub-picture (e.g., video segment) and / or a region-wise quality ordering region of a single layer representation. A device or metric server can derive a viewport quality (e.g., as indicated by a quality ordering value) as an average of region qualities. A quality of a region can be determined based on a quality ordering value of the region or a minimum or maximum quality ordering value among quality ordering values of regions that constitute the viewport. As an example, a quality of a viewport can be determined as a weighted average of qualities (indicated by quality ordering values) of each respective region or video segment that constitutes the viewport. The weights can correspond to an area, size, and / or portion of the viewport covered by each respective region. A quality of a viewport can be indicated by a quality level, a quality value, and / or a quality ordering.

[0180] Figure 12 is an example of a sub-picture based streaming. A sub-picture can correspond to one or more DASH video segments or one or more portions of a DASH video segment. In Figure 12 , a viewport can include 4 sub-pictures, where each sub-picture has its own quality (e.g., as indicated by a quality ordering value). For example, a quality ordering value of sub-picture A 1202 can be 1, a quality ordering value of sub-picture B 1204 can be 3, a quality ordering value of sub-picture C 1206 can be 2, and a quality ordering value of sub-picture D 1208 can be 5. A coverage of a sub-picture can be expressed in terms of a percentage of a viewport area covered by the respective sub-picture. As shown in Figure 12 , 70% of a viewport area can be covered by sub-picture A 1202, 10% of a viewport area can be covered by sub-picture B 1204, 15% of a viewport area can be covered by sub-picture C 1206, and 5% of a viewport area can be covered by sub-picture D 1208. As an example, a viewport quality can be calculated using Equation 3. Quality (viewport) can represent a quality ordering value of a viewport. Quality (sub-picture(n)) can represent a quality ordering value of sub-picture n. Coverage (sub-picture(n)) can represent a percentage of an area of sub-picture n in a viewport.

[0181]

[0182] Equation 4 provides an example of a quality ordering value of a viewport derived using Equation 3.Figure 12 Examples of viewport quality ordering values.

[0183] Quality (viewport) = Quality (sub-picture A) * Coverage (sub-picture A) + Quality (sub-picture B) * Coverage (sub-picture B) + Quality (sub-picture C) * Coverage (sub-picture C) + Quality (sub-picture D) * Coverage (sub-picture D) = 1 * 0.7 + 3 * 0.1 + 2 * 0.15 + 5 * 0.05 = 1.55 Equation 4

[0184] Figure 13 is an example of a region-level quality ranking (RQR) encoding scenario in which a viewport is simultaneously covered by a high quality region (e.g., region A 1306) and a low quality region (e.g., region B 1304). The region quality can be indicated by a corresponding RWQR value. The RWQR value for region A 1306 can be 1. The RWQR value for region B 1304 can be 5. Region A can cover 90% of the viewport and region B 1304 can cover 10% of the viewport. As an example, the viewport quality ordering value can be calculated using Equation 5.

[0185] Quality (viewport) = Quality (region A) * Coverage (region A) + Quality (region B) * Coverage (region B) = 1 * 0.9 + 5 * 0.1 = 1.4 Equation 5

[0186] A quality viewport switching event can be identified when the quality of a second viewport rendered by the device is higher than a threshold quality. For example, a quality viewport switching event can be identified when the quality ordering value of the second viewport being rendered is equal to or less than the quality ordering value of the first viewport being rendered prior to the switch. A quality viewport switching event can be identified when the quality ordering value of the second viewport being rendered is less than the quality ordering value of the first viewport prior to the viewport switching event.

[0187] An EQ viewport switching latency can be used to define a latency associated with an equal quality viewport switching event for a change in orientation. The EQ viewport switching latency can be a time interval between a time at which a device (e.g., a sensor of the device) detects a change in user orientation from a first viewport to a second viewport (as an example, the viewport switch here can be identified as an equal quality switching viewport) and a time at which the second viewport is fully rendered to the user at the same or similar quality as the first viewport.

[0188] Figure 14An example of an equal quality viewport switching event is shown. A device can receive and decode multiple sub-pictures (e.g., DASH video segments), e.g., sub-pictures A-F. Sub-pictures A-F can correspond to one or more respective video segments and respective qualities (e.g., indicated by one or more quality ranking values). The device can use sub-pictures B and C to produce viewport #1. The device can render (e.g., display) viewport #1 for a user on a display of the device. The device or a metrics server can derive a quality ranking value for viewport #1 from one or more quality ranking values of sub-pictures B and C. In a viewport adaptation streaming scenario, the quality ranking values of sub-pictures B and C can be (e.g., typically will be) lower than the quality ranking values of sub-pictures A, D, E, and F. As Figure 14 (b) shows that when the device moves and the user viewing orientation moves from viewport #1 to viewport #2, viewport #2 is covered by sub-pictures B and C. Sub-pictures B and C for viewport #2 can be the same or different than the sub-pictures B and C used to create viewport #1. Viewport #2 can have the same quality as the quality of viewport #1, or viewport #2 can have a different quality than the quality of viewport #1. The device or a metrics server can derive a quality for viewport #2 (e.g., as indicated by a quality ranking value) from one or more quality ranking values of sub-pictures B and C.

[0189] As Figure 14 (c) shows that when the device moves and causes the user viewing orientation to move from viewport #2 to viewport #3, viewport #3 is covered by sub-pictures C and D. Sub-pictures C and D used to generate viewport #3 can be the same or different than the sub-pictures C and D requested by the client when it rendered viewport #1 and / or viewport #2. Sub-pictures C and D were not used to generate viewport #1 and viewport #2. Sub-pictures C and D can be requested to handle a change (e.g., a sudden viewing orientation change). The device or a metrics server can derive a quality for viewport #3 from one or more quality rankings of sub-pictures C and D. The device can compare the quality of viewport #3 to the quality of viewport #1, viewport #2, and / or a quality threshold, where the quality threshold is greater than the quality of viewport #2 but less than the quality of viewport #1. For example, if the device determines that the quality of viewport #3 is less than the quality of viewport #1 or less than the threshold, the device can request a different (e.g., new) representation of sub-picture D for a better quality. As an example, the better quality can be indicated by a relatively lower quality ranking value than the original quality ranking value of sub-picture D. As an example, upon receiving and using the different (e.g., new) representation of sub-picture D to produce viewport #3, if the device determines that the quality of viewport #3 is greater than the quality threshold or equal to or greater than the quality of viewport #1 (e.g., an equal quality viewport switching event), the device can determine and / or report a quality viewport switching event.

[0190] Figure 15An example is shown for 360 scaling (e.g., for measuring isometric viewport switch latency during a scaling operation). The device can request and receive a plurality of sub-pictures, e.g., sub-pictures A, B, C, D, E, and F. As shown at 1510, the device can use sub-pictures B and C to generate a first viewport. Sub-pictures B and C can be high quality sub-pictures. As shown at 1520, the device can receive an input to scale and can generate and display a second viewport in response. The second viewport can include a spherical region spanning sub-pictures A, B, C, and D. Sub-pictures B and C can have high quality, and sub-pictures A and D can have low quality. As such, the second viewport can be defined by a lower quality than the first viewport. As shown at 1530, the device can request a different (e.g., new) representation of sub-pictures A and D with higher quality (e.g., lower quality ranking value), switch sub-pictures A and D from low quality to high quality, and generate a full second viewport with high quality as shown at 1540. Figure 15 (c) shown in (b).

[0191] As an example, depending on which of the following is used to determine an isometric viewport switch event, the device can generate an EQ viewport switch latency metric based on Figure 15 (a) the time interval between Figure 15 (c) or Figure 15 (b) the time interval between Figure 15 (c). The EQ viewport switch latency can be measured as the time (e.g., in milliseconds) from when the sensor detects a user viewing orientation in the position of the second viewport and the time when the second viewport content is fully rendered to the user at a quality that is equal to or higher than the quality at which the first viewport was previously rendered. The EQ viewport switch latency can be measured as the time (e.g., in milliseconds) from when the sensor detects a user viewing orientation in the position of the second viewport and the time when the second viewport content is fully rendered to the user at a higher quality than the quality of the second viewport when the first viewport was rendered.

[0192] The metric computation and reporting module can use the signaled representation bitrate and / or resolution in the DASH MPD to determine the relative quality of the respective representation (as an example, rather than the RWQR value). The device (e.g., the metric computation and reporting module) can derive the viewport quality from the bitrate and / or resolution of the representation (e.g., each representation) that is requested, decoded, or rendered to cover the viewport. In the absence of a representation, the quality of the representation can be derived as the highest quality ranking value by default.

[0193] A device (e.g., a VR device) can generate and report an initial latency metric and message. The device can use a metric to indicate a time difference between starting head motion and starting corresponding feedback in the VR domain. The metric can include an initial latency metric (e.g., when using an HMD). The initial latency can have a delay dependent on the HMD (e.g., based on sensors, etc.) compared to viewport switch latency. Table 7 shows an example of reporting on the initial latency metric. The device or metric server can derive the initial latency metric and / or one or more entries shown in Table 7 from OP3 and / or OP4.

[0194] Table 7 shows an example of reporting on the initial latency metric.

[0195]

[0196] Table 7

[0197] A device can use a metric to indicate a time difference between stopping head motion and stopping corresponding feedback in the VR domain. The metric can include a settling latency metric (e.g., when using an HMD). Table 8 shows an example of reporting on the settling latency metric. The device or metric server can derive the settling latency metric and / or one or more entries shown in Table 8 from OP3 and / or OP4.

[0198]

[0199] Table 8

[0200] As an example, a device (e.g., a VR device) can report latency metrics separately for each viewer's head motion and / or viewport changes. As an example, the device can calculate and / or report an average over a predetermined period (e.g., at certain intervals). A metric server can measure latency performance of different VR devices based on the latency metrics. The metric server can correlate the latency metrics with other data including one or more of: content type, total viewing time, and frequency of viewport changes to determine the extent to which one or more VR device-specific characteristics affect user satisfaction. As an example, the metric server can use the initial and settling latencies to perform benchmarking of HMDs and / or provide feedback to vendors, as examples.

[0201] A device (e.g., an HMD, a phone, a tablet, or a personal computer) can request one or more DASH video segments and display the DASH video segments as various viewports with different qualities. The device can determine different types of latency and can report the latency as one or more latency metrics. The device can determine the one or more latency metrics based on a time difference between displaying a current viewport and a different time point. The time point can include one or more of: the device starting to move, the device terminating to move, the processor determining that the device has started to move, the processor determining that the device has stopped to move, or displaying a different viewport and / or a different quality.

[0202] Figure 16Example latency intervals that can be reported by a DASH client (e.g., a VR device) are shown. At time to, the user's head / device (e.g., HMD) can start moving from viewport #1 to viewport #2. The user can be viewing viewport #1 (e.g., at HQ) at to. The device can detect the movement from viewport #1 to viewport #2, and / or can act on the movement of the HMD at ti (ti >= to). The movement from viewport #1 to viewport #2 can affect the user's QoE. As an example, the device can reflect this head motion (e.g., on the HMD display), and the quality of viewport #1 can decrease after ti. The latency between to and ti can vary depending on different devices (e.g., due to sensors, processors, etc.). If the motion (e.g., head movement) is fine (e.g., less than a threshold), then the device can detect the motion, but not act on the motion. At t2, the head or device can have settled at viewport #2. The device can detect viewport #2, and can act on viewport #2 at t3 (t3 >= t2). For example, at t3, the device can determine that the HMD has stopped at a location associated with viewport #2 for a predetermined amount of time (e.g., n milliseconds). At t4, the user can observe the viewport shift and / or quality change in the HMD (e.g., HMD display). t4 can be less than or before t2 in time (e.g., the user can keep moving his head). The viewport change can be complete (e.g., at t5), and the device can perform display at the initial low quality viewport at t5. The quality of the display on the device can be restored. For example, the device can request one or more higher quality video segments (e.g., DASH segments) to produce viewport #2. In this way, the quality of viewport #2 can be high quality (e.g., the same quality as the previous viewport #1 at t6 (t6 >= t5), or higher quality than the previous viewport #1). The device (e.g., a metrics collection and processing (MCP) module) can collect input timing information from OP3, and / or collect this information on OP4. As an example, the device (e.g., MCP component) can derive corresponding latency metrics based on one or more of times t1-t6 to report to a metrics server (e.g., as shown in the client reference model of Figure 7 .

[0203] As an example, a VR device can measure various latency parameters. The viewport switch latency can correspond to the time difference between t5 and ti. The equal quality viewport switch latency can correspond to the time difference between t6 and ti. The initial latency can correspond to t4-to. The settling latency can correspond to the time difference between t5 and t2. In one embodiment, the device can record and / or report the time points (e.g., to, ti, t2, t3, t4, t5, t6) and / or the corresponding latency metrics (e.g., initial latency, viewport switch latency, equal quality viewport switch latency, settling latency). Figure 16The device can use a metric to indicate user position. As an example, the metric can include a 6DoF coordinate metric, as shown in Table 9. The device (e.g., a VR device) can report the 6DoF coordinate metric (e.g., report it to a metrics server). Local sensors can provide user position. The device can extract 6DoF coordinates based on user position and / or HMD viewport data. The 6DoF position can be represented by 3D coordinates (e.g., X, Y, Z) of the user's HMD. The 6DoF position can be represented by the controller position in the VR space and one or more relevant spherical coordinates (e.g., yaw, pitch, and roll, the spherical coordinate center can be X, Y, and Z) of the user's HMD or controller. When the user is equipped with multiple controllers, the device or user can signal multiple 6DoF position signals. Table 9 shows an example of reporting on 6DoF coordinates. The device or metrics server can derive the 6DoF coordinate metric and / or one or more entries shown in Table 9 from OP3.

[0204] The device can use a metric to indicate user position. As an example, the metric can include a 6DoF coordinate metric, as shown in Table 9. The device (e.g., a VR device) can report the 6DoF coordinate metric (e.g., report it to a metrics server). Local sensors can provide user position. The device can extract 6DoF coordinates based on user position and / or HMD viewport data. The 6DoF position can be represented by 3D coordinates (e.g., X, Y, Z) of the user's HMD. The 6DoF position can be represented by the controller position in the VR space and one or more relevant spherical coordinates (e.g., yaw, pitch, and roll, the spherical coordinate center can be X, Y, and Z) of the user's HMD or controller. When the user is equipped with multiple controllers, the device or user can signal multiple 6DoF position signals. Table 9 shows an example of reporting on 6DoF coordinates. The device or metrics server can derive the 6DoF coordinate metric and / or one or more entries shown in Table 9 from OP3.

[0205]

[0206] Table 9

[0207] The device (e.g., a VR device) can report gaze data. The metrics server can use the gaze data to perform analysis (e.g., to target advertisements). Table 10 shows an example of reporting on gaze data. The device or metrics server can derive the gaze data metric and / or one or more entries contained in Table 10 from OP3. The device can use the gaze data to decide how to configure and / or boost tile-based streaming (e.g., select a tile configuration).

[0208]

[0209] Table 10

[0210] The device (e.g., a VR device) can report frame rate metrics. The rendering frame rate can be different from the native frame rate. Table 11 shows an example of reporting on rendering frame rate. The device or metrics server can derive the frame rate metric and / or one or more entries contained in Table 11 from OP4.

[0211]

[0212] Table 11

[0213] General system quality of service (QoS) metrics or parameters (e.g., packet / frame loss rate, frame error rate, and / or frame discard rate) can represent quality of service (e.g., for regular video streaming). General QoS metrics can not reflect VR user experience for certain streaming (e.g., viewport-dependent 360-degree video streaming, say tile-based streaming). For regular video streaming, some or all packets or frames can have similar or identical impact on user experience. For viewport-dependent 360-degree video streaming (e.g., tile-based streaming), user experience can be determined (e.g., partially or primarily determined) by the viewport presented to the user. For viewport-dependent 360-degree video streaming (e.g., tile-based streaming), packet or frame loss for other non-viewport tiles can not impact user experience. For example, the user does not view other non-viewport tiles. Packet or frame loss for one or more viewport tiles viewed by the user can impact (e.g., significantly impact) user experience. General QoS metrics can not reflect VR user experience.

[0214] The device can use the metrics to indicate one or more viewport loss events. As an example, the metrics can include the viewport loss metrics described in Table 12. The device can use the viewport loss metrics to report one or more viewport loss events for user experience analysis (e.g., for viewport-dependent 360-degree video streaming, say tile-based streaming). For instance, the client can determine whether one or more loss events impact the viewport displayed to the user, and / or can report only loss events that impact the displayed viewport with the viewport loss metrics. As shown in Table 12, the viewport loss metrics can include one or more timing information to indicate the time when the loss is observed, and / or other loss information. The device or the metrics server can derive the viewport loss metrics and / or one or more entries in Table 12 from OP1 and / or OP3.

[0215]

[0216] Table 12

[0217] The viewport loss metrics can report viewport segment loss events and / or provide information or additional information about the loss status. One entry (e.g., entry "sourceUrl") can indicate (e.g., specify) the URL of the lost viewport segment. An entry (e.g., entry "lossReason") can indicate one or more reasons for the segment loss (as an example, including one or more reasons described herein).

[0218] One entry of the viewport loss metrics (e.g., entry "server error") can indicate (e.g., specify) the reason why the server did not complete the viewport segment request (e.g., the particular viewport segment request). Here, the list of HTTP response status 5xx server error codes can be used, for example, 504 corresponds to "Gateway Timeout", 505 corresponds to "HTTP Version not supported", and so on. One entry of the viewport loss metrics (e.g., entry "client error") can indicate (e.g., specify) the client error when requesting (e.g., the particular) viewport segment at the client. The list of HTTP response status 4xx client error codes can be used, for example, 401 corresponds to "Unauthorized", 404 corresponds to "Not Found", and so on. One entry of the viewport loss metrics (e.g., entry "packet loss") can indicate (e.g., specify) that one or more packets of the viewport segment associated with the source Url failed to reach its destination. One entry of the viewport loss metrics (e.g., entry "packet error") can indicate (e.g., specify) that one or more packets of the viewport segment associated with the source Url are invalid. One entry of the viewport loss metrics (e.g., entry "packet dropped") can indicate (e.g., specify) that one or more packets of the viewport segment associated with the source Url are dropped. One entry of the viewport loss metrics (e.g., entry "error") can provide a detailed (e.g., more detailed) error message. For example, the field "error" can provide the HTTP response error code, as an example, 504 corresponds to "Gateway Timeout", 505 corresponds to "HTTP Version not supported", 401 corresponds to "Unauthorized", 404 corresponds to "Not Found", and so on. With the error metrics described here, the metrics server is able to analyze the impact of network conditions on the VR user experience and / or to track one or more root causes.

[0219] A device (e.g., a VR device) can generate and / or report metrics related to the HMD, such as precision and / or sensitivity. The device can use the metrics to indicate the angular positioning consistency between physical movement measured in degrees and visual feedback in the VR domain. The metric can include a precision metric. Table 13 shows an example of the report on HMD precision. The device or the metrics server can derive the precision metric and / or one or more entries contained in Table 13 from OP3 and / or OP5.

[0220]

[0221] Table 13

[0222] For example, at time t, the user's head can be pointing at coordinate A, however the image displayed to the user can reflect coordinate B. The accuracy metric can correspond to this difference (e.g., coordinate B minus coordinate A). As an example, the user can move their head 3 degrees to the right, and the image displayed on the HMD can reflect that the user moved 2.5 degrees to the right and 0.2 degrees down. The accuracy metric can indicate this measured difference, e.g., a 0.5 degree difference horizontally and a 0.2 degree difference vertically. The metric can report the difference. The difference can be in the form of a delta coordinate (e.g., a vector). The difference can be in the form of an absolute difference (e.g., a single value). The HMD and / or controller can provide the accuracy metric. The HMD can be included in a VR device. The delta value can vary depending on the distance between the sensor and the user.

[0223] The device can use a metric to indicate the ability of the HMD inertial sensor to sense small movements and subsequently provide feedback to the user. The metric can include a sensitivity metric. Table 14 shows an example of a report on HMD sensitivity. The device or metric server can derive the sensitivity metric and / or one or more entries contained in Table 14 from OP3 and / or OP5.

[0224]

[0225] Table 14

[0226] The device and / or metric server can use the accuracy and sensitivity in various ways. A content provider or server can provide one or more customized content suggestions to the user based on the accuracy and / or sensitivity of the HMD. The content provider, server, or device can collect accuracy and / or sensitivity metrics for VR device model / brand ratings, and / or share these metrics with the manufacturer (e.g., of the VR device), thereby addressing issues and / or improving the device.

[0227] The device (e.g., VR device) can send the accuracy and / or sensitivity as a PER message. The DANE can send the accuracy and / or sensitivity message to the DASH client (e.g., containing the VR device). The device or DANE can send the appropriate accuracy / sensitivity settings for the selected content to the configurable HMD. The viewer's experience of the selected content is enhanced. Table 15 shows an example of a report on accuracy messages. Table 16 shows an example of a report on sensitivity messages.

[0228]

[0229] Table 15

[0230]

[0231] Table 16

[0232] A device (e.g., a VR device) can generate a message related to ROI. The ROI message can be of PER type. The ROI message can indicate to the viewer the desired ROI. The DASH client can pre-fetch the desired ROI. The ROI message can indicate to the DASH client to pre-fetch the desired ROI. The ROI can change over time. The device or content provider can infer the ROI from viewport view statistics. The device or content provider can dictate the ROI (e.g., director's cut). The device or content provider can determine the ROI based on real-time analytics. The DANE can send this ROI information as a PER message to the DASH client. Based on the ROI information, the DASH client can determine the tile / segment / byte range to retrieve from the server. For example, the device or a metrics server can derive the ROI message from OP4.

[0233] ROI can indicate a VR region that has high priority in terms of attractiveness, importance, or quality (e.g., prompt the user to make a decision ahead of time). The user can view the front in a consistent manner. With the ROI message, the user can be prompted to view another region, but can still view the front of the video.

[0234] The DANE can receive the ROI information as a PED message in order to perform pre-fetching and boost cache performance. In the PED message, if the ROI is provided, the DANE can determine the tile / segment / byte range to retrieve. One or more specific URLs can be provided in the PED message. The ROI information message can be sent from the DANE to another DANE. Table 17 shows an example of the content for the ROI information message.

[0235]

[0236] Table 17

[0237] The ROI information can be carried in a timed metadata track or event stream element.

[0238] Figure 17A message flow showing SAND messages performed between a DANE and a DASH client or between a DASH client and a metrics server is shown. DANE 1704 (e.g., origin server) can send video segments to DANE 1706 (e.g., CDC / cache). DANE 1706 (e.g., CDC / cache) can send video segments to DASH client 1708. DASH client 1708 can send metrics to metrics server 1702. The metrics can include one or more of: viewport view, rendering device, viewport switch latency, initial latency, settling latency, 6DoF coordinates, gaze data, frame rate, precision, and / or sensitivity, among others. DANE 1704 (e.g., origin server) can send information (e.g., initial rendering orientation, precision, sensitivity, and / or ROI, among others) to DASH client 1708. DANE 1706 (e.g., CDC / cache) can send information (e.g., initial rendering orientation, precision, sensitivity, and / or ROI, among others) to DASH client 1708. DASH client 1708 can send information (e.g., rendering device, viewport switch latency, and / or frame rate, among others) to DANE 1706 (e.g., CDC / cache) or DANE 1704 (e.g., origin server).

[0239] The device can detect, derive, and / or report the metrics and / or metric changes to one or more metrics servers for performing analysis and / or calculations on the one or more metrics servers. The VR device, the metrics server, and / or the controller can perform one or more of the detecting, deriving, analyzing, and / or calculating.

[0240] While features and elements are described above in particular combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer readable medium for execution by a computer or processor. Examples of computer readable media include electronic signals (machine- readable storage) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as, but not limited to, a hard disk, floppy diskette, and a tape, optical media such as, but not limited to, a compact disc (CD) and a digital versatile disc (DVD), and the like. A processor in association with software can be used to implement a radio frequency transceiver for use in a WTRU, terminal, base station, RNC, or any host computer. The processor can be used to implement a radio frequency transceiver for use in a WTRU, terminal, base station, RNC, or any host computer.

Claims

1. A device for receiving and displaying media content, the device comprising: a first interface coupled to a sensor component, wherein the first interface is configured to receive first metric data associated with a user orientation from the sensor component; a second interface coupled to a rendering component, wherein the second interface is configured to receive second metric data associated with rendering a Dynamic Adaptive Streaming over HTTP (DASH) video segment from the rendering component; and a processor configured to: receive the first metric data associated with the user orientation via the first interface coupled to the sensor component, receive the second metric data associated with rendering the DASH video segment via the second interface coupled to the rendering component, wherein the second metric data is generated by using one or more of: color conversion, projection, or media composition, generate a latency metric based on the first metric data associated with the user orientation and the second metric data associated with rendering the DASH video segment, and send the latency metric to a metric server.

2. The device of claim 1, wherein the device further comprises at least one of: a third interface coupled to a network access component, the third interface configured to receive third metric data associated with requesting and receiving the DASH video segment from a network from the network access component; a fourth interface coupled to a media processing component, the fourth interface configured to receive fourth metric data associated with decoding the DASH video segment received from the network from the media processing component; and a fifth interface coupled to a control component, the fifth interface configured to receive fifth metric data associated with configuration parameters corresponding to the device from the control component, wherein the latency metric is further generated based on at least one of: the third metric data associated with requesting and receiving the DASH video segment from the network, the fourth metric data associated with decoding the DASH video segment received from the network, and the fifth metric data associated with the configuration parameters corresponding to the device.

3. The device of claim 1, wherein the latency metric comprises a comparable quality viewport switch latency. the processor comprises a metric processing component coupled to the first interface and the second interface, and the latency metric is generated by the metric processing component.

5. The device of claim 1, wherein the device comprises a third interface, a fourth interface, and a fifth interface, the processor comprises a metric processing component coupled to the first interface, the second interface, the third interface, the fourth interface, and the fifth interface, and the latency metric is generated by the metric processing component.

4. The apparatus of claim 1, wherein, ​ ​ 6. The apparatus of claim 1, wherein, The latency metric is further generated based on one of third metric data associated with requesting and receiving the DASH video segments from a network, fourth metric data associated with decoding the DASH video segments received from the network, or fifth metric data associated with configuration parameters corresponding to the device.

7. The apparatus of claim 1, wherein, The latency metric is selected from a group consisting of a viewport switch latency, a quality viewport switch latency, an initial latency, and a setup latency.

8. The apparatus of claim 1, wherein, A network access component is coupled to at least one of the first interface or the second interface over a communication network.

9. The apparatus of claim 1, wherein, The device comprises a virtual reality client device.

10. A method for receiving and displaying media content, the method comprising: receiving, via a first interface coupled to a sensor component, first metric data associated with a user orientation, wherein the first interface is configured to receive the first metric data associated with the user orientation from the sensor component; receiving, via a second interface coupled to a rendering component, second metric data associated with rendering HTTP Dynamic Adaptive Streaming over HTTP (DASH) video segments, wherein the second interface is configured to receive the second metric data associated with rendering the DASH video segments from the rendering component, and wherein the second metric data is generated by using one or more of: color conversion, projection, or media composition; generating a latency metric based on the first metric data associated with the user orientation and the second metric data associated with rendering the DASH video segments; and sending the latency metric to a metric server.

11. The method of claim 10, wherein, A metric processing component is coupled to the first interface and the second interface, and the latency metric is generated by the metric processing component.

12. The method of claim 10, wherein, The latency metric is further generated based on one of third metric data associated with requesting and receiving the DASH video segments from a network, fourth metric data associated with decoding the DASH video segments received from the network, or fifth metric data associated with configuration parameters corresponding to the device.

13. The method of claim 10, wherein, The latency metric is selected from a group consisting of a viewport switch latency, a quality viewport switch latency, an initial latency, and a setup latency.

14. The method of claim 10, wherein, A network access component is coupled to at least one of the first interface or the second interface over a communication network.

15. The method of claim 10, wherein, The latency metric comprises a comparable quality viewport switch latency.

16. The method of claim 10, wherein the method further comprises: receiving, via a third interface coupled to a network access component, third metric data associated with requesting and receiving the DASH video segments from a network, wherein the third interface is configured to receive the third metric data from the network access component; receiving, via a fourth interface coupled to a media processing component, fourth metric data associated with decoding the DASH video segments received from the network, wherein the fourth interface is configured to receive the fourth metric data associated with decoding the DASH video segments received from the network from the media processing component; and receiving, via a fifth interface coupled to the control component, fifth metric data associated with configuration parameters corresponding to the device, wherein the fifth interface is configured to receive, from the control component, the fifth metric data associated with the configuration parameters corresponding to the device, wherein the latency metric is generated based further on at least one of: the third metric data associated with requesting and receiving the DASH video segments from the network, the fourth metric data associated with decoding the DASH video segments received from the network, and the fifth metric data associated with the configuration parameters corresponding to the device.

Citation Information

Patent Citations

  • Method and terminal for displaying video state

    CN104735515A

  • Metrics and messages to improve experience for 360-degree adaptive streaming

    CN110622483A

  • Viewport dependent video streaming events

    CN112219406A

  • System and method for providing 360 degrees immersive video based on gaze vector information

    CN112335236A

  • Methods and devices for rendering a video on a display

    CN113994707A