System and method for predictive overfill for virtual reality

By predicting the user's future head position on the server side and rendering an overfilled image, the problems of latency and computing power limitations in VR services are solved, improving the stability and comfort of the user experience.

CN112313712BActive Publication Date: 2025-12-30INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980039734.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-19
Filing Date
2019-04-18
Publication Date
2025-12-30
Estimated Expiration
2039-04-18

AI Technical Summary

Technical Problem

In virtual reality (VR) services, especially remote VR services, processing latency and computing power limitations lead to inconsistent user experiences and motion sickness, affecting user satisfaction.

Method used

The system predicts the user's future head position on the server side, renders an image based on the predicted future head position and overfill factor, and sends the overfilled image to the client device for display to reduce latency and inconsistency.

Benefits of technology

Predictive overfill technology reduces latency and image inconsistency, improves user experience stability and comfort, and reduces the incidence of motion sickness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112313712B_ABST
    Figure CN112313712B_ABST
Patent Text Reader

Abstract

An exemplary disclosed method according to some embodiments includes receiving head tracking position information from a client device, the head tracking position information associated with a user at the client device; predicting a future head position of the user at a scan-out time for displaying a virtual reality (VR) video frame, wherein the VR video frame is displayed to the user via the client device; determining an overfill factor based on an expected error in the predicted future head position of the user; rendering an overfilled image based on the predicted future head position of the user and the overfill factor; and sending the VR video frame including the overfilled image to the client device for display to the user.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference section

[0002] This application is a non-provisional application of the following application and claims the benefit of it under 35 U.SC §119(e): U.S. Provisional Patent Application Serial No. 62 / 660,228, filed April 19, 2018, entitled "Systems and Methods Employing Predictive Overfilling for Virtual Reality," which is incorporated herein by reference in its entirety. Background Technology

[0003] In some virtual reality (VR) services, VR content can be processed in a separate processing entity (e.g., cloud, server farm, and local desktop PC) located separately from the client device (e.g., HMD), and the VR content can be transmitted to the client device (e.g., HMD) for display to the user. Because processing VR content typically consumes significant computing power, processing all VR content solely at the HMD may be inappropriate or impractical; the HMD may also have limitations in terms of computing power, battery life, and thermal budget.

[0004] Furthermore, latency is widely considered an obstacle, even in local VR services (where, for example, servers and client devices are directly connected via HDMI cables). Latency typically increases or becomes more severe in remote VR services due to additional processing (e.g., encoding and / or decoding) and / or network transmission latency. Generally, latency can be measured as the interval between the time a user moves and the time it takes for the user to see an image of the correspondingly changed view, and the latency measured as this interval is often referred to as motion-to-photon (MTP) latency. As this latency increases, the inconsistency between the generated VR video frames and the user's field of view (FOV) at the time of display also increases. Particularly in remote VR services, the significant additional latency from video compression and decompression, as well as network transmission, can lead to large inconsistencies. Large inconsistencies can cause significant motion sickness in users, resulting in an uncomfortable user experience and / or potentially reducing user satisfaction (e.g., causing users to quit the service). Summary of the Invention

[0005] According to some embodiments, a method performed at a server includes: receiving head tracking position information from a client device, the head tracking position information being associated with a user at the client device; predicting the user's future head position at a scan-out time for displaying a virtual reality (VR) video frame, wherein the VR video frame is displayed to the user via the client device; determining an overfill factor based on an expected error in the user's predicted future head position; rendering an overfilled image based on the user's predicted future head position and the overfill factor; and sending the VR video frame including the overfilled image to the client device for display to the user.

[0006] In some embodiments, the client device includes a head-mounted display (HMD). In some embodiments, the expected error in the user's predicted future head position is based on observations of head tracking position information received by the server. In some embodiments, the expected error in the user's predicted future head position is based on observations of network latency over time.

[0007] In some embodiments, the method further includes: rendering at least one other VR video frame containing another overfilled image, wherein the size of the at least one other VR video frame is different from the size of the VR video frame containing the overfilled image. In this regard, in some embodiments, at least one of the pixel size or aspect ratio is dynamically changed from one VR video frame to another based on changes in latency of the connection between the server and the client device. In some embodiments, at least one of the pixel size or aspect ratio is dynamically changed from one VR video frame to another based on changes in observed user head movement (e.g., head rotation).

[0008] In some embodiments, predicting the user's future head position at the scan output time includes: using at least part of the head tracking position information to predict the user's future head position at the scan output time. In some embodiments, the head tracking position information is based on IMU motion data. In some embodiments, the method further includes: receiving timing information from the client device in response to receiving the VR video frame including the overfilled image at the client device, the timing information including at least the scan output start time of the VR video frame containing the rendered overfilled image. Furthermore, in some embodiments, the method includes determining the rendering-to-scan-out delay distribution.

[0009] In some embodiments, determining the render-to-scan output latency distribution includes: at least partially determining the difference between the rendering start time of the VR video frame and the scan output start time of the VR video frame to calculate a render-to-scan output latency value for the VR video frame. In some embodiments, the render-to-scan output latency value of the VR video frame is added to a table configured to maintain a plurality of render-to-scan output latency values ​​associated with the rendered VR video frame; and the table is used to determine the render-to-scan output latency distribution. In some embodiments, each time a new render-to-scan output latency value is added to the latency table, an older render-to-scan output latency value is deleted from the latency table.

[0010] Furthermore, in some embodiments, using the head tracking position information at least in part to predict the user's future head position at the scan output time includes: predicting a first field of view (FOV) at a first time (T1), wherein the first predicted FOV is based on a predicted first fixation point; and predicting a second FOV at a second time (T2), wherein the second predicted FOV is based on a predicted second fixation point. In some embodiments, T1 and T2 are selected based on the render-to-scan output delay distribution. In some embodiments, T1 provides a lower bound for the expected scan output time of the VR video frame, and T2 provides an upper bound for the expected scan output time of the VR video frame. Furthermore, in some embodiments, T1 and T2 are selected such that the time interval between T1 and T2 includes the target probability of the delay distribution.

[0011] In some embodiments, the method further includes adding a first error tolerance associated with the predicted first gaze point to the first predicted FOV; and adding a second error tolerance associated with the predicted second gaze point to the second predicted FOV. In some embodiments, the second error tolerance is greater than the first error tolerance. Furthermore, in some embodiments, each of the first and second error tolerances associated with the predicted first and second gaze points is based on a prediction technique selected from the group consisting of constant rate (velocity)-based prediction (CRP) and constant acceleration-based prediction (CAP). In some embodiments, the method further includes confirming the values ​​of the first and second error tolerances in real time based on the head tracking position information received from the client device. In some embodiments, the first prediction error tolerance and the second prediction error tolerance are based on the error between the received head tracking position information and the predicted motion data.

[0012] In some embodiments, determining the overfill factor based on the expected error in the predicted head position of the user includes setting the overfill factor at least in part based on a first prediction error tolerance and a second prediction error tolerance. In some embodiments, the overfill factor includes a first overfill factor value for the horizontal axis and a second overfill factor value for the vertical axis, the first and second overfill factor values ​​being different from each other. In some embodiments, the method further includes determining a combined FOV associated with the overfilled image based on the first predicted FOV and the second predicted FOV. Determining the combined FOV based on the first predicted FOV and the second predicted FOV includes combining (i) a first adjusted predicted FOV and (ii) a second adjusted predicted FOV, wherein the first adjusted predicted FOV is determined by adding the first error tolerance to the first predicted FOV, and the second adjusted predicted FOV is determined by adding the second error tolerance to the second predicted FOV. In some embodiments, the combined FOV is determined by selecting a rectangular region that includes the first and second adjusted FOVs. In some embodiments, the combined FOV is determined by selecting a hexagonal shape that includes the first and second adjusted FOVs. In some embodiments, rendering the overfilled image includes applying the overfill factor relative to the center point of the combined FOV.

[0013] According to some embodiments, the method further includes determining a time T, where time T represents the predicted scan output time of the VR video frame containing the overfilled image. In some embodiments, the time T is predicted based on a render-to-scan output delay distribution. In some embodiments, the time T corresponds to the median or average of the render-to-scan output delay distribution. According to some embodiments, the method further includes: determining the extended FOV of the user within the time T based on the direction and speed of the user's head rotation; and aligning the center position of the extended FOV with the center position of the predicted FOV to produce a final extended FOV.

[0014] According to some embodiments, a method performed at a server includes: determining that a loss in the field of view (FOV) of a virtual reality (VR) frame transmitted to a client device has occurred; and adaptively adjusting an overfill factor weight based on the determination that the loss has occurred in the FOV. In some embodiments, the method further includes: increasing the overfill weight factor in response to determining that the loss in the FOV has occurred; and decreasing the overfill weight factor in response to determining that the loss in the FOV has not occurred. In some embodiments, the reduction of the overfill factor is performed once it is determined that no loss in the FOV has occurred in a given number of VR video frames.

[0015] Furthermore, in some embodiments, the method includes receiving feedback information from the client device when the loss in the FOV has occurred. In some embodiments, the feedback information includes an FOV loss rate. Furthermore, in some embodiments, the method includes determining a first overfill factor value for the horizontal axis and a second overfill factor value for the vertical axis, wherein determining the first and second overfill factors includes identifying the rotation speed of the user's head on the client side for each of the horizontal and vertical axes, and multiplying the weighting factor by the rotation speed. Furthermore, in some embodiments, the method includes performing a ping exchange with the client device to determine connection latency.

[0016] According to some embodiments, a method performed by a virtual reality (VR) client device includes: receiving a first virtual reality (VR) video frame from a server; in response to receiving the first VR video frame, sending timing information to the server, wherein the timing information includes at least a scan output start time of the received first VR video frame; sending motion information of a user of the VR client device to the server; receiving a second VR video frame from the server, wherein the second VR video frame contains an overfilled image based on (i) a predicted head position of the user at a scan output time of the second VR video frame for display to the user and (ii) an overfill factor; and displaying a selected portion of the overfilled image to the user, wherein the portion is selected based on the actual head position of the user at the scan output time of the second VR video frame, and wherein the predicted head position is based on the transmitted motion information of the user, and the overfill factor is based on an expected error in the predicted head position of the user.

[0017] In some embodiments, the motion information includes IMU-based motion data. In some embodiments, the client device includes a head-mounted display (HMD). Furthermore, in some embodiments, the frame size of the received second VR video frame, which includes the overfilled image, is different from the frame size of the received first VR video frame.

[0018] According to some embodiments, the aspect ratio of the received second VR video frame, including the overfilled image, is different from the aspect ratio of the received first VR video frame. In some embodiments, due to variations in connection latency between the client device and the server, at least one of the pixel dimensions or aspect ratio of the received second VR video frame, including the overfilled image, differs from at least one of the pixel dimensions or aspect ratio of the received first VR video frame. In some embodiments, due to variations in user head rotation, at least one of the pixel dimensions or aspect ratio of the received second VR video frame, including the overfilled image, differs from at least one of the pixel dimensions or aspect ratio of the received first VR video frame.

[0019] Furthermore, in some embodiments, the method includes receiving an indication from the server regarding at least one of the pixel dimension or the aspect ratio before the scan output time. In some embodiments, each of the received first VR video frame and second VR video frame includes a corresponding timestamp indicating the frame rendering time at the server. Furthermore, in some embodiments, the method includes time warping at least one of the first VR video frame or the second VR video frame. Furthermore, in some embodiments, the method includes tracking the user's actual head position.

[0020] Other embodiments include systems, servers, and VR client devices configured (e.g., having a processor and a non-transitory computer-readable medium storing a plurality of instructions for execution by the processor) to perform the methods described herein. In some embodiments, the VR client device includes a head-mounted display (HMD). Attached Figure Description

[0021] In the accompanying drawings, the same reference numerals denote the same elements, and wherein:

[0022] Figure 1A This is a system schematic diagram illustrating an exemplary communication system that can implement one or more embodiments disclosed;

[0023] Figure 1B It is shown that, according to the embodiment, it is possible to Figure 1A A schematic diagram of an exemplary wireless transmit / receive unit (WTRU) used within a communication system.

[0024] Figure 1C It is shown that, according to the embodiment, it is possible to Figure 1A The diagram shows an exemplary radio access network (RAN) and an exemplary core network (CN) used within the communication system.

[0025] Figure 1DIt is shown that, according to the embodiment, it is possible to Figure 1A The diagram shows another exemplary RAN and another exemplary CN used within the communication system.

[0026] Figure 2 This is a flowchart of an exemplary process used for local VR services.

[0027] Figure 3 An exemplary process for remote or cloud-based VR services is shown.

[0028] Figure 4 This is a schematic diagram illustrating an exemplary time warp of a video frame according to some embodiments.

[0029] Figure 5 This is a flowchart of an exemplary server-side time warping process according to some embodiments.

[0030] Figure 6 This is a flowchart of an exemplary client-side (e.g., HMD-side) time warping process according to some embodiments.

[0031] Figure 7 An example is shown of a scenario in which FOV loss occurs after time warping, according to some embodiments.

[0032] Figure 8 An exemplary overfilling process is shown.

[0033] Figure 9 An example of an overfilled rendering area is shown.

[0034] Figure 10 An exemplary virtual reality environment according to some embodiments is shown.

[0035] Figure 11 This is a flowchart of an exemplary process for predictive overfilling according to some embodiments.

[0036] Figure 12 This is a flowchart of an exemplary process according to some embodiments.

[0037] Figure 13 An exemplary delay table management process according to some embodiments is shown.

[0038] Figure 14 This is a schematic diagram illustrating the exemplary delay distribution probability according to some embodiments.

[0039] Figure 15 It is a graph showing the probability distribution of exemplary delays, including exemplary time intervals, according to some embodiments.

[0040] Figure 16This is a schematic diagram illustrating an exemplary prediction of gaze points according to some embodiments.

[0041] Figure 17A and 17B These are two graphs showing the experimental results of using two prediction techniques based on some embodiments.

[0042] Figure 18 Illustrations are shown according to some embodiments Figure 17A The graph shows the FOV and an example of predicted FOV.

[0043] Figure 19 Exemplary error tolerance configurations according to some embodiments are shown.

[0044] Figure 20 This is a perspective view of a VR HMD according to some embodiments.

[0045] Figure 21A An exemplary rendering area is shown according to some embodiments.

[0046] Figure 21B Another exemplary rendering area is shown according to some embodiments.

[0047] Figure 21C The example of an overfilled rendering area is shown.

[0048] Figure 22 A schematic diagram illustrating the potential FOV formation according to some embodiments is shown in more detail.

[0049] Figure 23 This is an illustrated representation of the use of exemplary calculated values ​​from the overfill factor according to some embodiments.

[0050] Figure 24 This is a flowchart of an exemplary process for predictive overfilling according to some embodiments.

[0051] Figure 25 This is a schematic diagram illustrating an exemplary delay distribution according to some embodiments.

[0052] Figure 26 This is a schematic diagram showing several FOVs according to some embodiments.

[0053] Figure 27 This is a schematic diagram illustrating the determination of the overfill factor according to some embodiments.

[0054] Figure 28 This illustrates how latency distribution can be changed during various stages of a VR service, according to some embodiments.

[0055] Figure 29AThis is a schematic diagram showing an exemplary output VR frame that takes into account changes in system latency.

[0056] Figure 29B This is a schematic diagram illustrating the effect of system latency on the output VR frame according to some embodiments.

[0057] Figure 30A This is a schematic diagram illustrating an exemplary output VR frame that takes into account the user's head rotation direction.

[0058] Figure 30B This is a schematic diagram illustrating the effect of the user's head rotation direction on the output VR frame according to some embodiments.

[0059] Figure 31 This is a message passing diagram illustrating an exemplary process according to some embodiments.

[0060] Figure 32 This is a schematic diagram illustrating an exemplary FOV loss region according to some embodiments.

[0061] Figure 33 The diagram illustrates the relationship between [various embodiments] and [other embodiments]. Figure 32 Exemplary processes related to the process

[0062] Figure 34 A schematic diagram illustrating an example of changing the delay from rendering to scan output, according to some embodiments, is shown.

[0063] Figure 35 This is a schematic diagram illustrating an exemplary delay table update according to some embodiments.

[0064] Figure 36 Exemplary delay distribution probabilities according to some embodiments are shown.

[0065] Figure 37A Exemplary time values ​​from some embodiments are shown. Figure 16 A schematic diagram.

[0066] Figure 37B Exemplary time values ​​are shown according to some embodiments. Figure 17B The curve graph.

[0067] Figure 38 Potential FOVs according to some embodiments are shown. In some implementations...

[0068] Figure 39 This is a schematic diagram illustrating the relationship between potential FOV and overfill factor according to some embodiments.

[0069] Figure 40 This is a flowchart of an exemplary method according to some embodiments.

[0070] Figure 41 This is a flowchart illustrating another exemplary method according to some embodiments.

[0071] Figure 42 This is a flowchart illustrating yet another exemplary method according to some embodiments.

[0072] Figure 43 Exemplary computing entities that may be used in embodiments of this disclosure are depicted. Detailed Implementation

[0073] Figure 1A This is a schematic diagram illustrating an exemplary communication system 100 that can implement one or more of the disclosed embodiments. The communication system 100 can be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 can use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT-Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtering OFDM, and Filter Bank Multicarrier (FBMC), etc.

[0074] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each of WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a wireless environment. For example, any of WTRUs 102a, 102b, 102c, and 102d can be referred to as a “station” and / or “STA”, and can be configured to transmit and / or receive wireless signals. They can include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics, and devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d can be interchangeably referred to as a UE.

[0075] The communication system 100 may further include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to enable access to one or more communication networks (e.g., CN106 / 115, Internet 110, and / or other networks 112) by wirelessly interfacing with at least one of WTRUs 102a, 102b, 102c, and 102d. For example, base stations 114a and 114b may be base transceiver stations (BTS), node B, e-node B, home node B, home e-node B, gNB, new radio (NR) node B, site controller, access point (AP), and wireless routers, etc. Although each of base stations 114a and 114b is described as a single component, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network components.

[0076] Base station 114a may be part of RAN 104 / 113, and the RAN may also include other base stations and / or network components (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies called cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide radio service coverage for a specific geographic area that is relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., each transceiver corresponds to one sector of the cell. In embodiments, base station 114a may use multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, by using beamforming, signals can be transmitted and / or received in a desired spatial direction.

[0077] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, wherein the air interface can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).

[0078] More specifically, as described above, the communication system 100 can be a multiple access system and can use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, and 102c can implement a certain radio technology, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), wherein the technology can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0079] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a certain radio technology, such as Evolved UMTS Terrestrial Radio Access (E-UTRA), wherein the technology may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.

[0080] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a radio technology that can establish an air interface 116 using a new radio (NR), such as NR radio access.

[0081] In this embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access (e.g., using the dual connectivity (DC) principle). Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).

[0082] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN), etc.

[0083] Figure 1ABase station 114b can be, for example, a wireless router, home node B, home e node B, or access point, and can use any suitable RAT to facilitate wireless connectivity in a local area, such as a business premises, residence, vehicle, campus, industrial facility, air corridor (e.g., for use by drones), and road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless local area network (WLAN) by implementing radio technology such as IEEE 802.11. In another embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless personal area network (WPAN) by implementing radio technology such as IEEE 802.15. In yet another embodiment, base station 114b and WTRUs 102c, 102d can establish a picocell or femtocell by using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). Figure 1A As shown, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b does not need to access the Internet 110 via CN 106 / 115.

[0084] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 102a, 102b, 102c, and 102d. This data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or can perform advanced security functions such as user authentication. Although in Figure 1A While not shown, it should be understood that RAN104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT or a different RAT as RAN 104 / 113. For example, in addition to connecting to RAN 104 / 113 which uses NR radio technology, CN 106 / 115 can also communicate with other RANs (not shown) that use GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technologies.

[0085] CN 106 / 115 can also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Simple Old-Style Telephone Service (POTS). The Internet 110 may include a global interconnected computer network equipment system using common communication protocols (e.g., TCP, UDP, and / or IP from the Transmission Control Protocol / Internet Protocol (TCP / IP) suite). Network 112 may include wired or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, wherein the one or more RANs may use the same RAT or a different RAT as RAN 104 / 113.

[0086] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capability (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers communicating with different wireless networks on different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a using cellular-based radio technology, and with base station 114b using IEEE 802 radio technology.

[0087] Figure 1B This is a system schematic diagram illustrating an exemplary WTRU 102. (See attached diagram.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmitter / receiver unit 122, a speaker / microphone 124, a numeric keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or peripheral devices 138. It should be understood that, while remaining consistent with the embodiments, WTRU 102 may also include any sub-combination of the foregoing components.

[0088] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, and transceiver 120 can be coupled to transmitting / receiving unit 122. Although Figure 1B While processor 118 and transceiver 120 are described as separate components, it should be understood that processor 118 and transceiver 120 can also be integrated together in a single electronic component or chip.

[0089] Transmit / receive component 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmit / receive component 122 may be an antenna configured to transmit and / or receive RF signals. As an example, in another embodiment, transmit / receive component 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmit / receive component 122 may be configured to transmit and / or receive RF and optical signals. It should be understood that transmit / receive component 122 may be configured to transmit and / or receive any combination of wireless signals.

[0090] Although Figure 1B While the transmit / receive component 122 is described as a single component, the WTRU 102 may include any number of transmit / receive components 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive components 122 (e.g., multiple antennas) that transmit and receive wireless signals via the air interface 116.

[0091] Transceiver 120 can be configured to modulate signals to be transmitted by transmitter / receiver 122 and demodulate signals received by transmitter / receiver 122. As described above, WTRU 102 can have multimode capability. Therefore, transceiver 120 can include multiple transceivers that allow WTRU 102 to communicate using various RATs (e.g., NR and IEEE 802.11).

[0092] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a numeric keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from these components. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 can include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 can include a subscriber identity module (SIM) card, memory stick, secure digital card (SD) memory card, etc. In other embodiments, the processor 118 can access and store information from memory that is not actually located in WTRU 102; for example, such memory could be located in a server or home computer (not shown).

[0093] The processor 118 can receive power from the power source 134 and can be configured to distribute and / or control power for other components in the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cell battery packs (such as nickel-cadmium (Ni-Cd), nickel-zinc (Ni-Zn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells, etc.

[0094] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) related to the current location of the WTRU 102. As a supplement or replacement to the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116, and / or determine its location based on signal timing received from two or more nearby base stations. It should be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable positioning method.

[0095] The processor 118 can also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, the peripheral device 138 may include an accelerometer, electronic compass, satellite transceiver, digital camera (for photos and / or video), Universal Serial Bus (USB) port, vibration device, television transceiver, hands-free headset, etc. Modules, FM radio units, digital music players, media players, video game console modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. The peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors, etc.

[0096] WTRU 102 may include a full-duplex wireless device, wherein the reception or transmission of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous for the wireless device. The full-duplex wireless device may include an interference management unit that reduces and / or substantially eliminates self-interference by means of hardware (e.g., choke coils) or by means of a processor (e.g., a separate processor (not shown) or by means of processor 118) for signal processing. In embodiments, WTRU 102 may include a half-duplex wireless device that transmits and receives some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception).

[0097] Figure 1C This diagram illustrates a system schematic of RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.

[0098] RAN 104 may include eNodeBs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNodeBs while remaining consistent with the embodiments. Each of eNodeBs 160a, 160b, and 160c may include one or more transceivers communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNodeBs 160a, 160b, and 160c may implement MIMO technology. Thus, for example, eNodeB 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0099] Each of the eNodeB 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. For example... Figure 1C As shown, nodes B160a, 160b, and 160c can communicate with each other via the X2 interface.

[0100] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing components is described as part of the CN 106, it should be understood that any of these components may be owned and / or operated by an entity other than the CN operator.

[0101] The MME 162 can connect to each of the eNodeBs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, performing bearer activation / deactivation processes, and selecting a specific serving gateway during the initial attach process of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) using other radio technologies (such as GSM and / or WCDMA).

[0102] The SGW 164 can connect to each of the eNodeBs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 typically routes and forwards user data packets to / from WTRUs 102a, 102b, and 102c. Furthermore, the SGW 164 can perform other functions, such as anchoring the user plane during handover between eNBs, triggering paging processes when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c, etc.

[0103] SGW 164 can be connected to PGW 146, which can provide packet-switched network (e.g., Internet 110) access for WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.

[0104] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to a circuit-switched network (e.g., PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and conventional landline communication equipment. For example, CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server), and the IP gateway may act as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRUs 102a, 102b, and 102c with access to the other network 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0105] Although Figure 1A-1D The WTRU is described as a wireless terminal; however, it should be understood that in some representative embodiments, such a terminal may use a wired communication interface (e.g., temporary or permanent) with the communication network.

[0106] In a representative embodiment, the other network 112 may be a WLAN.

[0107] A WLAN employing an Infrastructure Basic Services Set (BSS) model may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may access or interface with a distributed system (DS) or other types of wired / wireless networks that send traffic into and / or out of the BSS. Traffic originating outside the BSS and destined for a STA can be delivered to the STA via the AP. Traffic originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, provided that the source STA can send traffic to the AP and the AP can deliver traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be sent between the source and destination STAs (e.g., directly therebetween) using Direct Link Establishment (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Channelized DLS (TDLS). For example, a WLAN using the Standalone BSS (IBSS) mode does not have an access point (AP) and is located within the IBSS or the STAs using the IBSS (e.g., all STAs) can communicate directly with each other. Here, the IBSS communication mode is sometimes referred to as an "ad-hoc" communication mode.

[0108] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel (e.g., the primary channel). The primary channel can have a fixed width (e.g., a 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish connections with the AP. In some representative embodiments, carrier-sense multiple access with collision avoidance (CSMA / CA) can be implemented (e.g., in an 802.11 system). For CSMA / CA, STAs, including the AP (e.g., each STA), can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, then that particular STA can back off. In a given BSS, at any given time, there is only one STA (e.g., only one station) transmitting.

[0109] High-throughput (HT) STAs can communicate using a 40MHz wide channel (e.g., by combining a 20MHz wide main channel with adjacent or non-adjacent 20MHz wide channels to form a 40MHz wide channel).

[0110] Very High Throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels or by combining two non-consecutive 80MHz channels (this combination may be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data is transmitted and passed through a segmented parser that splits the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed individually on each stream. The streams can be mapped onto two 80MHz channels, and the data can be transmitted by the STA performing the transmission. On the receiver of the STA performing the reception, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0111] 802.11af and 802.11ah support sub-1 GHz operating modes. Compared to 802.11n and 802.11ac, 802.11af and 802.11ah utilize reduced channel bandwidth and carriers. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support instrument-type control / machine-type communication, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include a battery with a battery life exceeding a threshold (e.g., for maintaining a very long battery life).

[0112] For WLAN systems that can support multiple channels and channel bandwidths (e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah), these systems include a channel that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a single STA, which is derived from all STAs operating in the BSS supporting the minimum bandwidth operating mode. In the example of 802.11ah, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the width of the primary channel can be 1MHz for STAs that support (e.g., only support) the 1MHz mode (e.g., MTC type devices). Carrier sensing and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy (e.g., because an STA (which only supports the 1MHz operating mode) is transmitting to the AP), then the entire available band can be considered busy even if most of the available band remains idle and available.

[0113] In the United States, the available frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. Depending on the country code, the total bandwidth available for 802.11ah is 6MHz to 26MHz.

[0114] Figure 1D This diagram illustrates a system schematic of RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.

[0115] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. Each of gNBs 180a, 180b, and 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may use beamforming to transmit and / or receive signals to and / or from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and receive radio signals from WTRU 102a. In embodiments, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c may implement Cooperative Multipoint (CoMP) technology. For example, WTRU 102a can receive cooperative transmissions from gNB 180a and gNB 180b (and / or gNB 180c).

[0116] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digital configurations. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can be different for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., containing different numbers of OFDM symbols and / or varying absolute durations).

[0117] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNodeBs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobile anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can use signals in unlicensed frequency bands to communicate with gNBs 180a, 180b, and 180c. In a non-standalone configuration, WTRUs 102a, 102b, and 102c communicate / connect with gNBs 180a, 180b, and 180c simultaneously with other RANs (e.g., eNodeBs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNodeBs 160a, 160b, and 160c, by implementing DC principles. In a non-standalone configuration, eNodeBs 160a, 160b, and 160c can act as mobile anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRUs 102a, 102b, and 102c.

[0118] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support network slicing, dual connectivity, implement interoperability processing between NR and E-UTRA, route user plane data to User Plane Functions (UPF) 184a and 184b, and route control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB180a, 180b, and 180c can communicate with each other via the Xn interface.

[0119] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and may include data network (DN) 185a, 185b. While each of the foregoing components is described as part of CN 115, it should be understood that any of these components may be owned and / or operated by an entity other than a CN operator.

[0120] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different needs), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, and mobility management, etc. AMF 182a and 182b can use network slicing to customize the CN support provided to WTRU 102a, 102b, and 102c based on the service types used by WTRU 102a, 102b, and 102c. As an example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and / or services for Machine-Type Communication (MTC) access, etc. AMF 182 can provide control plane functions for switching between RAN 113 and other RANs (not shown) using other radio technologies (e.g., LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi).

[0121] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and can configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating WTRU or UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications, etc. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0122] UPF 184a and 184b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface, thus providing WTRU 102a, 102b, and 102c with access to a packet-switched network (e.g., Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring, etc.

[0123] CN 115 can facilitate communication with other networks. For example, CN 115 may include or can communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to DNs 185a and 185b via the N3 interface connected to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and local data networks (DNs) 185a and 185b.

[0124] In view of Figure 1A-1D And about Figure 1A-1D The corresponding descriptions herein refer to one or more of the following functions, which can be performed by one or more emulation devices (not shown): WTRU 102a-d, Base Station 114a-b, eNodeB 160a-c, MME 162, SGW164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN185a-b, and / or one or more other devices described herein. These emulation devices can be one or more devices configured to simulate one or more of the functions described herein. For example, these emulation devices can be used to test other devices and / or simulate network and / or WTRU functions.

[0125] The simulation device may be designed to perform one or more tests on other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices may perform one or more functions while being implemented and / or deployed, wholly or partially, as part of a wired and / or wireless communication network, to test other devices within the communication network. The one or more simulation devices may perform one or more functions while being implemented or deployed temporarily as part of a wired and / or wireless communication network. The simulation device may be directly coupled to other devices to perform tests, and / or may use over-the-air wireless communication to perform tests.

[0126] One or more simulation devices can perform one or more functions, including all functionalities, without being implemented or deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used in a test laboratory and / or a test scenario where a wired and / or wireless communication network is not deployed (e.g., under test) to perform tests on one or more components. The one or more simulation devices can be test equipment. The simulation device can transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., the circuitry may include one or more antennas).

[0127] Figure 2 This is a flowchart of an exemplary process used for local VR services. Figure 2 An example of latency in a local VR environment is illustrated, where the user device has a wired connection to a local computer. At 202, information regarding the user's motion and / or head orientation (e.g., inertial measurement unit (IMU) sensor data) can be measured at the user device (also referred to in some places as a "client device"), and at 204, this is transmitted to the local computer (e.g., a desktop computer, for example, via a USB connection). At 206, the local computer can then generate a new VR video frame based on the latest motion information, and at 208, this new VR video frame can be sent to the user device (e.g., typically via an HDMI cable). Because the user's motion may continue to change during this process, when a frame including the new image is displayed to the user on the user device at time 212, the frame received and processed by the user device at 210 (including, for example, pixel switching) may show an FOV different from the user's actual field of view (FOV).

[0128] To reduce motion-to-photon (MTP) latency (e.g., MT delay), some VR systems or devices employ user head prediction methods. The current field of view (FOV) is reported by a sensor that provides the user's current viewing direction. The VR computer then renders an image corresponding to the expected position of the user's head, rather than an image corresponding to the position where the user's head is currently positioned. Some current VR systems use prediction methods based on constant acceleration to predict future head position. However, prediction errors can become significant when the user's motion changes rapidly.

[0129] In remote or cloud-based VR service environments, or in VR service environments where the HMD is wirelessly connected to a local server, latency is typically greater than in local VR service environments using wired connections (e.g., ...). Figure 2 (As shown in the diagram). To illustrate, due to the long distance between the cloud server and the HMD, network latency can be added in both transmission directions. Unlike the example local VR environment where the HMD and local computer can communicate via USB and / or HDMI, in a remote VR service environment, network latency can be more problematic than other types of latency due to some variable characteristics of network latency.

[0130] Figure 3 An exemplary process 300 for remote or cloud-based VR services is shown. For example, Figure 3 The user's head position at time T0(318) is shown. Figure 3 (In Chinese, this is represented as "movement"). Furthermore, such as... Figure 3 As shown, network latency 304 may occur during the transmission of IMU sensor data 302 (e.g., data corresponding to samples taken at time T0) from the HMD (in the form of a client device) to the server via the network. At 306, the server may generate a service frame. In some embodiments, the service frame (or, for example, a service image) refers to VR content provided by the user device (e.g., the HMD, as in this example) for displaying to the user. In some embodiments, the service frame generated by the user device may reflect the user's position and head orientation (e.g., the user's viewpoint). Frame generation may cause processing latency 308, and... Figure 3 The image 314 shows the user's view assuming the user's head remains in the same position after a delay of 308. At 310, the server sends the frame to the client device via the network. This frame transmission causes additional network latency 312. As shown, when the rendered frame is received at the client device, the user's head position (at...) Figure 3The motion (represented as "motion") can change at time T0+ΔT (320). In some embodiments, ΔT represents the amount of time including the sum of network latency 304, processing latency 308, and network latency 312. Therefore, due to the user's head movement, the user's FOV 316 at time T0+ΔT is different from the FOV at time T0, and thus the user cannot see the entire image 314 sent from the server.

[0131] Figure 4 This is a schematic diagram illustrating an example timewarping 400 of a video frame according to some embodiments. Typically, timewarping is a technique that immediately shifts a frame before it is scanned and output to a display in a VR video frame. For example... Figure 4 As illustrated in the example, time warp takes into account the difference between the FOV 402 of the received frame and the current FOV 404 of the user 406 at or near the display time, to produce a shifted frame 408 (shifted by time warp). In some cases, time warp can reduce or prevent motion sickness (e.g., VR motion sickness) that may result from inconsistencies between two FOVs.

[0132] Figure 5 This is a flowchart of an example server-side time warp process 500 according to some embodiments. In some embodiments, server-side time warp can be used in a VR environment where the HMD and server computer are directly connected, for example, via an HDMI cable. Figure 5 As shown, at time T0, the server can perform content simulation and rendering. At time T0+ΔT, the user wearing the HMD can move (e.g., the user can perform head movements at time T0+ΔT, such as...). Figure 5 (As shown). Therefore, at 506, the server can perform time warping and can send the shifted frame to the HMD, for example, via an HDMI cable. The HMD can receive the information via, for example, the HDMI cable, and can display the received information at 508. The received information includes a shifted frame, which is time-warped by the server based on the changed head angle of the user at 504 during or after the rendering of the frame but before the frame is time-warped. Figure 5 As shown, user 504 can move at time T0+ΔT+ε, where ε represents the MTP (motion-to-photon) delay corresponding to the time warp on the server side. In some embodiments, this movement may be disregarded in the frame displayed to the user at the HMD because the latency of the HDMI transmission may be relatively small, and no detectable change may occur, for example, in the user's FOV.

[0133] Figure 6This is a flowchart of an example client-side (e.g., HMD-side) time warp process 600 according to some embodiments. As mentioned above, in some embodiments, time warping performed by the server typically does not reflect additional changes in the user's field of view (FOV) that may occur during encoding and frame transmission from the server. Server-side time warping can therefore result in greater MTP latency than, for example, client-side time warping, and thus in some cases, may cause motion sickness in users in VR service environments with network latency. Therefore, in some embodiments, the client-side time warping may be more appropriate when the client and server are wirelessly connected or, for example, a cloud server is used for content rendering.

[0134] like Figure 6 As shown, at 602, at time T0, the server can perform content simulation and rendering. Figure 6 The diagram shows the head position of user 604 wearing the HMD at time T0. At 606, the server can encode the content, and at 608, a (Tx) (VR video) frame can be sent to the HMD via a communication medium (e.g., a wireless connection). Content processing at the server prior to frame transmission may introduce an additional time amount ΔT, as well as an additional delay ΔTX during the frame's transmission from the server to the client, such as... Figure 6 As shown. During the time interval from time T0 to time T0+ΔT+ΔTX, the user 604 wearing the HMD may move (e.g., the user 604 may perform head movements, resulting in a change in head position at time T0+ΔT+ΔTX, such as...). Figure 6 (As shown). The HMD can receive (Rx) frames at 610 and decode them at 612. At 614, the HMD can perform time warping to generate shifted frames for display to the user. This time warping may take into account (e.g., compensate for) user motion during the time period from time T0 to time T0+ΔT+ΔTX, as measured at the client device (HMD in this example). However, the client's time warping process may incur an additional time amount ΔTP before displaying the shifted (e.g., time-warped) frame at 616. Figure 6 As shown, in some embodiments, user 604 may move during the additional time ΔTP, resulting in a further change in head position at time T0+ΔT+ΔTX+ΔTP. In some embodiments, ΔTX+ΔTP represents the component of MTP delay that remains (or is not compensated for) after server-side time warping. In some embodiments, ΔTP represents the component of MTP delay remaining after client-side time warping (or not compensated for by client-side time warping). Therefore, in some cases, the component of MTP delay associated with network latency (e.g., ΔTX) can be compensated by performing client-side time warping.

[0135] Figure 7 An example of scenario 700, in which a time warp occurs, is shown according to some embodiments. Figure 7 In the example, the rendered image 704 (e.g., showing a portion of the virtual world 702) can have a size substantially the same as the size of the HMD displaying that image. For example, in scene 700 where a loss of FOV 714 occurs after a time warp, the HMD can be configured to display an image with a specific resolution (e.g., an image included in a frame), and that image is rendered (e.g., as rendered image 704) at that specific resolution. As combined above... Figure 4 As described above, in time warp, frames can be shifted based on user movement that occurs at or after the moment the frame is rendered. As shown in the figure, as... Figure 7 In the example, if image 704 has already been based on the predicted scan output time T P (Or, in other words, the predicted scan output timing) the gaze point 706 is rendered, but due to the delay in change, the rendered image 704 is rendered at the actual time T. A (Or, in other words, the actual scan output timing) is scanned and output, then the rendered image 704 is shifted by time warp before being scanned and output, so as to be synchronized with the time T. A The fixation point 710 is aligned. Figure 7 It shows that at time T A The effective FOV of the user at the location is 708 and at time T A The user's FOV 712. In some embodiments, the user's FOV (e.g., the user's FOV 712) refers to the value that should be available at time T. A The image area provided to the user, while the user's effective FOV (e.g., user's effective FOV 708) refers to the image area that the VR system can actually provide to the user. If time T P and T A If the difference between the two is relatively large, and the user wearing the HMD engaged in significant physical activity during that time period, the degree of time distortion shift may also be relatively large. Therefore, in some embodiments, the shifted image may not completely fill the display, and for example, areas of the display without an image may appear black to the user. Figure 7 As shown, the user's effective FOV 708 only includes a portion of FOV 712. This may cause the user to perceive a loss of FOV, for example... Figure 7 The field of view (FOV) is reduced by 714, and the user's immersion may be compromised as a result.

[0136] To minimize FOV loss, typical VR rendering techniques can support overfill, through which the image is rendered larger than the display size. Figure 8 An example of an exemplary overfill process is shown. Figure 8 As shown, the rendered image 804 (showing, for example, a portion of virtual world 802) can be rendered larger than the display size 806, where a tolerance is uniformly added for all directions at the boundaries of image 806. To determine the size of this tolerance, a scaling factor parameter can be defined. Figure 8 This shows the predicted scan output time T. P (Or, in other words, the gaze point 808 of the predicted scan output timing) at actual time T A (Or, in other words, the actual scan output timing) fixation point 810, at time T A The user's effective FOV is 812 and at time T A The user's FOV 814. With overfill tolerance applied, the user's FOV 814 at scan output is more likely to be within the rendered image 804 with overfill compared to when there is no overfill (e.g., ...). Figure 8 (As shown). Therefore, the display can show essentially the time T. A The gaze point 810 is aligned with the rendered image 804 without a black border.

[0137] Overfill factor is a factor that determines how many more pixels to render than the number of pixels the user will eventually see through the service image to prevent FOV loss due to time warp. Some current VR systems typically use user motion information identified by the latest IMU data to render images. Because the image is displayed when the next service frame scan outputs (e.g., VSync (vertical synchronization)), these current VR systems predict the user's later FOV based on the latest IMU data.

[0138] Therefore, in such a system, rendering is performed on the predicted orientation, and the rendering is performed on a rendering area that is as large as the overfill factor (which is the same for both axes) to prevent the FOV loss problem after time warp. Figure 9 An example of overfilling the rendering area is shown. For example... Figure 9 As shown, rendering area 902 (e.g., a portion of virtual world 900 is shown) includes areas rendered by applying, such as... Figure 9 The overfill factor shown is the user FOV 904 extending along both the x and y axes relative to the user's gaze point 906, where R_H represents the vertical resolution (along the height direction) and R_W represents the horizontal resolution (along the width direction).

[0139] Overfilling can reduce or eliminate FOV loss after time warp, but this comes at the cost of increased computational load on the server, for example. An example technique for estimating the computational load used for rendering is to count the number of pixels in the resulting image. Thus, this estimation technique proportionalizes the increased computational load from overfilling to the increased tolerance area during the overfilling process. As the tolerance area increases, the probability of the user experiencing FOV loss decreases, but at the cost of increased computational load. As an example, if the VR system uses an overfill factor of two (2) for a typical example overfilling process, the VR system will render four (4) times the number of pixels to be provided to the user (two (2) times in both the horizontal and vertical directions). In addition to the increased rendering load at the server, the network load for transferring the overfilled image from the server to the client also increases.

[0140] The high price of powerful processors capable of handling the high FPS (frames per second) and resolution (which may increase for dual monitors) of current VR systems makes immersive VR applications inaccessible to the average user. If computational power is consumed during overfilling, the expected performance of devices used in VR applications employing overfilling will be even higher. On the other hand, if the VR system uses a small overfill factor in its overfilling mechanism to reduce the computational power consumed by rendering, a time-warped FOV loss problem may occur.

[0141] Furthermore, typical overfilling usually assumes that service frames will be delivered to the user based on the VSync signal, thus it can be considered a strictly time-based system. For cloud-based VR systems or other VR systems using wireless connections to the HMD, there is a probability that service frames will not be delivered until after the target VSync due to the nature of network latency. In such scenarios, existing overfilling methods may not work well or may use very large overfilling factors.

[0142] According to some embodiments, the systems and methods described herein use a dynamic overfill factor for each of the horizontal and vertical image resolutions. Conversely, in cases such as Figure 9In the typical overfilling methods described herein, the same overfill factor is effectively used for both the horizontal and vertical image resolutions. In some embodiments, the overfill factor is calculated based on both the render-to-display latency distribution and motion error tolerance measured at two representative timing points covering, for example, most latency distributions. In some embodiments, a predictive overfilling technique is employed, wherein FOV loss is minimized or eliminated. Utilizing the various techniques disclosed herein, according to some embodiments, the associated computational load can be reduced compared to some typical overfilling techniques. In some embodiments, a predictive overfilling technique adapted to system latency rather than target time is employed.

[0143] As mentioned above, shifting the rendered VR image in an attempt to align it with the latest foveation point before the scan output can lead to a loss of field of view (FOV) in at least some cases, and this can worsen as latency from VR image processing and network delivery increases. Over-rendering VR images larger than the display size can help reduce this FOV loss, but usually at the cost of increased processing load.

[0144] According to some embodiments, the systems and methods described herein can adjust over-rendered regions based on user head movement information and rendering-to-scan output latency measurements. Some example embodiments expand the rendering region based on the predicted movement path of the gaze point, thereby minimizing the FOV loss problem with potentially reduced computational load. Some example embodiments select a time range within which a rendered image may be scanned and output despite latency variations. This image can then be rendered such that the resulting image covers the FOVs at both ends of the timing range, the FOVs including various prediction error tolerances to generate a potential FOV. Since the prediction errors at both ends of the range will differ (e.g., errors typically increase with longer expected time), the example embodiments apply different error tolerances to each FOV. Note that in some embodiments, the expected time represents the difference between the current time point and a future time point on which the prediction is performed. Performing a prediction at a future time point may result in prediction errors, which typically increase when predictions are made at more distant future time points. In some cases, the prediction error increases as the time point associated with the prediction moves further away from the capture time of the data on which the prediction is based.

[0145] Note that in the various embodiments disclosed herein, according to some embodiments, the HMD and / or other client devices associated with the user may be implemented as one or more WTRUs as described herein.

[0146] Figure 10 An example virtual reality environment 1000 according to some embodiments is shown. Figure 10In some embodiments, a portion of the virtual reality environment 1000 utilizes the predicted overfill described herein for rendering. In some embodiments, an overfilled image, such as image 1012 (rendered image 1012), is rendered using predicted overfill, which may be overfilled in a directional manner (e.g., the overfill is not applied uniformly in all directions). For illustration, just before rendering the image, for example, the server may predict the user's future gaze point at two time points T1 and T2. Figure 10 The following are shown: (i) the predicted fixation point 1002 at time T1 and the associated predicted FOV 1008 at time T1; and (ii) the predicted fixation point 1004 at time T2 and the associated predicted FOV 1010 at time T2. Furthermore, Figure 10 The trajectory 1006 of the gaze point is illustrated. In some embodiments, T1 provides a lower limit for the expected scan output time, and T2 provides an upper limit for the expected scan output time. In some embodiments, T1 and T2 are selected based on the assumption that the actual (e.g., real) scan output time may be within the range of [T1, T2]. Furthermore, in some embodiments, T1 may be based on the corresponding predicted minimum scan output delay, and T2 may be based on the corresponding predicted maximum scan output delay. In some embodiments, the rendered image 1012 is sized such that it completely covers the predicted FOV 1008 at the gaze point at time T1 and the predicted FOV 1010 at the gaze point at time T2.

[0147] In some embodiments, as will be described in detail, a potential FOV (e.g., a combined FOV) 1018 can be formed based on the FOV 1008 at time T1, the FOV 1010 at time T2, the error tolerance 1014 for T1, and the error tolerance 1016 for T2. In this regard, in some embodiments, the midpoint (or center / center point) of the potential FOV 1018 of the rendering region can be identified. Figure 10 As shown, the potential FOV 1018 may include FOV 1008 for T1 and FOV 1010 for T2. For example, by comparing (i) the number of pixels from the center point of (potential FOV 1018) to the horizontal and vertical (width and height) boundaries of the potential FOV 1018 with (ii) the resolution of the service frame to be presented to the user, a separate overfill factor can be determined for each of the horizontal and vertical axes.

[0148] Figure 11 This is a flowchart of an exemplary process 1100 for predictive overfilling according to some embodiments. This exemplary process 1100 can be implemented by a remote computer such as server 1102 and a user device such as HMD 1104. Figure 11An example method for obtaining times T1 and T2 (hereinafter referred to as "T1" and "T2" for brevity) is shown. At 1106, server 1102 can track the time when it begins rendering a service frame and the time spent scanning that service frame out to the user. To perform this tracking, in some embodiments, server 1102 can receive scan output timing information from, for example, HMD 1104 at 1108. For example, server 1102 can receive feedback from HMD 1104 regarding the time of scanning out the service frame. In some embodiments, this feedback can be feedback other than that the server can receive from HMD 1104, such as, for example, user motion information (1110) including head tracking position information. In some embodiments, the user motion information can include IMU data collected by performing IMU processing 1122 at HMD 1104. In some embodiments, at 1106, server 1102 determines the delay distribution (also... Figure 11 (Illustrated graphically). This latency distribution can be based on the start rendering time (e.g., determined by server 1102) and scan output time (e.g., received by server 1102 from HMD 1104) of the frames rendered by server 1102 and scanned out by HMD 1106. In some embodiments, the determined latency distribution will have a specific distribution due to the time spent rendering and the distribution of network latency.

[0149] at 1112 (also Figure 11 (Figured in the image) The choice of T1 and T2 (e.g., the updated choice if a previous choice has been made) can be based on the start time of rendering the video frame at the server and a determined delay distribution. For example, T1 and T2 can be chosen such that the probability that the scan output time of the video frame at the client will fall between times T1 and T2 is higher than a threshold. T1 can be the start time of rendering at the server plus a lower limit for the rendering-to-scan output delay, and T2 can be the start time of rendering at the server plus an upper limit for the rendering-to-scan output delay. This is in... Figure 11 The diagram is illustrated in the figure. The choice of T1 and T2 can be used to change the degree of overfilling applied to the prediction. In other words, the overfilling performance can be adjusted by changing the interval between T1 and T2. For example, in some embodiments, as the time interval between T1 and T2 becomes longer, the overfilling performance increases (e.g., the probability of FOV loss decreases), but the processing and data overhead increases because the overfilling factor increases. In 1114, based on T1 and T2 (also... Figure 11The various predicted FOVs (graphically shown in the figure) are used to determine the potential FOV and to determine (e.g., set) various overfill factors, as described in more detail later. On average, the error in predicting the gaze point of T2 will be greater than the error in predicting the gaze point of T1. Therefore, in some embodiments, corresponding error tolerances are applied to the predicted FOVs of T1 and T2 (e.g., where a larger tolerance can be applied to T2). At 1116, server 1102 can simulate content and render VR video frames including overfilled images based on (i) the user's predicted future head position (e.g., which is determined based on user motion information (1110) received from HMD 1104 to predict the potential FOV) and the overfill factor determined based on the expected error in the user's predicted future head position. As a result, in some embodiments, the rendering region (with the overfill factor applied) is a region that includes the prediction error tolerance for each predicted gaze point and the distance between the expected gaze points of T1 and T2. More specifically, when determining the rendering area, in some embodiments, the VR system (e.g., server 1102 in this case) includes the field of view (FOV) at all time points between T1 and T2, which is resolved by determining a rendering area that includes the FOV of T1 and T2.

[0150] At 1118, the rendered VR video frame can be sent as part of the VR video stream to HMD 1104. At HMD 1104, and at 1120, HMD 1120 can perform time warping on the received frame in some embodiments, and then scan the frame and output it to the user.

[0151] The exemplary embodiments of this disclosure can reduce the size of images rendered using typical overfill, thereby reducing computational load and also reducing or eliminating FOV loss. The exemplary embodiments can adjust the overfill factor based on the user's head movement characteristics. For example, in some embodiments, because the overfill factor can be increased only for the direction of the user's head rotation (e.g., the vertical overfill factor is close to 1 when the user rotates their head horizontally and maintains it in the same or nearly the same vertical position), content generation and / or transmission can be more efficient.

[0152] In some scenarios, significant latency in VR is caused by the time required to render high-resolution images. For example, in some embodiments, for a computer used as a server, reducing the computational load on rendering images, all else being equal, will result in reduced rendering time, thereby reducing VR latency. Furthermore, in some embodiments, because the process for determining T1 and T2 takes into account the latency of the VR system, it adaptively addresses the FOV loss problem even when network latency changes.

[0153] Furthermore, the example embodiments employing predictive overfilling disclosed herein use less processing power than some omnidirectional overfilling techniques. In this regard, in some embodiments, the proposed method using [T1, T2] renders only the potential FOV, which in some embodiments may be a region merged from the FOVs predicted at T1 and T2.

[0154] Figure 12 This is a message passing diagram of an exemplary process 1200 according to some embodiments. As shown, this exemplary process is executed between a VR content server 1202 and a user (client) device 1204 (e.g., an HMD or another user device), and can be executed iteratively (e.g., frame-by-frame, periodically only, such as every given number of frames, etc.). In some embodiments, Figure 12 Beginning at 1206, the VR content server 1202 renders a new service frame (e.g., a new overfilled frame). In some embodiments, once the overfill factor is set, the exemplary process described below in a more comprehensive manner can be repeated. For example, the exemplary process can be repeated while VR content is being streamed. In some embodiments, the process can be repeated continuously, such that multiple steps of the process are performed simultaneously at any given point in time. Other options are also possible.

[0155] Referring back to reference 1206, the VR content server 1202 can record the start time (T) of rendering the new service frame. R At 1208, the generated (rendered) service frames (e.g., in the form of service packets) can be sent to user device 1204. For example, the transmission of rendered frames from VR content server 1202 to user device 1204 may include network latency (e.g., packetization or packet delay).

[0156] In some embodiments, at 1210, user equipment 1204 may perform time warping on frames received from VR content server 1202, and scan out the service frames at 1212. User equipment 1204 may also determine timing information to be provided to server 1202 as feedback. In this regard, in some embodiments, user equipment 1204 may record the scan output start time (T0) of the received frames. S Then, at 1214, user equipment 1204 can, for example, via timing information feedback processing, transmit the scan output start time T. S Send to VR content server 1202.

[0157] In some embodiments, at 1216, a render-to-scan output delay distribution is determined or refreshed. The render-to-scan output delay of a frame can be determined, for example, by means of the received T... SSubtract the recorded T from the middle R To measure (therefore, the rendering to scan output delay = T) S -T R In some embodiments, T S and T R Synchronized, or T adjusted according to time offset S and T R One or both of these. (This could be, for example, the clock difference between the server and the client). Latency measured for frames can be added to the table. Figure 13 An example delay table management process according to some embodiments is shown.

[0158] Latency can be measured for some or all frames. For example... Figure 13 As shown, for example, the size of table 1300 can be maintained by replacing older delay data 1304 with updated data samples, such as new delay data 1302. For example, as... Figure 13 As shown, when new delay data 1302 becomes available, older delay data 1304 can be deleted from table 1300, and new delay data 1302 can be added to table 1300. The distribution of delay values ​​stored in the delay table 1300 can be graphically represented. Figure 14 This is a schematic diagram illustrating an exemplary delay distribution probability 1400 according to some embodiments.

[0159] In some embodiments, at 1218, time intervals [T1, T2] are selected. T1 and T2 may be, for example, the result of estimating the rendering-to-scanning output delay of the VR system. In some embodiments, T1 and T2 are selected based on the recorded delay distribution. Figure 15 This is a graph illustrating the probability distribution of example delays, including example time intervals, according to some embodiments. For example... Figure 15 As shown, various time intervals can have corresponding probabilities (e.g., 80% probability and 95% probability, as illustrated). In some embodiments, the time interval between T1 and T2 can be selected to include a target probability of the delay distribution. If [T1, T2] is selected as, for example... Figure 15 As shown in T1_95 and T2_95, the probability of a service frame scan being output to the user within the time interval T1 to T2 will be 95%. If an 80% probability distribution is included, then [T1, T2] will be a smaller interval (smaller than for 95%), such as T1_80 and T2_80. In some embodiments, a sample of a certain percentage X in the delay table 1300 exists between T1 and T2. According to some embodiments, the percentage X can be determined by the stability of the prediction overfill technique disclosed herein. Note that... Figure 15The interval selection shown in the figure effectively selects the lower and upper limits of the predicted render-to-scan output delay values, which can then be converted into times T1 and T2 (the lower and upper limits of the predicted scan output time of the current VR video frame) by adding the server rendering start time of the current VR video frame to the determined delay limits.

[0160] The reason for determining the time interval [T1, T2] based on the probability between these two times, rather than setting the interval to some arbitrary "wide enough" value, is likely that the overhead of overfilling (e.g., from processing and data volume) increases with the size of the time interval. To illustrate, as the time interval between T1 and T2 increases, predicting overfilling requires rendering a wider range of the virtual world. This phenomenon is described in more detail below.

[0161] In some embodiments, if the system's rendering-to-scan output delay is stable (e.g., a narrow distribution of delay), the interval between T1 and T2 will remain relatively shorter than that of an unstable system (e.g., with substantially the same performance). Thus, even if T1 and T2 (as chosen using a probabilistic method) are relatively close together, the probability that the actual delay will fall within the narrow time interval between them is high.

[0162] In some embodiments, at 1222, FOV is predicted. In some embodiments, FOV for T1 and FOV for T2 are predicted using the latest motion information provided, for example, by user device 1204 at 1220. In some embodiments, the method for predicting user FOV may employ current IMU data-based prediction methods. If it is assumed that the user's motion is maintained, the location of the user's gaze point in the future (e.g., at time T1 and at time T2) can be predicted from the gaze point (e.g., head orientation) associated with the latest IMU feedback (e.g., the latest user motion information provided by user device 1204). In some embodiments, the user's FOV is a frame-size space centered on the gaze point.

[0163] Figure 16 This is a schematic diagram illustrating an example of predicting gaze points based on some embodiments. Figure 16 The gaze point within the 1600-degree virtual world section is shown. Figure 16The diagram illustrates the current or current user's gaze point 1602, the predicted user gaze point 1604 at time T1, and the predicted user gaze point 1606 at time T2. The current user's gaze point 1602 is indicated by arrow 1608 (dashed line), representing observed user motion. In some embodiments, the user motion is observed, for example, using the IMU of a user-worn HMD. The observed user motion may include data indicating the user's direction and velocity, rate, and / or acceleration. The predicted user gaze point 1604 at T1 may have a corresponding prediction error range 1610, and the predicted user gaze point 1606 at T2 may have a corresponding prediction error range 1612, both of which are within... Figure 16 As shown in the figure. In some embodiments, the range of each prediction error can be determined based on the type of prediction technique used.

[0164] In some embodiments, at 1224, a potential FOV (e.g., a rendering region) is determined. For example, the potential FOV (e.g., as a combined FOV) can be determined based on the predicted overfill by combining two separate predicted FOVs (in this case, the FOV of T1 and the FOV of T2). However, the foveation points predicted for T1 and T2 may have errors, possibly due to errors in the prediction method used itself. Therefore, in some embodiments, when determining the potential FOV (e.g., the rendering region), a corresponding error tolerance for each predicted FOV (e.g., added to each respective predicted FOV) can be used. In this respect, in some embodiments, the potential FOV is determined by combining (also referred to herein as "merging") (i) a first adjusted predicted FOV of T1 (wherein the first adjusted FOV is determined by adding a first error tolerance to the predicted FOV of T1) and (ii) a second adjusted predicted FOV (wherein the second adjusted FOV is determined by adding a second error tolerance to the predicted FOV of T2). The first and second error tolerances may be different to reflect different expected errors associated with the predicted FOVs at T1 and T2. For example, the second error tolerance may be greater than the first error tolerance to reflect a larger expected error when predicting the FOV at a future time. Furthermore, as described above, in some embodiments, combining two predicted FOVs at two corresponding time points may include forming the FOVs at all time points between T1 and T2, which is addressed by determining a rendering region that includes the FOVs of T1 and T2. An example of merging two predicted FOVs will be described in more detail later.

[0165] Experiments were conducted using two prediction techniques currently used in some VR systems: (1) prediction based on constant rate (velocity) (CRP); and (2) prediction based on constant acceleration (CAP). Figure 17A and 17BThese are two graphs, 1700 and 1750, showing experimental results using two prediction techniques based on some embodiments. More specifically, Figure 17A and 17B The experimental results illustrating the prediction errors of CRP and CAP prediction techniques over time are illustrated graphically. The experiment was conducted by rotating the head in the yaw direction using an IMU sensor attached to the user's head, and comparing the measured velocity and acceleration with the predicted values ​​(using CAP and CRP prediction techniques). Figure 17A The graph shows the MSE of orientation (in degrees). 2 The relationship between time (in milliseconds (ms)) and expected time. Figure 17B The graph shows the relationship between the maximum orientation error (in degrees) and the expected time.

[0166] The VR content (streaming) server 1202 can obtain error graphs for one or more individual VR users in real time. In some embodiments, the server 1202 may do the following: (1) the server receives internal IMU motion data of the VR HMD at a certain time; (2) the server calculates future IMU motion values ​​using CAP and / or CRP prediction methods; (3) the server then receives actual IMU motion data, records the error between the predicted motion data and the actual motion data, and updates the average and / or maximum error graphs in real time.

[0167]

[0168] Table 1: Example of Real-Time Prediction Error Verification

[0169] As shown in Table 1 above, the technique for identifying the real-time prediction error can use the identified IMU values ​​to predict future head orientation.

[0170] The CRP technique used in the example in Table 1 assumes that the user's head rotation speed (0.04 degrees / ms) will be maintained, and predicts the user's head orientation (prediction(t)) based on the current orientation (99.97 degrees). The system then compares the received IMU feedback data with the predicted value in real time, confirming the prediction error of the prediction method according to the expected time. The information obtained regarding the prediction error (e.g., continuously in real time) is accumulated (e.g., averaged using the most recent 100 values) or processed in different ways to update the error table.

[0171] In some embodiments, the updated error table allows the system to continuously update to calibrate the values ​​used for error tolerance (for overfilling) during VR content playback. In the embodiments described above, predictions and error assessments are given from 1 to 6 ms. In some embodiments, the system assesses the error of the target delay interval (e.g., error assessment may be performed for each of the yaw and pitch axes).

[0172] As the experimental results above show, the longer the predicted target time, the greater the prediction error tends to be. In some embodiments, the rendering area is defined to include a lower error tolerance for relatively near predicted scan output times (e.g., a lower error tolerance added for the FOV corresponding to the lower limit T1) and a larger error tolerance for relatively far predicted scan output times (e.g., a larger error tolerance added for the FOV corresponding to the upper limit T2). In some embodiments, as described above, errors in the prediction method are identified empirically during service, thus it is feasible to apply yaw prediction errors and pitch prediction errors that reflect individual characteristics of the user's head movement.

[0173] Figure 18 Illustrations based on some embodiments Figure 17A The graph shows the predicted FOV and an example. More specifically, Figure 18 Two example time values ​​for T1 and T2 derived from the graph are shown (where T1 = 70 ms and T2 = 120 ms), and the corresponding predicted FOVs 1810 and 1812 for two fixations at T1 and T2, i.e., predicted fixation point 1802 at T1 and predicted fixation point 1806 at T2. Furthermore, Figure 18 Predicted FOVs 1814 and 1816 are shown, with corresponding error tolerances 1804 and 1808 added thereto to produce a first adjusted FOV 1810 and a second adjusted FOV 1812. Specifically, an error tolerance of 1804 in the form of 7.8% over-full fill is added to predicted FOV 1814, and an error tolerance of 1808 in the form of 20% over-full fill is added to predicted FOV 1816. Figure 18 As shown, the corresponding overfill amount associated with T1 and T2 can be determined from the MSE curve.

[0174] Figure 19 Exemplary error tolerance configurations according to some embodiments are shown. Figure 19 As shown, the pitch error tolerance is 1902 (for example, as...). Figure 19 The upper and lower error tolerances in the middle) and the yaw error tolerance 1904 (for example, such as Figure 19 The left and right error tolerances in the FOV 1900 are added. Figure 19 As shown, the error tolerances are different from each other.

[0175] Figure 20 This is a perspective view of a VR HMD 2000 according to some embodiments. Figure 20 The diagram illustrates pitch, yaw, and roll motion in a coordinate system. In some embodiments, the VR HMD hardware may include multiple microelectromechanical systems (MEMS) or other sensors, such as gyroscopes, accelerometers, and magnetometers. Furthermore, in some embodiments, the HMD 2000 may include sensors for tracking the position of the head-mounted device. Information from each of these sensors can be combined through a sensor fusion process to determine the user's head movement in the real world and synchronize the user's view in real time. In some embodiments, such as... Figure 20 As shown, the coordinate system uses the following conventions: the x-axis is positive to the right; the y-axis is positive upwards; and the z-axis is positive backwards.

[0176] In some embodiments, rotation is maintained as a unit quaternion, but it can also be reported in pitch-yaw-roll form. When viewed from the negative direction of each axis, positive rotation is counter-clockwise (CCW). Figure 20 (The direction of the rotating arrows in the diagram). Pitch is the rotation about the x-axis, which has a positive value when looking upwards. Yaw is the rotation about the y-axis, which has a positive value when turning left. Roll is the rotation about the z-axis, which has a positive value when tilting left in the XY plane.

[0177] In some embodiments, two or more FOVs may be combined to determine a potential FOV (e.g., a combined FOV), which in some embodiments corresponds to a rendering region. Figure 21A An example rendered region 2102 is shown according to some embodiments. In some embodiments, an example overfill prediction technique is used according to the example described herein. Figure 21A The rendering region 2102 in the image corresponds to the potential field of view (FOV). For example... Figure 21A As shown, the potential FOV is generated by combining (i) FOV 2104 (which is centered at the predicted gaze point 2108 at time T1 and has a first corresponding error tolerance added thereto to produce a first adjusted FOV) and (ii) FOV 2106 (which is centered at the predicted gaze point 2110 at time T2 and has a second corresponding error tolerance added thereto to produce a second adjusted FOV). At this point, for example, a rectangular region containing both FOVs can be selected to generate the rendering region 2102, such as... Figure 21A As shown.

[0178] Figure 21B Another example of a rendering region 2102 according to some embodiments is shown. Similar to Figure 21B In some embodiments, the overfill prediction technique is based on the examples described herein. Figure 21BThe rendering region 2102 in the image corresponds to the potential field of view (FOV). For example... Figure 21B As shown, the potential FOV is generated by combining (i) FOV 2114 (which centers at time T1 on the predicted fixation point 2118 and has a first corresponding error tolerance added thereto to produce a first adjusted FOV) and (ii) FOV 2116 (which centers at time T2 on the predicted fixation point 2120 and has a second corresponding error tolerance added thereto to produce a second adjusted FOV). However, with Figure 21A Unlike other methods, these two FOVs can be combined in a hexagonal manner. That is, in this example, the hexagonal region / shape can be selected as a more economical area to contain both FOVs. In some embodiments, a more economical area refers to a smaller rendering area that still captures both FOVs. Generally, the larger the rendering area used to capture both FOVs, the greater the overhead of the VR system (e.g., the overhead associated with processing and transmission). Therefore, from the perspective of the VR system, making the rendering area smaller is more economical.

[0179] on the contrary, Figure 21C The example shows a rendering area 2100 that is overfilled (e.g., typical). Figure 21C The rendering region includes a field of view (FOV) of 2124 centered at the gaze point 2126 at a given time T. Furthermore, as... Figure 21C As shown, overfill is applied uniformly to FOV 2124 in all directions (e.g., by using the same overfill factor for each axis) without considering the user's future head position. This differs from embodiments of this disclosure, which, among other factors, consider the user's predicted future head position for dynamic or adaptive overfill determination.

[0180] Figure 22 This is a more detailed illustration of the potential FOV formation according to some embodiments. (See diagram for example.) Figure 22 As shown, the potential FOV 2200 is generated by combining (i) the predicted FOV 2202 at time T1 (which is supplemented with a first corresponding error tolerance 2208 of T1) (a first adjusted FOV) and (ii) the predicted FOV 2204 at time T2 (which is supplemented with a second corresponding error tolerance 2210 of T2) (a second adjusted FOV). Furthermore, Figure 22 Arrow 2206 is shown indicating the direction of the user's head rotation to show the displacement of the second predicted FOV (2204) along that direction. Figure 22 In the example, as described above, the two predicted FOVs (and their corresponding error tolerances) are combined in a rectangular manner, where a rectangular region containing the two FOVs can be selected to produce rendering region 2102, such as... Figure 21A As shown.

[0181] Return to reference Figure 12 In some embodiments, an overfill factor setting is determined at 1226. In some embodiments, different overfill factors are determined for each of the horizontal and vertical axes. Typically, in some embodiments, the overfill factor for each axis can be determined by comparing the number of pixels present in the potential FOV with the resolution to be provided to the user device 1204. In some embodiments, the ratio of each axis of the potential FOV determined above differs from, for example, the ratio of each axis of the resolution of the HMD. Different overfill factors (which may also be referred to herein as "overfill factor values") can be applied for each axis because, in some embodiments as described above, the potential FOV is determined taking into account or based on the user's head rotation direction.

[0182] To calculate the overfill factor of the horizontal axis, in some embodiments, the following example equation 1 is used:

[0183] Overfill factor = [(max(HEM)] T1 HEM T2 -abs(V×(T2-T1)×sinx))+HEM T2 Equation 1

[0184] abs(V×(T2-T1)×sinx)+MVA horizontal )] / MVA horizontal

[0185] To calculate the overfill factor of the vertical axis, in some embodiments, the following example Equation 2 is used.

[0186] Overfill factor = [(max(VEM)] T1 VEM T2 -abs(V×(T2-T1)×cosx))+VEM T2 + Equation 2

[0187] abs(V×(T2-T1)×cosx)+MVA vertical )] / MVA vertical

[0188] In Equation 1, the variable HEM T1 and HEM T2 These represent the error tolerances for time T1 and time T2, respectively, for the yaw (horizontal) axis. Variable V represents the user's head rotation speed (e.g., in degrees per second). Variable x represents the user's head rotation direction (e.g., zero at the 12 o'clock position). Additionally, variable MVA... horizontal This indicates the monocular horizontal field of view of the service frame.

[0189] In Equation 2, the variable VEM T1 VEM T2 These represent the error tolerances for time T1 and time T2 on the pitch (vertical) axis, respectively. Variable V represents the user's head rotation speed (e.g., in degrees per second). Variable x represents the user's head rotation direction (e.g., zero at the 12 o'clock position). Additionally, variable MVA... vertical It is the monocular vertical viewing angle of the service frame.

[0190] In some embodiments, when the time difference between T1 and T2 is relatively small (therefore, for example, there is no significant change in the FOV between T1 and T2), the max function described above, represented by equations 1 and 2, will return EM. T2 –abs(V×(T2–T1)×sinx), for example, Figures 21A-21B As shown in the example.

[0191] Below are example calculations of the overfill factor (or first overfill factor value) for the horizontal axis and the overfill factor (or second overfill factor value) for the vertical axis. In this example, the prediction errors for both axes are essentially the same. Figure 23 This is a graphical representation of the use of example values ​​calculated from the overfill factor according to some embodiments. More specifically, Figure 23 The values ​​included in the example calculations using the overfill factor in Equations 1 and 2 are illustrated graphically. Note that... Figure 23 All values ​​shown are assumed to be in degrees.

[0192] User's head rotation speed = 100 degrees / second

[0193] The user's head rotation direction is 30 degrees to the right and down (x = 120 degrees).

[0194] The horizontal and vertical prediction error for T1 (e.g., T1 = 70 ms, as described in the example above) is 7.8 degrees.

[0195] The horizontal and vertical prediction error for T2 (e.g., T2 = 120 ms, as described in the example above) is 20 degrees.

[0196] The horizontal field of view of a single eye in an HMD (e.g., according to HMD specifications) is 90 degrees.

[0197] The vertical field of view of a single eye in an HMD (e.g., according to HMD specifications) is 96.73 degrees.

[0198] Overfill factor (level) = [(max(7.8,20-abs(100×(0.05)×sin(120)))+20+abs(100×(0.05)×sin(120))+90] / 90

[0199] Overfill factor (horizontal) = (15.7 + 20 + 4.3 + 90) / 90 ≈ 1.44 Overfill factor (vertical) = (max(7.8, 20 - abs(100 × (0.05) × cos(120))) + 20 + abs(100 × (0.05) × cos(120)) + 96.73 / 96.73

[0200] Overfill factor (vertical) = (17.5 + 20 + 2.5 + 96.73) / 96.73 ≈ 1.41

[0201] Figure 23 The diagram illustrates predicted fixation points 2302 and 2304 for T1, predicted FOV 2306 for T1, predicted FOV 2308 for T2, a potential FOV 2310 without error tolerance, and a potential FOV 2312 with added error tolerance. The potential FOV 2312 is determined based on the predicted FOVs 2306 and 2308 and their associated corresponding error tolerances. Figure 23 As shown, in some embodiments, elements 2310 and 2312 are monocular.

[0202] In some embodiments, the overfill factor (for each axis) determined above is used for rendering along the center of the potential FOV (the point where the centers of each axis intersect), such as... Figure 23 As in the example.

[0203] In some embodiments, the predictive overfill technique proposed herein can effectively reduce the loss of processing power typically required for omnidirectional overfill. Figure 24 This is a message passing diagram illustrating another exemplary process 2400 for predicting overfill according to some embodiments. This exemplary process 2400 can be executed between a VR content server 2402 and a user (client) device 2404 (e.g., an HMD). Figure 24 The exemplary process shown utilizes the predicted scan output time T and, for example, renders only the region that extends accordingly based on the user's head rotation direction from the predicted FOV, rather than performing typical omnidirectional overfill.

[0204] In some embodiments, Figure 24 Steps 2406-2416 of the exemplary process can be combined with the above. Figure 12 Steps 1206-1216 of the exemplary process are essentially the same. As described above... Figure 12 As described, and applicable to Figure 24The exemplary process described herein can be repeatedly performed while VR is being served. In some implementations, the process can be repeated continuously, such that multiple steps of the process are performed simultaneously at any given point in time.

[0205] In some embodiments, at 2418, a time T representing the predicted scan output time is determined. In some embodiments, the rendering-to-scan output latency of the VR system can be predicted with reference to a recorded past latency distribution. In some embodiments, the system can set the predicted scan output time T to correspond to the center point of the latency distribution (e.g., at T, the latency distribution is divided into 50:50, or generally in half). Therefore, the VR content server 2402 can, for example, select the median or average of one or more latency distributions as a representative rendering-to-scan output latency value and can add the rendering start time of a specific VR video frame to convert the representative latency value into a time value T representing the predicted scan output time of the VR video frame. In some embodiments, latency variation characteristics are not considered. Therefore, the exemplary process 2400 may be more suitable for a given VR system where the latency distribution is stable near the representative latency value (e.g., near the value of T).

[0206] Figure 25 This is a schematic diagram illustrating an exemplary delay distribution 2500 according to some embodiments. For example... Figure 25 As shown, it is possible to predict and Figure 25 The approximate center point of the delay distribution map corresponds to a representative render-to-scan output delay 2502 (showing probability relative to time). The predicted scan output time T can be calculated by adding the render start time of the VR video frame (e.g., server-side render start time) to the representative delay value, such that time T corresponds to the scan output time associated with the representative delay 2502.

[0207] In some embodiments, at 2422, an overfill factor setting is determined. Typically, in some embodiments, the overfill factor is determined taking into account the user's head rotation speed, which can be determined, for example, based on user motion information fed back from user device 2404 to server 2402 at 2420 (e.g., ...). Figure 24 (As shown).

[0208] Figure 26 This is a schematic diagram illustrating multiple fields of view (FOVs) according to some embodiments. That is, Figure 26 The diagram illustrates a predicted FOV 2608 for time T, an expanded FOV 2610 for time T, and a final rendering region 2604 for time T, according to some embodiments. In some embodiments, the direction and speed of the user's head rotation are taken into account (in... Figure 26The rendering area 2608, extended by arrow 2602 (represented by the "user's head rotation direction") (e.g., in vector form), includes the predicted FOV of the HMD (and / or the user) at time T. In some embodiments, this is achieved by considering variations in the system's computational capabilities or the recorded latency distribution (e.g., as shown in the image). Figure 25 (As shown) to determine the degree of expansion. The first step in expanding the FOV can be to expand the predicted FOV 2608 based on the user's head / HMD orientation and velocity, as shown. Figure 26 As shown on the left. Figure 26 The diagram illustrates how the expansion can be performed using, for example, a value 2606 (“+ / - user’s head rotation direction / 2”) expressed as a vector. The center position of the expanded FOV 2610 can then be aligned (e.g., calibrated to) with the center position of the predicted FOV 2608 at time T. Horizontal and vertical error tolerances for the FOV can be added to the expanded FOV 2610 to determine the final expanded FOV at time T. Figure 26 The final rendering area 2604 is shown as time T. Figure 26 (See right side) shows how to add pitch error tolerance 2612 (e.g., as...) Figure 26 The upper and lower error tolerances in the middle) and the yaw error tolerance 2614 (e.g., as Figure 26 (Left and right error tolerances in the text). Figure 26 As shown, the error tolerances 2612 and 2614 are different from each other.

[0209] Figure 27 The illustration is a schematic diagram illustrating the determination of the overfill factor according to some embodiments. Figure 27 The predicted gaze point 2702, the predicted FOV 2704, and the rendering region 2708 for T are shown. Furthermore, as... Figure 27 As shown, arrow 2710 represents, for example, a value expressed as a vector: "+ / - user head rotation direction / 2". Figure 27 As shown, in some embodiments, the rendering region 2708 is monocular. Furthermore, in some embodiments, the overfill factor can be calculated according to the following equation:

[0210] Overfill factor (X) = RA x / MVA x Equation 3, where RA x Indicates the rendering angle, while MVA x This represents the monocular viewpoint. Note that Equation 3 applies to calculating the overfill factor in both the horizontal and vertical directions (or the horizontal and vertical axes), where X represents the horizontal or vertical direction.

[0211] In some embodiments, feedback, such as from the client device, is used to shape the render-to-scan output latency distribution. For example, the render-to-scan output latency distribution may be unstable shortly after or immediately after the service begins between the server and the client device. For a stable system, a stable distribution can be formed as the service continues. Figure 28 This is a schematic diagram illustrating an exemplary delay distribution 2800. Figure 28 This illustrates how the latency distribution (e.g., the recorded latency distribution) changes during various phases of a VR service, according to some embodiments. As an example, Figure 28 The delay distributions for the early stage, the middle stage, and the late stage are shown.

[0212] As discussed in conjunction with the various embodiments above, latency values ​​can be recorded in a latency table. In some embodiments, the more data accumulated in the latency table, the more stable the probability distribution may become. However, in some embodiments, the problem is that if the system uses a large latency table, it may not be able to adapt quickly enough when the service environment changes. For example, if the processing of the VR content server is slow due to heat, or if transmission delays occur in a system that includes network latency, a large latency table will have different probability distributions depending on the characteristics of the changed system. Therefore, in some embodiments, the VR system may need to maintain an appropriate number of samples in the latency table to observe and react to changes in latency (e.g., network latency).

[0213] As described above, in some embodiments, times T1 and T2 can be selected by considering a trade-off between overfill stability and rendering overhead. Furthermore, in some embodiments, the interval between the selected times T1 and T2 is a factor for determining the overfill factor (e.g., when the user moves).

[0214] According to some embodiments, one or more options may be available for adjusting the overfill stability available for selecting T1, T2, which may be done, for example, by taking into account the processor’s computational power tolerance.

[0215] As mentioned above, predicting user behavior (or gaze point) may have varying errors, depending on factors such as the technique used for prediction and / or the target time interval. These prediction errors (one or more) can be applied as tolerances (one or more) to determine the rendering area.

[0216] As described above, the example prediction error has been determined experimentally. Some embodiments determine the prediction error and check it in real time during service. For example, using IMU data that can be fed back at, for example, regular intervals (e.g., every 1ms), a VR content server can evaluate one or more errors of the prediction method(s) it is using.

[0217] In some embodiments, the VR content server sets an error update interval and predicts the user's gaze path based on a timeline, from the minimum render to scan output delay to the maximum render to scan output delay (which is recorded data in a delay table). Subsequently, in some embodiments, for example every 1 ms (1000Hz feedback), the user's actual gaze point identified from the feedback IMU data is compared with the predicted gaze point to check for errors in the prediction method over time.

[0218] In some cases, depending on the graphics card's rendering method, the various embodiments presented herein can be applied more effectively. Typically, when using a device that only supports rectangular shape rendering, a rectangular model (e.g., the combination of the above) is used. Figure 21A As described above). When supporting various types of rendering, a more efficient potential FOV can be utilized (such as via hexagonal methods (e.g., the combination of the above)). Figure 21B (The FOV obtained as described). In this case, two or more overfill factors can be set and utilized.

[0219] Furthermore, the various embodiments described herein allow for great flexibility in the size of the output frames, depending on system latency characteristics (e.g., stability) and / or the user's head rotation speed.

[0220] Figure 29A This is a schematic diagram illustrating an example output VR frame that takes into account changes in system latency. As shown, as the system latency 2900 changes (e.g., from a steady state to an unstable state), including exemplary typical overfilling (e.g., as...), Figure 29A As shown, the frame size (and therefore shape) of the output VR frame 2902 with an overfill factor of two (2) remains the same (e.g., with a resolution of 4320x2400, as shown). In some embodiments, this resolution may refer to the pixel dimension (the number of pixels in the horizontal direction (width) × the number of pixels in the vertical direction (height)).

[0221] on the contrary, Figure 29B This is a schematic diagram illustrating the effect of system latency on the size of the output VR frame according to some embodiments. Figure 29BThe example assumes that the client device (e.g., HMD) used by the user has a resolution of 2160x1200. As shown, the size of the output frame (e.g., resolution and / or aspect ratio) changes as system latency changes (e.g., increases) (e.g., from a steady state to an unstable state). More specifically, the frame size (and therefore shape) of the output VR frame 2954, including the predicted overfilled frame, can have a resolution of 2400×1250 (and therefore a corresponding aspect ratio) during a stable system latency period. As system latency increases, each of the following output VR frames 2960 and 2966 will have different sizes as system latency increases and as the user turns his / her head in the yaw direction (as shown). For example, as shown, frame 2960 can have a resolution of 3000×1280 (and therefore a corresponding aspect ratio), while frame 2966 can have a different resolution of 4000×1300 (and therefore a corresponding aspect ratio). Furthermore, Figure 29B This illustrates how a combination of predicted FOV 2956 of T1 and predicted FOV 2958 of T2 can be used to configure each of frames 2954, 2960, and 2966 (where FOV 2956 and 2958 may include their respective error tolerances). According to some embodiments, the difference in the output VR frames based on the user's head rotation speed can have an effect similar to the latency stability of the (VR) system. For example, in some embodiments, user head movement can cause additional variations in the output VR frames.

[0222] Figure 30A This is a schematic diagram illustrating an example VR frame based on the user's head rotation direction. As shown, when the user's head rotation direction changes by 300°, it includes typical overfilling (e.g., as...). Figure 30A The frame size (and therefore shape) of the output VR frame 3002 (shown as an overfill factor of two (2)) remains the same (e.g., with a resolution of 4320x2400 as shown).

[0223] on the contrary, Figure 30B This is a schematic diagram illustrating the effect of the user's head rotation direction on the size of the output VR frame according to some embodiments. Figure 30BThe example assumes the user's client device (e.g., HMD) has a resolution of 2160x1200. As shown, the size of the output VR frame (e.g., resolution and / or aspect ratio) changes as the user's head rotation direction changes. More specifically, the frame size (and therefore shape) of the predicted overfilled output VR frame 3058 can have a resolution of 2200×2200 (and therefore a corresponding aspect ratio). As the head rotation direction changes, the following output VR frames 3060 and 3062 will each have different sizes as the head rotation further changes. For example, as shown, frame 3060 can have a resolution of 3000x2000 (and therefore a corresponding aspect ratio), while frame 3062 can then have a different resolution of 4000x1300 (and therefore a corresponding aspect ratio). Furthermore, Figure 30B This demonstrates how to configure each frame 3058, 3060, and 3062 using a combination of predicted FOV 3054 for T1 and predicted FOV 3056 for T2.

[0224] Variations of the above-described example methods and systems according to some embodiments will now be described.

[0225] In some embodiments, in addition to the parameters and content described above, additional parameters may be signaled from the VR content server to the client device. For example, in some embodiments, the client device may receive an indication from the server, prior to the scan output time, of at least one of the pixel dimensions or aspect ratio of the VR video frame (e.g., an overfilled, rendered VR video frame). In some embodiments, each VR video frame received from the server may include a rendering timestamp indicating the time the frame was rendered at the server. In some embodiments, each VR video frame received from the server may include additional timestamps, such as a decoding timestamp indicating the time the frame should be decoded at the client, and / or a playback timestamp indicating the time the frame should be scanned and output at the client device.

[0226] In some embodiments, the server may, for example, indirectly signal information indicating adaptively changing overfill regions. More specifically, the client device may learn the frame size of the received VR video frame during the decoding process of the (encoded) frame. The received frame will contain frame information (e.g., resolution, aspect ratio, etc.), which will typically be required for the decoding process. Therefore, in some embodiments, the client device may calculate an overfill factor for the received frame by comparing the size of the received frame with the size of the service frame generated at the client device. In some embodiments, the server may provide parameters indicating a predicted gaze point T or a set of gaze points {T1, T2} used to generate the rendered overfilled image. In some embodiments, the server may provide parameters indicating error tolerance or overfill regions within the rendered overfilled image. In some embodiments, the server may provide parameters indicating how the rendered overfilled image is aligned with the coordinate system used by the client to render the VR content to the user. For example, this coordinate system may be a spherical coordinate system.

[0227] In some embodiments, without considering the latency of the VR system, each of the overfill factors applied to each axis can be adjusted based on the user's head rotation speed and direction, for example, by multiplying the overfill weight factor by the head rotation speed in each axis direction.

[0228] In this context, in some embodiments, the system may not need to adapt to system latency, but it can still achieve a reduction in rendering overhead due to the different overfill factors applied to each axis. In some embodiments, the proposed solution utilizes information about the approximate render-to-scan output latency of the system, so it can use, for example, ping swapping or pre-configured values.

[0229] Figure 31 This is a message passing diagram illustrating an exemplary process 3100 according to some embodiments. This exemplary process can be executed between a VR content server 3102 and a user device (e.g., an HMD) 3104. In some embodiments, the exemplary process can be executed iteratively, such as, for example, on a frame-by-frame basis, or after a given number of frames.

[0230] As shown in the figure, in some embodiments, at 3106, a ping exchange (e.g., connection latency test) is performed between server 3102 and user device 3104. In some embodiments, the VR system may, for example, measure the connection latency between VR content server 3102 and user device 3104 before starting a service routine. In some embodiments, the VR system may use latency information determined by the ping exchange, and / or use a preset value based on the content server-user device connection type. In some embodiments, if the overfill weight factor becomes too high (e.g., relative to a given threshold) when serving content, another ping exchange may be performed to update the latency information. The reason for remeasuring latency when the overfill weight factor is high may be that an excessively high overfill weight factor indicates a relatively large difference between the expected latency and the actual latency.

[0231] In some embodiments, at 3108, VR content server 3102 can render VR video frames, such as overfilled frames. In this regard, in some embodiments, VR content server 3102 determines the target time point, for example, by summing the connection latency and the time spent rendering a frame. VR content server 3102 can then, for example, predict the user's gaze point at the target time point. VR content server 3102 can then apply an overfill factor to each axis to render the frame. VR content server 3102 can render the area around the predicted gaze point, taking into account the overfill factor set for each axis. The rendered frame can then be placed in a frame buffer.

[0232] In some embodiments, at 3110, user equipment 3104 may send user motion data (e.g., IMU-based data) to VR content server 3102. At 3112, rendered frames (e.g., service packets) may be sent to user equipment 3104. At 3114, user equipment 3104 may time-warp the received frames to generate service frames, and at 3116, the service frames are scanned and output. As described above, according to some embodiments, the service frames generally refer to frames that have been properly formatted or processed for display to a user via user equipment 3104. In this respect, frames received by user equipment 3104 from server 3102 may include overfilled images larger than the scanned output resolution. Therefore, user equipment 3104 selects and extracts data from appropriate portions of the received frames to generate the service frames. Additionally, user equipment 3104 may apply time warping to the received frames. Furthermore, in some embodiments, user equipment 3104 may determine and report FOV loss information. That is, in some embodiments, user device 3104 generates a service frame by performing a time warp based on the user's current FOV. The region in the created service frame where FOV loss occurs (after the time warp) is then measured and sent to VR content server 3104 at 3118. In some embodiments, this feedback information may be in the form of an FOV loss rate (e.g., ...). Figure 32 (as shown) and / or another form of information regarding overfill accuracy.

[0233] Right now, Figure 32 This is a view illustrating an example FOV loss region according to some embodiments. More specifically, Figure 32 The diagram illustrates the user's field of view (FOV) 3200, received frame 3202, and service frame 3204. As described above, in some embodiments, the received frame 3202 refers to a frame received at the user equipment before the user equipment performs time warping on that frame according to the user's current FOV 3200. Conversely, the service frame 3204 refers to a frame generated after time warping. Furthermore, Figure 32 The illustration shows a FOV loss 3206 (e.g., a 25% FOV loss) occurring in service frame 3204. Therefore, if the region corresponding to the user's current FOV 3200 is not included in the received frame 3202 (e.g., when the applied overfill factor is relatively small), the service frame 3204 generated using time warp may not meet the original frame size and will only contain data in some regions. In some embodiments, the FOV loss represents the percentage of the service frame 3204 (which is generated by the user equipment using the received frame 3202) that does not contain data. In some embodiments, when the region with the FOV loss is displayed to the user via the user equipment, this region is represented as a black space.

[0234] In some embodiments, at 3120, VR content server 3102 adjusts the overfill weight factor. In some embodiments, VR content server 3102 adjusts the overfill weight factor, for example, based on received FOV loss information (e.g., based on the FOV loss rate). In some embodiments, the content server may increase the overfill weight factor when it detects, for example, that the FOV loss has occurred in the service frame via feedback data, and may decrease the overfill weight factor when it detects, for example, that no FOV loss has occurred in a certain number of display frames (e.g., a certain number of consecutive display frames at user device 3104). Through this process, in some embodiments, the overfill weight factor can be controlled empirically and can have a stable weight factor in a stable latency environment. Note that in some embodiments, if the weight factor increases beyond a threshold due to repeated FOV loss, VR content server 3102 can remeasure the connection latency by performing a ping exchange again (see 3106) and can initialize the weight factor.

[0235] In some embodiments, where there is no FOV loss, such as for a specific number of frames (e.g., a specific number of consecutively displayed frames as described above), the overfill weight factor can be set to, for example, 0.9 (or some other value between 0 and 1), such that the new overfill factor is 0.9 of the previous overfill factor and is therefore reduced. For example, the overfill weight factor can be reduced from a set value (e.g., 0.9) and / or further reduced to another value (e.g., 0.8).

[0236] In some embodiments, the following example Equation 4 is used to determine the new overfill factor.

[0237]

[0238] In equation 4, the variable This refers to adjusting or modifying the intensity parameter. In some embodiments, Equation 4 provides a method for increasing the overfill factor in the event of FOV loss. According to this embodiment, the overfill factor is increased whenever the FOV loss rate is positive. A new overfill factor greater than or equal to the previous overfill factor is generated because the condition for modifying the overfill factor using the equation only occurs when FOV loss occurs. FOV loss indicates that the current overfill factor is insufficient for the current network latency and / or the dynamics of the VR user. In some embodiments, if the increased overfill factor eventually exceeds a certain threshold, the new overfill factor can be effectively reset based on the measured network latency. As mentioned above, if the overfill weight factor becomes too high (e.g., relative to a given threshold) when content is served, another latency test (e.g., ping exchange) can be performed to update the latency information, and the overfill factor can be initialized to a new value based on the result of the newly performed latency test. In some embodiments, for example, for large If FOV loss is detected in a traffic frame, the new overfill factor (e.g., determined from Equation 4) is greater than the previous overfill factor (e.g., the overfill factor determined for the frame immediately preceding the frame for which the new overfill factor is determined). In some embodiments, a larger one is used. This allows, for example, the system to dynamically adjust the overfill factor to, for example, respond quickly to FOV loss. In some embodiments, the... It can be selectively adjusted, for example, frame by frame.

[0239] If the overfill weighting factor and / or the overfill factor is due to repeated FOV

[0240] If the loss increases beyond the threshold, the content server can, for example, remeasure the connection latency by re-performing the ping exchange, and / or initialize the overfill factor.

[0241] In alternative embodiments, the overfill factor can be adjusted relative to a FOV loss rate threshold. In some implementations, Equation 5 can be used. According to Equation 5, if the FOV loss rate is higher than the value "thresh", the new overfill factor will be greater than the previous overfill factor. According to Equation 5, if the FOV loss rate is lower than the threshold, the new overfill factor will be less than the previous overfill factor, thus providing attenuation of the overfill factor when no FOV loss occurs. For example, in some implementations, Equation 5 can be used, for instance, to periodically reduce the overfill factor, at least in part based on the thresh term, even when the FOV loss rate is zero (0). This can, for example, allow for continuous attempts to minimize overfill rendering overhead.

[0242]

[0243] In some embodiments, at 3124, an overfill factor is set. Prior to this, at 3122, updated user actions may be received at the VR content server 3102. In some embodiments, the rotational speed of the user's head with respect to each axis may be identified. In some embodiments, the VR content server 3102 determines the overfill factor for each axis by multiplying a weighting factor by the head rotational speed. In other words, if the user rotates their head relatively quickly, for example, in the horizontal direction, the overfill factor in that horizontal direction will be relatively larger than the overfill factor in the vertical direction. By simultaneously or concurrently considering the overfill factor and the user's head rotational speed, some embodiments can also prevent FOV loss even if the actual scan output timing does not precisely match the target scan output timing (e.g., due to varying network latency).

[0244] Figure 33 The diagram illustrates the relationship between [various embodiments] and [other embodiments]. Figure 32 Example process 3300 related to the process. Figure 33 In some embodiments, it is shown that Figure 32 This is an example of how the process is performed in operation. At 3302 and 3304, the user device and the VR content server can participate in ping exchanges (as described in more detail earlier) to measure latency. Because ping signaling is carried in both the transmit (Tx) and receive (Rx) directions, network latency 3306 is doubled (or as...). Figure 33 (twice as in). At 3308, the user equipment can provide IMU data to the server. The server can then render VR video frames (e.g., the first overfilled frame) based on the target time 3314.

[0245] As previously stated, the target time 3314 may include processing latency 3310 (e.g., the time spent rendering a frame) and network latency 3312 (e.g., half of latency 3306) between the server and the user equipment. In some embodiments, the target time 3314 may include additional time for other predictable latency, such as decoding and processing of the service frame at the user equipment (e.g., the client device). The target time 3314 may include any latency expected to occur between rendering at the server and scanning output at the client (e.g., any latency included in the rendering-to-scanning-output delay). At 3316, the rendered frame is transmitted to the user equipment to substantially meet, for example, the target time point. At 3318, the user equipment may perform a time warping process on the frame received from the server to generate a service frame. At 3322, the user equipment scans and outputs the service frame 3324 to display to the user. Furthermore, as previously stated, the user equipment may detect whether any FOV loss has occurred. Figure 33 The example assumes that such a loss has occurred, and at 3320, the user equipment sends FOV loss information to the server. Furthermore, at 3326, the user equipment may also provide the server with the latest IMU data. On the server side, at 3330, the server adjusts the overfill weight factor based on the FOV loss information and sets the overfill factor for each axis accordingly. Therefore, in some embodiments, the adjusted overfill weight factor can be used for one or more subsequent frames (e.g., the next frame) rendered by the server.

[0246] Examples of use cases based on, for example, the embodiments described above will now be described.

[0247] In some embodiments, taking into account that the user's movement will change while providing VR services, the VR systems described herein apply overfilling and time warping. According to some embodiments, predictive overfilling can reduce overhead and can be applied to systems with varying latency.

[0248] Figure 34 A schematic diagram 3400 is shown, illustrating an example of changing the rendering-to-scan output delay according to some embodiments. Figure 34The example can be applied to scenarios involving remote VR services between a cloud VR server and an HMD. As shown in the figure, at 3402, the server can begin rendering a service frame. The delay 3410 from the start of rendering to the actual transmission time can include a varying processing delay 3408. At 3412, the rendered frame can be transmitted. Due to, for example, varying network latency, the arrival time of the transmitted frame at the HMD can be expected such that the scan output time of the frame will occur somewhere between T1 and T2 (e.g., time interval 3414), where in some embodiments, T1 and T2 respectively provide a lower limit and an upper limit for the expected scan output time. Furthermore, Figure 34 The user's head movements 3416 at time T1 and time T2 are shown as the user's head movements (represented by 3418) may change.

[0249] An exemplary process for calculating time intervals [T1, T2] to determine the predicted overfilled rendering area of ​​the user's FOV will now be described.

[0250] Such as combination Figure 12 To select two predictive gaze points corresponding to T1 and T2 for predictive rendering, the actual delay in serving the corresponding service frame from the server to the user device can be obtained or determined. To know the system's render-to-scan-output delay, in some embodiments, the server records the render start time information for some frames or each frame and compares this render start time information with the actual service time of the corresponding frame (e.g., the scan-to-output time), which can be identified by feedback data from the user device (e.g., HMD). As described herein, the recorded render start time and the corresponding associated service time (or scan-to-scan-output time) can be used to construct the render-to-scan-output delay distribution.

[0251] Because the rendering-to-scan output latency may vary over time (e.g., due to processor temperature, the complexity of the scene to be rendered, the amount of overfilling, and / or network latency), the latency distribution table can be managed by swapping “old” data entries with “new” data entries. Figure 35 This is a schematic diagram illustrating an exemplary delay table update according to some embodiments. For example... Figure 35 As shown, in some embodiments, if the most recently measured render-to-scan output latency data 3502 (e.g., 83.2 ms) is measured, the oldest data 3504 (e.g., 76.5 ms) in table 3500 is deleted, and the corresponding most recently measured data 3502 is added to table 3500 to update it. Since the data in table 3500 only covers the latency of a specific number of frames (e.g., 100 measurements), predictive overfilling according to some embodiments disclosed herein can adaptively set the appropriate rendering foveation as the system environment changes.

[0252] In some embodiments, the server can determine and / or examine the latency distribution using data stored in a latency table. If the data in the latency table has been updated, the probabilities will also change. Figure 36 An exemplary delay distribution probability table 3600 according to some embodiments is shown. As illustrated by way of example, if the delay distribution is calculated at 2ms intervals, then... Figure 36 As shown, the probability of each delay interval can be determined. According to Table 3600, if the stability of the predicted overfill is set to 60% for the system, then time T1 corresponding to a delay value of 78ms and time T2 corresponding to a delay value of 88ms can be selected.

[0253] The fixation points are the predicted user FOV points corresponding to previously determined T1 (78ms) and T2 (88ms). The method used to calculate the two fixation points corresponding to times T1 and T2 can utilize CAP or CRP as methods for predicting user FOV, as done in the example VR system. The VR server can utilize IMU information from the HMD's feedback data to examine the user's current FOV and movement characteristics. The CAP or CRP technique can be applied to determine the fixation points at the predicted scan output times T1 and T2, for example, as... Figure 36 As shown.

[0254] As mentioned above, T2 represents a more distant future prediction time than T1, therefore the prediction for T2 (88ms) is more uncertain than the prediction for T1 (78ms). In some embodiments, the VR content server performs rendering over a larger area around the gaze point for T2 compared to the gaze point for T1. If the user's head rotation speed is, for example, 100 degrees / second, then the distance between the gaze points for T1 and T2 is 1 degree (e.g., using CRP). Figure 37A The illustration shows a similar approach to two specific time values ​​according to some embodiments. Figure 16 Example Figure 3700. Figure 37B This illustrates, according to some embodiments, two specific example time values ​​(i.e., a first time value with T1 = 78 ms and a second time value with T2 = 88 ms) from Figure 17A The example MSE of orientation error versus expected time is shown in Figure 3750.

[0255] Figure 38 Potential FOVs according to some embodiments are shown. Figure 38A portion of a virtual world 3800 is shown, including a user gaze point 3802 predicted for T1 and a user gaze point 3804 predicted for T2. As described in more detail above, the predicted user gaze point 3802 may have a corresponding prediction error range 3806, and the predicted user gaze point 3804 may have a corresponding prediction error range 3808. Furthermore, Figure 38 The diagram illustrates a predicted FOV 3810 adjusted for T1 (a first predicted FOV with an added error tolerance (e.g., 8.8 degrees)) and a predicted FOV 3812 adjusted for T2 (a second predicted FOV with an added error tolerance (e.g., 11.7 degrees)). In some embodiments, a potential FOV 3814 (e.g., a combined FOV, such as the target rendering area) is determined by merging the predicted FOVs 3810 and 3812 adjusted for T1 and T2. Furthermore, as... Figure 38 As shown, the location information of the center point of the potential FOV 3814 (as shown in the figure) can be used as information for rendering.

[0256] Figure 39 This is a schematic diagram illustrating the relationship between the potential FOV and the overfill factor according to some embodiments. The potential FOV 3900 (e.g., the target rendering area) includes the service FOV 3902 of the HMD. The overfill factor for each corresponding axis (i.e., the x-axis (horizontal direction) and the y-axis (vertical direction)) can be calculated according to the following example equations 6 and 7:

[0257] Overfill factor (level) = R_width / S_width Equation 6

[0258] Overfill factor (vertical) = R_height / S_height Equation 7

[0259] Figure 40 This is a flowchart of an example method 4000 according to some embodiments. In some embodiments, the processing may be performed by a server. In step 4002, the server receives head tracking position information from a client device, which is associated with a user at the client device. In step 4004, the server predicts the user's future head position at the time of scan output for displaying virtual reality (VR) video frames, which are displayed to the user via the client device. In step 4006, the server determines an overfill factor based on the expected error in the user's predicted future head position. In step 4008, the server renders an overfilled image based on the user's predicted future head position and the overfill factor. Then, in step 4010, the server sends the VR video frame, including the overfilled image, to the client device for display to the user.

[0260] Figure 41 This is a flowchart of another example method 4100 according to some embodiments. In some embodiments, this process may be performed by a server. In step 4102, the server determines that a loss in the field of view (FOV) of a virtual reality (VR) frame sent to the client device has occurred. Then, in step 4104, the server adaptively adjusts the overfill factor weights based on the determination that the loss in the FOV has occurred.

[0261] Although examples of various features of the example methods and systems have been described regarding servers and execution by servers (see, for example, see...), Figure 11 , 12 (24 and 31), but this description and these example methods are not limited to such implementations, and various features can be performed by, for example, a client device (e.g., an HMD). Furthermore, in any step or embodiment involving a server, the server can be a cloud-based server communicating with the client (user) device via a network, or the server can be a local computer communicating with the client (user) device using a local network, wireless communication protocols, and / or a wired connection.

[0262] For example, Figure 42 This is a flowchart of another example method 4200 according to some embodiments. In some embodiments, this process may be performed by a client device. In step 4202, the client device receives a first virtual reality (VR) video frame from a server. In response to receiving the first VR video frame, in step 4204, the client device sends timing information to the server, wherein the timing information includes at least the scan output start time of the received first VR video frame. In step 4206, the client device sends motion information of the user of the VR client device to the server. In step 4208, the client device receives a second VR video frame from the server, wherein the second VR video frame includes an overfilled image based on (i) the user's predicted head position at the scan output time when the second VR video frame is displayed to the user, and (ii) an overfill factor. Then, in step 4210, the client device displays a selected portion of the overfilled image to the user, wherein the selected portion is based on the user's actual head position at the scan output time of the second VR video frame, and wherein the predicted head position is based on the transmitted motion information of the user, and the overfill factor is based on the expected error in the user's predicted head position.

[0263] Figure 43Example computing entity 4300 that can be used in embodiments of this disclosure is depicted, for example, as a local or remote VR content server, as part of such a VR content server, or as part of multiple entities that can together perform as a VR service system. Figure 43 As depicted, the computing entity 4300 includes a communication interface 4302, a processor 4304, and a non-transitory data storage 4306, all of which are communicatively linked via a bus, network, or other communication path 4308.

[0264] Communication interface 4302 may include one or more wired communication interfaces and / or one or more wireless communication interfaces. Regarding wired communication, as an example, communication interface 4302 may include one or more interfaces, such as an Ethernet interface. Regarding wireless communication, communication interface 4302 may include multiple components, such as one or more antennas, one or more transceivers / chipsets designed and configured for one or more types of wireless (e.g., LTE) communication, and / or any other components deemed suitable by those skilled in the art. Furthermore, regarding wireless communication, communication interface 4302 may be proportionally configured and have a configuration suitable for operation on the network side (as opposed to the client side) of wireless communication (e.g., LTE communication, WiFi communication, etc.). Therefore, communication interface 4302 may include appropriate equipment and circuitry (potentially including multiple transceivers) for multiple mobile stations, UEs, or other access terminals in a service coverage area.

[0265] Processor 4304 may include one or more processors of any type that a person skilled in the art would deem appropriate, some examples of which include general-purpose microprocessors and dedicated DSPs.

[0266] Data storage 4306 may take the form of any non-transitory computer-readable medium or a combination of such media. Some examples include flash memory, read-only memory (ROM), and random access memory (RAM), to name just a few, as any one or more types of non-transitory data storage may be used as deemed appropriate by a person skilled in the art. Figure 43 As depicted herein, according to some embodiments, data memory 4306 includes program instructions 4310 that can be executed by processor 4304 to implement various combinations of the various functions described herein.

[0267] In addition, various (e.g., related) embodiments have been described above.

[0268] According to some embodiments, a method for rendering an overfilled image based on predicted HMD location and predicted scan output time is disclosed, wherein the overfill factor is adaptively calculated.

[0269] According to some embodiments, a method for adaptively performing overfilling based on observed user head movement is disclosed.

[0270] According to some embodiments, a method for performing overfilling based on the observed distribution of network latency is disclosed.

[0271] According to some embodiments, a method for rendering an overfilled image based on two predicted scan output times (e.g., interval [T1, T2]) is disclosed.

[0272] According to some embodiments, a method for adjusting the overfill factor using FOV loss information and connection latency testing via ping switching is disclosed. In some embodiments, using the disclosed method, actual scan output time may not be required.

[0273] According to some implementations, a method for rendering an image to a VR user may include: determining a time T, where T is a predicted scan output time for displaying the current frame at a client device; predicting the HMD position and / or associated field of view (FOV) at time T, at least in part based on IMU motion data of the client device; determining an overfill factor based on at least one of: an expected error or confidence level of the predicted scan output time T, an expected error or confidence level of the predicted HMD position, or a recent measurement of the user's head motion; rendering an overfilled image based on the predicted HMD position and the overfill factor; and displaying a portion of the overfilled image to the VR user at the client device, the portion being determined based on the position of the HMD at the scan output time.

[0274] In some embodiments, the time T is determined based on the amount of network latency observed between the server and the client device. In some embodiments, the time T is determined based on the amount of prediction time required to render the current frame at the client device. In some embodiments, the method may further include constructing a render-to-scan output latency distribution, wherein the time T is determined based on this latency distribution. Furthermore, in some embodiments, rendering the overfilled image based on the predicted HMD location and the overfill factor may include: rendering a first overfilled image for a first time T1 using a first overfill factor; and rendering a second overfilled image for a second time T2 using a second overfill factor different from the first overfill factor.

[0275] According to some implementations, a method for rendering an image to a VR user may include: constructing a rendering-to-scan-output delay distribution; selecting times T1 and T2 based on the delay distribution, wherein T1 provides a lower limit for the expected scan-output time and T2 provides an upper limit for the expected scan-output time; predicting the HMD position and / or associated FOV for each of times T1 and T2 based on IMU motion data from a client device; determining an overfill factor for each of times T1 and T2; rendering an overfilled image based on the predicted HMD position and the overfill factor for time T1 and the predicted HMD position and the overfill factor for time T2; and displaying a portion of the overfilled image to the VR user at the client device, the portion being determined based on the position of the HMD at the scan-output time.

[0276] In some embodiments, the first overfilled image generated for a first scan output time has a different resolution than the second overfilled image generated for a second scan output time. In some embodiments, the first overfilled image generated during the first scan output time has a different aspect ratio than the second overfilled image generated during the second scan output time. In some embodiments, the overfill factor for each of times T1 and T2 is determined based on the amount of expected HMD position prediction error or confidence at HMD position prediction at times T1 and T2, respectively. Furthermore, in some embodiments, the rendering-to-scan output delay distribution is constructed at least in part based on the rendering start time (Tr) of the start rendering frame. R ) and the scan output time (T) used to display the frame at the client device S The difference between them.

[0277] According to some embodiments, a method may include: receiving image scan output time data from a head-mounted display (HMD); determining a render-to-scan output delay distribution based at least in part on the image scan output time data; receiving motion data from the HMD; determining expected image scan output time data based at least in part on the render-to-scan output delay distribution; estimating a first field of view at a first time and a second field of view at a second time based at least in part on the motion data, the first time and the second time corresponding to at least corresponding portions of the expected scan output time data; determining an overfill factor for each of the first time and the second time; rendering a first overfilled image based on the first field of view, the second field of view, and the corresponding overfill factor; and sending the first overfilled image to the HMD.

[0278] In some embodiments, estimating a first field of view for the first time and a second field of view for the second time, based at least in part on the motion data, may include estimating a first fixation point for the first time and a second fixation point for the second time. In some embodiments, estimating the first fixation point for the first time and the second fixation point for the second time may include performing a constant rate (velocity) prediction (CRP). In some embodiments, estimating the first fixation point for the first time and the second fixation point for the second time may include performing a constant acceleration prediction (CAP).

[0279] Furthermore, in some embodiments, rendering the overfilled image based on the first field of view, the second field of view, and the corresponding overfill factor may include: adjusting the first field of view based on the overfill factor at the first time; and adjusting the second field of view based on the overfill factor at the second time. In some embodiments, the method may further include: receiving newer image scan output time data from the HMD; and updating the rendering-to-scan output delay distribution based on the newer image scan output time data.

[0280] Furthermore, in some embodiments, the method may further include: rendering a second overfilled image; and transmitting the second overfilled image to the HMD, wherein the image scan output time data received from the HMD includes the second overfilled image scan output time data. In some embodiments, the method may further include time warping the first overfilled image. The method may also further include displaying the first overfilled image at the HMD.

[0281] Furthermore, in some embodiments, the method may further include: recording the start time for rendering the first overfilled image. In some embodiments, the method may further include determining the resolution display capability of the HMD; determining a prediction error for the first time; determining a prediction error for the second time; and comparing the resolution display capability with each of the prediction errors. In some embodiments, the first overfilled image is rendered as having a substantially rectangular shape. In some embodiments, the first overfilled image is rendered as having a substantially hexagonal shape.

[0282] Additionally, in some embodiments, the first overfilled image is rendered as having a substantially rectangular shape. In some embodiments, the overfill factor at the first time may include different values ​​for the horizontal and vertical dimensions, and the overfill factor at the second time may include different values ​​for the horizontal and vertical dimensions.

[0283] According to some embodiments, a method may include rendering a plurality of overfilled images, each overfilled image being overfilled based on user head movement information and a plurality of render-to-scan output delay measurements, and having error tolerances for at least horizontal and vertical resolutions.

[0284] According to some embodiments, a system may include a virtual reality content server configured to predictively overfill a video frame and send the predictively overfilled video frame to a virtual reality (VR) device.

[0285] According to some embodiments, a method may include: receiving motion data from a head-mounted display (HMD); predicting a first future orientation of the HMD at a first time based at least in part on the motion data; predicting a second future orientation of the HMD at a second time based at least in part on the motion data; estimating a first prediction error corresponding to the first future orientation and a second prediction error corresponding to the second future orientation; determining a first predicted field of view (FOV) corresponding to the first future orientation based at least in part on the first prediction error; determining a second predicted FOV corresponding to the second future orientation based at least in part on the second prediction error; determining a potential FOV, the potential FOV being dimensionalized at least in part based on the first predicted FOV and the second predicted FOV; and rendering an overfilled image based at least in part on the potential FOV.

[0286] According to some embodiments, a method may include: adjusting a first overfill factor for the first axis based at least in part on head-mounted display (HMD) motion data corresponding to a first axis of the image and render-to-scan output delay data; adjusting a second overfill factor for the second axis based at least in part on HMD motion data corresponding to a second axis of the image and the render-to-scan output delay data; and rendering an overfilled image based on the first overfill factor and the second overfill factor.

[0287] According to some embodiments, a system may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed on the processor, operate to perform any of the methods disclosed herein.

[0288] Note that the various hardware elements of the one or more embodiments described are referred to as “modules,” and their implementation (i.e., execution, operation, etc.) is in conjunction with the various functions described herein for each module. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include instructions executable to perform one or more functions described as being performed by the corresponding module, and note that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, and / or software instructions, or include them, and may be stored in any suitable non-transitory computer-readable medium or medium, such as media or media commonly referred to as RAM, ROM, etc.

[0289] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in computer programs, software, or firmware embedded in a computer-readable medium and executed by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multipurpose discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method performed at a server, comprising: receiving head-tracked position information from a client device, the head-tracked position information associated with a user at the client device; predicting a future head position of the user at a scan-out time for displaying a VR video frame, wherein the VR video frame is displayed to the user via the client device; determining an overfill factor based on an expected error in the predicted future head position of the user; rendering an overfilled image based on the predicted future head position of the user and the overfill factor; and sending the VR video frame including the overfilled image to the client device for display to the user, wherein the expected error in the predicted future head position of the user is calculated based on a latency profile, the latency profile including latency metric data for a plurality of VR video frames that were previously scan-out for display by the client device.

2. The method of claim 1, wherein the client device comprises a head-mounted display (HMD).

3. The method of any one of claims 1-2, wherein the expected error in the predicted future head position of the user is based on an observation of the head-tracked position information received by the server.

4. The method of any one of claims 1-2, wherein the expected error in the predicted future head position of the user is based on an observation of network latency over time. rendering at least one other VR video frame containing another overfilled image, wherein a size of the at least one other VR video frame is different than a size of the VR video frame containing the overfilled image.

5. The method of any one of claims 1-2, further comprising:

6. The method of claim 5, wherein at least one of a pixel size or an aspect ratio is dynamically changed from one VR video frame to another VR video frame according to changes in latency of a connection between the server and the client device.

7. The method of claim 5, wherein at least one of a pixel size or an aspect ratio is dynamically changed from one VR video frame to another VR video frame according to changes in head rotation of the user.

8. The method of any one of claims 1-2, wherein predicting the future head position of the user at the scan-out time comprises: predicting the future head position of the user at the scan-out time using, at least in part, the head-tracked position information.

9. The method of claim 8, wherein the head-tracked position information is based on IMU-based motion data.

10. The method of claim 8, wherein predicting the future head position of the user at the scan-out time using, at least in part, the head-tracked position information comprises: predicting a first FOV at a first time Tl, wherein a first predicted FOV is based on a predicted first gaze point; and predicting a second FOV at a second time T2, wherein a second predicted FOV is based on a predicted second gaze point. ​ ​ 11. The method of claim 10, further comprising: selecting Tl and T2 based on the latency profile.

12. The method of claim 11, wherein Tl and T2 are selected such that a time interval between Tl and T2 contains a target probability with respect to the latency profile.

13. The method of claim 10, wherein Tl provides a lower bound of an expected scan-out time of the VR video frame and T2 provides an upper bound of an expected scan-out time of the VR video frame.

14. The method of claim 10, further comprising: adding a first error margin associated with the predicted first gaze point to the first predicted FOV; and adding a second error margin associated with the predicted second gaze point to the second predicted FOV.

15. The method of claim 14, wherein the second error margin is greater than the first error margin.

16. The method of claim 14, wherein each of the first error margin and the second error margin associated with the predicted first and second gaze points, respectively, is based on a prediction technique selected from the group consisting of: constant rate based prediction (CRP) and constant acceleration based prediction (CAP). confirming values of the first error margin and the second error margin in real-time based on the head tracking position information received from the client device.

17. The method of claim 14, further comprising:

18. The method of claim 17, wherein the first error margin and the second error margin are based on an error between the head tracking position information and the predicted motion data. setting the overfill factor based at least in part on the first error margin and the second error margin.

19. The method of claim 14, wherein determining the overfill factor based on the expected error in the predicted future head position of the user comprises:

20. The method of claim 19, wherein the overfill factor comprises a first overfill factor value for a horizontal axis and a second overfill factor value for a vertical axis, the first and second overfill factor values being different from each other. determining a combined FOV associated with the overfilled image based on the first predicted FOV and the second predicted FOV.

21. The method of claim 14, further comprising: combining (i) a first adjusted predicted FOV and (ii) a second adjusted predicted FOV, wherein the first adjusted predicted FOV is determined by adding the first error margin to the first predicted FOV and the second adjusted predicted FOV is determined by adding the second error margin to the second predicted FOV.

22. The method of claim 21, wherein determining the combined FOV based on the first predicted FOV and the second predicted FOV comprises:

23. The method of claim 22, wherein the combined FOV is determined by selecting a rectangular region that includes the first adjusted predicted FOV and the second adjusted predicted FOV.

24. The method of claim 22, wherein the combined FOV is determined by selecting a hexagonal region that includes the first adjusted predicted FOV and the second adjusted predicted FOV. applying the overfill factor with respect to a center point of the combined FOV.

25. The method of claim 21, wherein rendering the overfilled image comprises:

26. The method of claim 8, further comprising: ​ determining a time T, wherein the time T represents a predicted scan-out time of the VR video frame containing the over-filled image.

27. The method of claim 26, wherein the time T is predicted based on the latency distribution.

28. The method of claim 26, wherein the time T corresponds to a mean or median value of the latency distribution.

29. The method of claim 26, further comprising: predicting a field of view (FOV) of the user at the time T based on the future head position of the user.

30. The method of claim 29, further comprising: determining an extended FOV of the user at the time T based on a direction and speed of head rotation of the user; and aligning a center position of the extended FOV with a center position of the FOV to produce a final extended FOV.

31. The method of claim 30, wherein rendering the overfilled image comprises: applying the over-fill factor to the final extended FOV, the over-fill factor having a first over-fill factor value for a horizontal axis and a second over-fill factor value for a vertical axis, wherein the first and second over-fill factor values are different from each other.

32. A method performed at a server, comprising: receiving head tracking position information from a client device, the head tracking position information associated with a user at the client device; predicting a future head position of the user at a scan-out time for displaying a virtual reality (VR) video frame, wherein the VR video frame is displayed to the user via the client device, and wherein predicting a future head position of the user at the scan-out time comprises predicting a future head position of the user at the scan-out time using, at least in part, the head tracking position information; determining an over-fill factor based on an expected error in the predicted future head position of the user; rendering an over-filled image based on the predicted future head position of the user and the over-fill factor; sending the VR video frame including the over-filled image to the client device for display to the user; in response to receiving the VR video frame including the over-filled image at the client device, receiving timing information from the client device, the timing information including at least a scan-out start time of the VR video frame containing the over-filled image; and determining a render-to-scan-out latency distribution, wherein determining the render-to-scan-out latency distribution comprises determining, at least in part, a difference between a render start time of the VR video frame and the scan-out start time of the VR video frame to calculate a render-to-scan-out latency value for the VR video frame.

33. The method of claim 32, further comprising: adding the render-to-scan-out latency value for the VR video frame to a table configured to hold a plurality of render-to-scan-out latency values associated with rendered VR video frames; and using the table to determine the render-to-scan-out latency distribution.

34. The method of claim 33, wherein each time a new render-to-scanout latency value is added to the latency table, an older render-to-scanout latency value is removed from the latency table.

35. A method performed by a VR client device, comprising: receiving a first VR video frame from a server; in response to receiving the first VR video frame, sending timing information to the server, wherein the timing information includes at least a scanout start time of the first VR video frame; sending motion information of a user of the VR client device to the server; receiving a second VR video frame from the server, wherein the second VR video frame contains an overfill image, the overfill image being based on (i) a predicted head position of the user at a scanout time of the second VR video frame for display to the user and (ii) an overfill factor, displaying a selected portion of the overfill image to the user, wherein the portion is selected based on an actual head position of the user at the scanout time of the second VR video frame, wherein the predicted head position is based on the motion information of the user, and the overfill factor is based on an expected error in the predicted head position of the user, and wherein the expected error in the predicted future head position of the user is calculated based on a latency distribution, the latency distribution including latency metric data for a plurality of VR video frames that were previously scanned out for display by the client device.

36. The method of claim 35, wherein the motion information comprises IMU-based motion data.

37. The method of any one of claims 35-36, wherein the client device comprises a head-mounted display (HMD).

38. The method of any one of claims 35-36, wherein a frame size of the second VR video frame comprising the overfill image is different than a frame size of the first VR video frame.

39. The method of claim 38, wherein an aspect ratio of the received second VR video frame comprising the overfill image is different than an aspect ratio of the received first VR video frame.

40. The method of claim 38, wherein at least one of a pixel size or an aspect ratio of the received second VR video frame comprising the overfill image is different than at least one of a pixel size or an aspect ratio of the received first VR video frame due to a change in a connection latency between the client device and the server.

41. The method of claim 38, wherein at least one of a pixel size or an aspect ratio of the received second VR video frame comprising the overfill image is different than at least one of a pixel size or an aspect ratio of the received first VR video frame due to a change in a head rotation of the user.

42. The method of any one of claims 40-41, further comprising: receiving an indication from the server regarding at least one of the pixel size or the aspect ratio prior to the scanout time.

43. The method of claim 38, wherein each of the received first and second VR video frames includes a respective timestamp indicating a frame rendering time at the server.

44. The method of any one of claims 35-36, further comprising: time warping at least one of the first or second VR video frames.

45. The method of any one of claims 35-36, further comprising: tracking an actual head position of the user.

46. A system for VR, comprising, a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the system to perform the method of any of claims 1-31.

47. A system for VR, comprising, a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the system to perform the method of any of claims 32-34.

48. A system for VR, comprising, a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the system to perform the method of any of claims 35-45.

49. A server, comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the server to: receive head tracking position information from a client device, the head tracking position information associated with a user at the client device; predict a future head position of the user at a scan-out time for displaying a VR video frame, wherein the VR video frame is displayed to the user via the client device; determine an overfill factor based on an expected error in the predicted future head position of the user; render an overfilled image based on the predicted future head position of the user and the overfill factor; and send the VR video frame including the overfilled image to the client device for display to the user, wherein the expected error in the predicted future head position of the user is calculated based on a latency profile, the latency profile including latency metric data for a plurality of VR video frames previously scanned out for display by the client device.

50. A VR client device, comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the VR client device to: receive a first VR video frame from a server; in response to receiving the first VR video frame, send timing information to the server, wherein the timing information includes at least a scan-out start time for the first VR video frame; send motion information of a user of the VR client device to the server; receiving a second VR video frame from the server, wherein the second VR video frame includes an overfill image, the overfill image based on (i) a predicted head position of the user at a scan-out time of the second VR video frame for display to the user and (ii) an overfill factor; and displaying a portion of the overfill image to the user based on an actual head position of the user at the scan-out time of the second VR video frame, wherein the predicted head position is based on the motion information of the user and the overfill factor is based on an expected error in the predicted head position of the user, and wherein the expected error in the predicted future head position of the user is calculated based on a latency distribution, the latency distribution including latency metric data for a plurality of VR video frames previously scanned out for display by the client device.

51. The VR client device of claim 50, wherein the VR client device comprises a head-mounted display (HMD).

Citation Information

Patent Citations

  • Remote rendering for virtual images

    US20170115488A1

  • Optimized Display Image Rendering

    US20180047332A1